Methods and compositions relating to ecdna biogenesis
The development of linear reporter polynucleotides as biosensors addresses the limitations of current ecDNA biogenesis studies, enabling efficient detection and inhibition of ecDNA formation, thereby providing a therapeutic strategy to target cancer cell adaptation and drug resistance.
Patent Information
- Application Number
- PCT/US2024/023815
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-10
- Publication Date
- 2025-10-16
AI Technical Summary
Current methods for studying extrachromosomal circular DNA (ecDNA) biogenesis are limited, and there is a need for improved mechanisms to identify therapeutic targets that can alter ecDNA biogenesis, particularly in aggressive cancers, as ecDNA presence thwarts durable cancer treatment and leads to shorter patient survival.
Development of linear reporter polynucleotides that function as biosensors for ecDNA biogenesis, which do not require pre-integration into the host genome and utilize high transfection efficiency to detect and inhibit ecDNA formation, allowing identification of factors and compounds that regulate ecDNA biogenesis.
The biosensors provide a reliable and efficient method for detecting ecDNA formation and identifying factors that regulate its biogenesis, offering a novel therapeutic approach to inhibit ecDNA in cancer cells, potentially suppressing cancer cell adaptation and drug resistance.
Smart Images

Figure US2024023815_16102025_PF_FP_ABST
Abstract
Description
METHODS AND COMPOSITIONS RELATING TO ECDNABIOGENESISSEQUENCE LISTING
[0001] The official copy of the sequence listing is submitted electronically via Patent Center as an XML formatted sequence listing with a file named DU8048PCT_1427979_SL.xml, created on April 10, 2024, and having a size of 73 kb. The sequence listing contained in this XML formatted document is part of the specification and is herein incorporated by reference in its entirety.BACKGROUND
[0002] Extrachromosomal circular DNA (ecDNA) is commonly produced within the nucleus to drive genome dynamics and heterogeneity. Notably, ecDNA serves as a common form of massive oncogene amplification, which enables cancer cells to rapidly adapt and evolve and is particularly common in certain aggressive cancer types (for example, glioblastoma, sarcoma, and esophageal cancers). The presence of ecDNA in cancer cells, in particular cells related to aggressive cancer types, thwart the efforts to achieve durable cancer treatment. Consequently, patients whose cancers harbor ecDNA have significantly shorter survival compared to patients with cancers that do not harbor ecDNA.
[0003] Despite the importance of ecDNA biogenesis in propelling cancer initiation and progression, little is known about the mechanisms behind ecDNA biogenesis. There is a need for improved mechanisms for studying ecDNA biogenesis, which can help identify therapeutic targets that may alter ecDNA biogenesis. Furthermore, targeting molecular drivers of ecDNA biogenesis can suppress ecDNA biogenesis in cancer cells that can harbor ecDNA, particularly in aggressive cancers, providing a promising avenue for novel and effective cancer therapeutics.SUMMARY
[0004] This summary provides a high-level ovendew of various aspects of the disclosure and introduces some of the concepts that are described and illustrated in the present document and the accompanying figures. The summary is not intended to identify key or essential features ofthe claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. Covered embodiments of the disclosure are defined by the claims, not this summary. The subject matter should be understood by reference to appropriate portions of the entire specification, any or all figures, and each claim. Some of the exemplary embodiments of the present disclosure are discussed below.
[0005] In one aspect, provided is a method of detecting circular DNA formation in a cell, the method comprising: (i) contacting a plurality of cells with a reporter polynucleotide, (ii) culturing the plurality of cells comprising the polynucleotide, and (iii) detecting a first signal from the first reporter polypeptide. In some embodiments, detection of the first signal indicates circular DNA formation in the plurality' of cells.
[0006] In some embodiments, the reporter polynucleotide comprises a first nucleotide sequence encoding a first reporter polypeptide followed by a first promoter, wherein circularization of the linear polynucleotide leads to the first promoter being operably linked to the first nucleotide sequence.
[0007] In some embodiments, the first reporter polypeptide comprises a fluorescent protein. In some embodiments, the reporter polynucleotide comprises a nucleotide sequence of at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 1. In some embodiments, the nucleotide sequence comprises the first nucleotide sequencing encoding the first reporter polypeptide followed by the first promoter element.
[0008] In some embodiments, the reporter polynucleotide further comprises a second promoter operably linked to a second nucleotide sequence encoding a second reporter polypeptide, wherein the second promoter and the second nucleotide sequence are positioned in the linear polynucleotide between the first nucleotide sequence and the first promoter.
[0009] In some embodiments, the first reporter polypeptide and the second reporter polypeptide are different proteins. In some embodiments, the first promoter and the second promoter each comprise a eukaryotic promoter nucleotide sequence. In some embodiments, the first promoter element and the second promoter element are different.
[0010] In some embodiments, the method further comprises detecting a second signal from the second reporter polypeptide, wherein detection of the second signal from the second reporter polypeptide indicates reporter polynucleotide presence in the plurality7of cells.
[0011] In some embodiments, the reporter polynucleotide further comprises a third nucleotide sequence encoding a selection marker, wherein the third nucleotide sequence encoding the selection marker is operably linked to the second promoter element. In some embodiments, the reporter polynucleotide comprises a nucleotide sequence comprising at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 2.
[0012] In some embodiments, the nucleotide sequence, from its 5 '-end to its 3 ’-end, comprises (i) the first nucleotide sequence encoding the first reporter polypeptide, (ii) the second promoter element, (iii) the second nucleotide sequence encoding the second reporter polypeptide, (iv) the third nucleotide sequence encoding the selection marker, and, (v) the first promoter element. In some embodiments, the second nucleotide sequence encoding the second reporter polypeptide and the third nucleotide sequence encoding the selection marker are operably linked to the second promoter element.
[0013] In some embodiments, the method further comprises culturing the plurality of cells with a compound that corresponds to a selectable marker.
[0014] In some embodiments, the method further comprises (i) culturing the plurality7of cells with a compound of interest, and (ii) detecting a third signal from the first reporter polypeptide. In some embodiments, the compound of interest decreases circular DNA presence if the third signal from the first reporter polypeptide is less than the first signal from the first reporter polypeptide. In some embodiments, the the plurality of cells is cultured with the compound of interest before the plurality of cells is contacted with the reporter polynucleotide.
[0015] In some embodiments, the first signal from the first reporter polypeptide, second signal from the second reporter polypeptide, and / or third signal from the first reporter polypeptide are fluorescence signals.
[0016] In another aspect, provided is a linear reporter polynucleotide comprising a first nucleotide sequence encoding a first reporter polypeptide followed by a first promoter. In some embodiments, circularization of the linear reporter polynucleotide leads to the first promoter being operably linked to the first nucleotide sequence.
[0017] In some embodiments, the first reporter polypeptide comprises a fluorescent protein. In some embodiments, the linear reporter polynucleotide comprises a nucleotide sequence of at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%,at least 99%, or 100% sequence identity with SEQ ID NO: 1. In some embodiments, the nucleotide sequence comprises the first nucleotide sequencing encoding the first reporter polypeptide followed by the first promoter element.
[0018] In some embodiments, the linear reporter polynucleotide further comprises a second promoter operably linked to a second nucleotide sequence encoding a second reporter polypeptide, wherein the second promoter and the second nucleotide sequence are positioned in the linear polynucleotide between the first nucleotide sequence and the first promoter. In some embodiments, the first reporter polypeptide and the second reporter polypeptide are different proteins. In some embodiments, the first promoter and the second promoter each comprise a eukary otic promoter nucleotide sequence. In some embodiments, the first promoter element and the second promoter element are different.
[0019] In some embodiments, the linear reporter polynucleotide further comprises a third nucleotide sequence encoding a selection marker. In some embodiments, the third nucleotide sequence encoding the selection marker lies between the second nucleotide sequence encoding the second reporter polypeptide and the first promoter element. In some embodiments, the third nucleotide sequence encoding the selection marker is operably linked to the second promoter element.
[0020] In some embodiments, the linear reporter polynucleotide comprises a nucleotide sequence comprising at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 2, wherein the nucleotide sequence, from its 5'-end to its 3’-end, comprises (i) the first nucleotide sequence encoding the first reporter polypeptide, (ii) the second promoter element, (iii) the second nucleotide sequence encoding the second reporter polypeptide, (iv) the third nucleotide sequence encoding the selection marker, and, (v) the first promoter element. In some embodiments, the second nucleotide sequence encoding the second reporter polypeptide and the third nucleotide sequence encoding the selection marker are operably linked to the second promoter element.
[0021] In some embodiments, the linear reporter polynucleotide further comprises a fourth nucleotide sequence encoding a cleavage site. In some embodiments, the fourth nucleotide sequence encoding the cleavage site lies between the second nucleotide sequence encoding the second reporter polypeptide and the third nucleotide sequence encoding the selection marker.
[0022] In another aspect, provided is a method of determining whether a compound impairs circular DNA formation in a cell, the method comprising: (i) contacting a plurality of cells with any reporter polynucleotide described above, (ii) culturing the plurality of cells comprising the polynucleotide, (iii) detecting a first signal from a first reporter polypeptide expressed from the polynucleotide, (iv) contacting the plurality of cells with a compound of interest, (v) detecting a second signal from the first reporter polypeptide, and (vi) comparing the second signal from the first reporter polypeptide to the first signal from the first reporter polypeptide.
[0023] In some embodiments, the compound impairs circular DNA formation if the second signal is less than the first signal.
[0024] In some embodiments, the polynucleotide comprises the sequence as set forth in SEQ ID NO: 1 or SEQ ID NO: 2.
[0025] In some embodiments, the method further comprises culturing the plurality of cells with a compound that corresponds to a selectable marker.
[0026] some embodiments, the method further comprises detecting a third signal from a second reporter polypeptide expressed from the polynucleotide, wherein detection of the second signal from the second reporter polypeptide indicates polynucleotide presence in the plurality of cells.
[0027] In some embodiments, the plurality of cells is a plurality of cancer cells. In some embodiments, the plurality of cells is a plurality of plant cells.
[0028] In another aspect, provided is a kit for detecting circular DNA presence in a cell. In some embodiments, the kit comprises any linear reporter polynucleotide described above and instructions for use.
[0029] In another aspect, provided is a method of inhibiting ecDNA biogenesis, comprising contacting a plurality’ of cells with an ecDNA inhibitor. In some embodiments, the ecDNA inhibitor suppresses the activity of Lig4, XCCR4, or a protein of the BRCA1-A complex. In some embodiments, the protein of the BRCA1-A complex is one of BABAM2, ABRAXAS 1, or Rap80.
[0030] In some embodiments, the plurality of cells is in a subject having or suspected of having a cancer or a plant.
[0031] In another aspect, provide is a method of screening for ecDNA formation inhibition in cells, comprising delivering any reporter polynucleotide described above to a plurality of cells, contacting a plurality of cells with an ecDNA inhibitor, thereby producing a population of treated cells, detecting an amount of ecDNA formation in the population of treated cells; and comparing the amount of ecDNA formation in the population of treated cells to a control population of untreated cells.
[0032] In some embodiments, the ecDNA inhibitor suppresses the activity’ of Lig4, XCCR4, or a protein of the BRCA1-A complex. In some embodiments, the protein of the BRCA1-A complex is one of BABAM2, ABRAXAS 1, or Rap80.
[0033] In some embodiments, the method further comprises measuring the level of a reporter
[0034] In another aspect, provided is a composition comprising an inhibitor of ecDNA biogenesis. In some embodiments, the inhibitor of ecDNA biogenesis suppresses the activity of Lig4, XCCR4, or a protein of the BRC Al -A complex. In some embodiments, the inhibitor of ecDNA biogenesis targets Lig4, XCCR4, or a protein of the BRCA1-A complex. In some embodiments, the inhibitor of ecDNA is present in an effective amount to reduce the formation of ecDNA in a plurality of cells or a subject compared to a baseline level of ecDNA formation in the subject without the inhibitor. In some embodiments, the composition further comprises a pharmaceutically acceptable carrier. In some embodiments, the inhibitor is present in a unit dose formulation.BRIEF DESCRIPTION OF THE DRAWINGS
[0035] FIGS. 1A-1F show exemplary' biosensors (a version 1, FIG. 1A and a version 2, FIG. ID) for monitoring ecDNA formation. FIG. 1A shows a schematic of the Version 1 ecDNA biosensor. Circularization brings the CAG promoter upstream of the eGFP coding sequence. FIG. IB shows testing of the Version 1 biosensor in HEK293T cells. Introduction of a linear eGFP coding sequence only or pre-ligating the biosensor into a circle did not produce eGFP expression, suggesting that eGFP expression from the biosensor relies on the circularization process, not random integration of the biosensor into the genome. FIG. 1C shows a PCR assay to measure the production of ecDNA from the biosensor. Exonuclease treatment removes linear DNA, as evidenced by the signals from the YWHAZ gene. FIG. ID shows a schematic of the Version 2 ecDNA biosensor. The EF-la promoter drives the expression of DsRed and the puromycin resistance gene, e.g. the pac gene encoding a puromycin N-acetyl-transferase, enabling selection of cells harboring the biosensors. FIG. IEshows testing of the Version 2 biosensor in HEK293T cells. FIG. IF shows testing the Version2 biosensor in three cancer cell lines - HeLa, PC9, and HCT116.
[0036] FIGS. 2A-2D relates to an exemplars’ CRISPR screen for identifying factors that regulate ecDNA biogenesis utilizing an eGFP biosensor (not VI or V2 above) for monitoring ecDNA production. FIG. 2A shows an illustration of applying an exemplary' genome-wide CRISPR screen with an exemplary ecDNA biosensor in cells, which was performed in 3 biological replicates. FIG. 2B shows a volcano plot to show the regulators of ecDNA biogenesis. The factors discussed in the Examples are labelled. FIG. 2C shows an interactome analysis of the factors that drive ecDNA biogenesis. Lig4, PRKDC, and XRCC4 were identified as drivers of ecDNA biogenesis and are components of the Lig4 complex. ABRAXAS 1, BABAM2, and UIMC1 were also identified as drivers of ecDNA biogenesis and are components of the BRCA1-A complex. Factors identified as drivers of ecDNA biogenesis in FIG. 2B are WDR92, SYMPK, C3orfl8, DPPA4, DCXR, ATP5MC1, HSPE1, PRIME H3- 3A, MAGIX, KPTN, TMEM241, GLMP, CD48, and CLRN3. Factors that show coessentiality are connected with dashed red lines. FIG. 2D shows a snake plot to show the potential function of factors from distinct DNA break repair pathways in ecDNA biogenesis.
[0037] FIGS. 3A-3B. Lig4 catalyzes ecDNA biogenesis from genomic fragments. Individually mutating the I.IG4. XRCC4, PRKDC, UIMC1 genes abolishes ecDNA production, as indicated by the loss of eGFP expression from the pre-integrated reporter (data not shown). FIG. 3A is a schematic of the reporter construct that was utilized, which is based on a prior CRISPR-C reporter and not the VI or V2 biosensors described herein. FIG. 3B is a plot of droplet digital PCR (ddPCR) results, quantifing ecDNA production from the pre-integrated reporter. The bars report mean ± standard deviation from three biological replicates (n=3). P- values were calculated with a two-tailed, two-sample unequal variance t test. For each gene, the mutant cell line generated by sgRNA-1 was picked for ddPCR. Re-introducing wild-ty pe LIG4 rescues ecDNA production from the pre-integrated reporter (data not shown). Mutant versions of LIG4 — either catalytically dead (E331A or K273A) or with the XRCC4-binding BRCT domains deleted — fail to rescue. Each rescue construct achieves a similar level of protein expression (data not shown). Lig4 drives ecDNA production in cancer cells (HCT116, HeLa, and PC9) regardless of the nucleotide composition at the two ends (data not shown). A version 2 reporter (FIG. ID) flanked by 6 random nucleotides was used.
[0038] FIGS. 4A-4D. Lig4 drives natural ecDNA biogenesis in vivo. FIG. 4A is a cartoon depicting the "Onion skin" model of chorion gene amplification occurring in Drosophila ovarian follicle cells. Dashed circles stand for DNA breaks. FIG. 4B shows ecDNA-Seq and Genome-Seq results that measure the amplification of the chorion locus on the 3rdchromosome and ecDNA production from this region. FIG. 4C are results of a PCR-based assay to measure the production of chorion ecDNA from Drosophila ovaries. Fly bodies without ovaries were used as the carcass for DNA extraction, serving as a negative control. FIG. 4D illustrates immunostaining results, showing EdU incorporating into the replicated DNA in Drosophila ovarian follicle cells. LIG4 mutation has no impact on DNA amplification, as indicated by EdU immunostaining.
[0039] FIGS. 5A-5F. Lig4 drives ecDNA-mediated cancer cell evolution. A CRISPR-C approach related to previous technology not described herein was utilized to generate megabases DHFR or EGFR ecDNA. This approach was used to generate the data presented in panels FIGS. 5A-5C. FIG. 5A ddPCR to quantify ecDNA production from the DHFR or EGFR region in HeLa or PC9 cells, respectively. The bars report mean ± standard deviation from three biological replicates (n=3). -values were calculated with a two-tailed, two-sample unequal variance t test. FIG. 5B is a graph showing bell growth curve upon methotrexate (MTX) treatment. CRISPR-C approach w as applied to generate DHFR ecDNA in HeLa cells before MTX exposure. FIG. 5C are micrographs showing DNA-FISH staining to measure DHFR ecDNA in parental and MTX-resistant cells. FIG. 5D is a schematic design of using cancer treatment drugs to induce natural ecDNA production in cell culture. This approach was used to generate data presented in panel FIG. 5E and FIG. 5G. FIG. 5E shows DNA-FISH staining to measure DHFR ecDNA in HeLa cells and EGFR ecDNA in PC9 cells. Upon drug treatment, wild-type cells formed and accumulated ecDNA to acquire drug resistance. FIG. 5F is representative of two plots showing cell growth curves upon methotrexate (MTX) (top panel) or Osimertinib treatment (bottom panel). Mutating LIG4 abolishes cancer cell evolution to adapt to drug treatment.
[0040] FIGS. 6A-6B show bioinformatic analysis of hits from the CRISPR screen. FIG. 6A shows gene ontology (GO) analysis for characterizing the pathways that potentially regulate ecDNA biogenesis. FIG. 6B shows an interactome analysis of the factors that suppress ecDNA biogenesis. MRN complex components are NBN, ATM, RAD50, and MRE1 1. Factors that show co-essentiality are connected with dashed lines.
[0041] FIG. 7 shows validation of Lig4 as ecDNA biogenesis regulator as identified in the CRISPR screen. Using eGFP expression from the reporter as the proxy for ecDNA production, immunofluorescence analysis of LIG4- / - HEK293T cell line generated in the screen showed no eGFP expression.
[0042] FIGS. 8A-8B show that suppressing the function of MRN complex enhances ecDNA production. FIGS. 8A-8B are plots showing that blocking the MRN complex function by using Mirin enhances ecDNA production. Both plots show the quantification of three biological replicates. The bars report mean ± standard deviation from three biological replicates (n=3). P- values were calculated with a two-tailed, two-sample unequal variance t test.
[0043] FIGS. 9A-9B. Lig4 drives natural ecDNA production from the Drosophila ovarian follicle cells. FIG. 9A demonstrates plot illustrating data mining from the published ecDNA- Seq data from the Drosophila ovary and shows ecDNA production from the two chorion loci. FIG. 9B shows plots utilizing ecDNA-Seq and Genome-Seq to measure the amplification of the chorion locus on the X chromosome and ecDNA production from this region.
[0044] FIGS. 10A-10B. Suppressing the function of the MRN complex accelerates ecDNA- mediated cancer cell adaptation. FIG. 10A is a plot showing quantification of DHFR ecDNA number in HeLa cells when the cells were treated with either methotrexate (MTX) alone or in combination with Mirin. FIG. 10B is a plot showing cell grow th curves when HeLa cells were treated with either methotrexate (MTX) alone or in combination with Mirin.
[0045] FIG. 11 is a plot illustrating the mutation of the BRCA1-A complex core component UIMC1 also leads to failure of cancer cells to evolve resistance to methotrexate.DETAILED DESCRIPTION
[0046] For the purposes of promoting an understanding of the principles of the present disclosure, reference will now be made to preferred embodiments and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of the disclosure is thereby intended, such alteration and further modifications of the disclosure as illustrated herein, being contemplated as would normally occur to one skilled in the art to which the disclosure relates.I. Introduction
[0047] Disclosed herein are constructs, methods, and kits relating to biosensors that are useful for studying the formation of circular DNA, e.g., extrachromosomal circular DNA(ecDNA). Biosensors as described herein are linear DNA reporter constructs (also referred to herein as “reporter construct polynucleotides7') that can produce a detectable signal, e.g., a fluorescence signal, when the linear construct is ligated to become a circular DNA molecule, thereby indicating the presence of circular DNA formation (e.g., ecDNA) in a cell. Further described herein are screening tools that allow for the study of therapeutics that can affect intracellular ecDNA formation, as well as therapeutic agents (also referred to herein as “ecDNA inhibitors”), pharmaceutical compositions compnsing ecDNA inhibitors, and methods of treatment.
[0048] The formation of ecDNA by retrotransposons in a genome has been studied (Y ang, Fu et al. Nature (2023)). Retrotransposon-derived ecDNA is not frequently produced in nature because retrotransposons are often silenced in the host genome (Wells, et al. Annual Review of Genetics (2020); Kazazian, et al. The New England Journal of Medicine (2017); Bums. KH. Nature Reviews. Cancer (2017)). A more common process of ecDNA biogenesis is the direct circularization of one or more DNA fragments generated from chromosomal breaks (Wu, et al. Annual Review of Pathology (2022); Noer, et al. Trends in Genetics Trends in Genetics (2022); Ilic. et al. Chromosoma (2022). This process appears to be more prevalent in cancer cells (in particular those relating to aggressive cancers), especially when those cells undergo chromosomal rearrangement (Shoshani, Ofer et al. Nature (2021); Rosswog, et al. Nature Genetics (2021)).
[0049] Current tools for identifying and observing ecDNA are limited. Previously, a CRISPR-C approach was developed that employs CRISPR-Cas9 cleavage to generate linear DNA fragments containing a reporter cassette (i.e., a fluorescent reporter, for example, eGFP). See Moller HD, et al. (2018). However, this CRISPR-C approach requires pre-integration of the reporter cassette into the genome, in addition to the presence of a CRISPR-Cas9 complex and sgRNAs in the cells. Moreover, the circularization efficiency of this eGFP reporter cassette is poor.
[0050] As described herein, the Inventors have developed reporter construct polynucleotides that function as biosensors for studying ecDNA biogenesis. The reporter construct systems, methods, and kits relating to these systems are superior to existing approaches, such as the CRISPR-C system discussed above. As disclosed herein, the provided reporter construct polynucleotides are linear DNA molecules that do not require pre-integration into a host cell genome or the presence of CRISPR-Cas9 or sgRNAs in the same cell. The reporter constructpolynucleotides provided herein have high transfection efficiency, thereby requiring lower transfection doses, and can provide more reliable transfection success. The reporter construct polynucleotides provided herein can also be easily modified to comprise any of a variety of promoters and coding sequences for reporter polypeptides. Reporter construct polynucleotides and systems as provided herein are useful in identifying ecDNA biogenesis factors from different DNA repair mechanisms. Use of the reporter constructs and systems as described herein are advantageous in that they can identify ecDNA factors that would otherwise go undetected with previous methods (i.e., CRISPR-C methods) because such methods are tied specifically to the formation and repair of double-stranded DNA breaks.
[0051] The reporter construct polynucleotides disclosed herein are linear DNA molecules that comprise at least a reporter polypeptide coding sequence followed by a promoter and require circularization of the linear DNA construct for the promoter to drive reporter expression. While the reporter construct polynucleotide is in its linear form, the reporter polypeptide is not expressed. When the reporter construct polynucleotide is circularized, the reporter polypeptide coding sequence and the promoter to become operably linked, allowing the promoter to drive reporter expression within a cell. Thus, the reporter polypeptide is expressed when the reporter construct polynucleotide is circularized.
[0052] In some embodiments, the reporter construct polynucleotides are useful for detecting the formation of circular DNA, e.g., ecDNA, in a cell. In some embodiments, the reporter constructs are useful for identifying regulators, e.g.. promoters or suppressors, of circular DNA, e.g.. ecDNA, formation. In some embodiments, the reporter constructs are useful for identifying compounds (e.g., chemicals or drugs) that can block circular DNA, e.g., ecDNA, formation, for use in the treatment of diseases, such as cancer.
[0053] Oncogene amplification of ecDNA appears to be a common event in cancer cells. Being able to express oncogenes at a massive level, ecDNA formation appears to drive tumorigenesis, cancer cell-acquired drug resistance, and tumor recurrence. As such, targeting ecDNA biogenesis can be a novel therapeutic approach for cancer treatment. As described in the Examples of this disclosure, the inventors performed genome-wide CRIPSR screening to identify factors that mediate ecDNA formation and found that the process does not appear to be controlled solely by a single DNA repair pathway. Rather, selective factors from different DNA repair steps appear to orchestrate ecDNA generation. The inventors’ findings suggest that upon DNA fragmentation, the end-processing complexes with opposite end-resectionfunctions antagonize each other to funnel the un-resected ends for a LIG4-catalyzed ecDNA production event. The inventors’ findings not only delineate a mechanism that likely frequently drives ecDNA production from genomic fragments, but also serve as a solid anchor point to compare the similarities or differences of distinct ecDNA formation processes.
[0054] Given that ecDNA is frequently generated in cancer cells to drive tumor evolution and adaptation, targeting ecDNA biogenesis represents a novel strategy for developing new cancer therapies. The inventors have identified that LIG4 serves as an important factor for catalyzing ecDNA formation. As described in the Examples, the inventors found that cancer cells without LIG4 lost their capability of initiating ecDNA-mediated adaptation. These findings highlight LIG4 as a potential target for cancer therapy. Humans appear to be able to tolerate the loss of LIG4 at a high level. Mutations in LIG4 leads to LIG4 syndrome, a disease from which patients show developmental microcephaly and growth retardation but only have manageable immune deficiency in adult life. Additionally, the inventors found that LIG4 normally is not required for cell viability, but only appears to be critical when cancer cells need to adapt to drug treatment. These findings suggest that the potential toxicity from targeting LIG4 should be minimal and controllable, suggesting a that LIG4 is a druggable target. Thus, provided in this disclosure are methods of determining whether a compound impairs circular DNA formation in a cell, methods of screening for ecDNA formation inhibition, and methods of inhibiting ecDNA biogenesis. In some instances, the methods are performed using the reporter construct polynucleotides as described in this disclosure.II. Terminology
[0055] A number of terms and concepts are discussed below. They are intended to facilitate the understanding of various embodiments of the invention in conjunction with the rest of the present document and the accompanying figures. These terms and concepts may be further clarified and understood based on the accepted conventions in the fields of the present invention, as well as the descnption provided throughout the present document and / or the accompanying figures. Some other terms can be explicitly or implicitly defined in other sections of this document and in the accompanying figures and may be used and understood based on the accepted conventions in the fields of the present invention, the description provided throughout the present document and / or the accompanying figures. The terms not explicitly defined can also be defined and understood based on the accepted conventions in the fields of the present invention and interpreted in the context of the present document and / or the accompanying figures.
[0056] Unless otherwise defined, all terms of art, notations, and other scientific or medical terms or terminology used herein are intended to have the meanings commonly understood by those of ordinary skill in the art. In some cases, terms with commonly understood meanings are defined herein for clarity and / or for ready reference, and the inclusion of such definitions herein should not be construed as representing a substantial difference over the definition of the term as generally understood in the art.
[0057] As used herein, the term “extrachromosomal DNA,” "ecDNA.'’ “extrachromoromal circular DNA,” or “eccDN A" refers to a double-stranded, closed-circle DNA molecule. Natural ecDNA molecules are typically extrachromosomal in structure and of endogenous chromosomal origin. EcDNA is present across different species and may comprise small polydispersed circular DNA (spcDNA) (100-10,000 bp), episomes (submicroscopic size range), microDNA (200-3000 bp), telomeric circles (t-circles) (100-30,000 bp), double minutes (DMs) (100 kb3 Mb), and cancer-specific circular extrachromosomal DNA (ecDNA) (mega-base-pair amplified). EcDNAs are discussed in detail in, for example, Zhao, Yiheng et al. eLife vol. 11 e81412. 18 Oct. 2022, doi: 10.7554 / eLife.81412; Noer, Julie B et al. Trends in genetics : TIG vol. 38,7 (2022): 766-781. doi: 10.1016 / j.tig.2022.02.007; Zuo, Shanru et al. Frontiers in cell and developmental biology vol. 9 792555. 6 Jan. 2022, doi:10.3389 / fcell.2021.792555.
[0058] Articles “a” and “an” are used herein to refer to one or to more than one (i.e. at least one) of the grammatical obj ect of the article. By way of example, “an element” means at least one element and can include more than one element.
[0059] “About” is used to provide flexibility to a numerical range endpoint by providing that a given value may be “slightly above” or “slightly below” the endpoint without affecting the desired result.
[0060] The use herein of the terms “including,” “comprising,” or “having,” and variations thereof, is meant to encompass the elements listed thereafter and equivalents thereof as well as additional elements. As used herein, “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations where interpreted in the alternative (“or”).
[0061] As used herein, the transitional phrase “consisting essentially of' (and grammatical variants) is to be interpreted as encompassing the recited materials or steps "and those that do not materially affect the basic and novel characteristic(s)" of the claimed invention. Thus, theterm “consisting essentially of as used herein should not be interpreted as equivalent to comprising / ’
[0062] Recitation of ranges of values herein are merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, unless otherwise indicated herein, and each separate value is incorporated into the specification as if it were individually recited herein. For example, if a concentration range is stated as 1% to 50%, it is intended that values such as 2% to 40%, 10% to 30%, or 1% to 3%, etc., are expressly enumerated in this specification. These are only examples of what is specifically intended, and all possible combinations of numerical values between and including the low est value and the highest value enumerated are to be considered to be expressly stated in this disclosure.
[0063] The terms “sample” and “biological sample” are used interchangeably herein and include, but are not limited to, a sample containing tissues, cells, and / or biological fluids. Such samples may be isolated from a subject, produced and / or maintained in culture (e.g., cell line), or obtained from a subject and then maintained in culture. Examples of biological samples include, but are not limited to, tissues, cells, biopsies, blood, lymph, serum, plasma, urine, saliva, mucus and tears. A biological sample may be obtained directly from a subject (e.g., by blood or tissue sampling) or from a third party (e.g.. received from an intermediary, such as a healthcare provider or lab technician).
[0064] “Contacting” as used herein, e.g., as in “contacting a sample” refers to contacting a sample directly or indirectly in vitro, ex vivo, or in vivo (i.e. within a subject as defined herein). Contacting a sample may include addition of genetic material (e.g., a reporting construct as provided herein) to a sample (e.g.. cell culture, biological sample, etc.), or administration to a subject. Contacting encompasses administration to a solution, cell, tissue, mammal, subject, patient, or human. Further, contacting a cell includes (i) adding genetic material (e.g. a reporting construct as provided herein) to a cell culture as well the acts of transfection and / or transformation.
[0065] As used herein, the terms “transfect” or “transfection” refer to the intracellular introduction of one or more encapsulated materials (e.g., nucleic acids and / or polynucleotides encapsulated by a virus) into a cell, or preferably into a target cell. The term “transfection efficiency” refers to the relative amount of such encapsulated material (e.g., polynucleotides encapsulated by a virus) up-taken by, introduced into and / or expressed by the target cell which is subject to transfection. In some embodiments, transfection efficiency may be estimated bythe amount of a reporter polynucleotide product produced by the target cells following transfection. In some embodiments, a transfer vehicle has high transfection efficiency. In some embodiments, a transfer vehicle has at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% transfection efficiency.
[0066] As used herein, the term “transformation” refers to the specific process where exogenous genetic material is directly taken up and incorporated by a cell through its cell membrane.
[0067] Linear nucleic acid molecules are said to have a “5’-terminus” (or “5’ end”) and a “3 ’-terminus” (or “3’ end”) because nucleic acid phosphodiester linkages occur at the 5’ carbon and 3’ carbon of the sugar moi eties of the substituent mononucleotides. The end nucleotide of a polynucleotide at which a new linkage would be to a 5’ carbon is its 5’ terminal nucleotide. The end nucleotide of a polynucleotide at which a new linkage would be to a 3’ carbon is its 3’ terminal nucleotide. A “terminal nucleotide,” as used herein, is the nucleotide at the end position of the 3’- or 5’-terminus.
[0068] As used herein, the term “circularization efficiency” refers to a measurement of the rate of formation of amount of resultant circular polyribonucleotide as compared to its linear starting material.
[0069] The expression sequences in the polynucleotide construct may be separated by a “cleavage site” sequence which enables polypeptides encoded by the expression sequences, once translated, to be expressed separately by the cell. A “self-cleaving peptide” refers to a peptide which is translated without a peptide bond between two adjacent amino acids, or functions such that when the polypeptide comprising the proteins and the self-cleaving peptide is produced, it is immediately cleaved or separated into distinct and discrete first and second polypeptides without the need for any external cleavage activity.
[0070] As used herein, “coding element,” “coding sequence,” “coding nucleic acid,” or “coding region” is region located within the expression sequence and encodings for one or more proteins or polypeptides (e.g., a reporter protein, a therapeutic protein, etc.). As used herein, a “noncoding element,” “noncoding sequence,” “non-coding nucleic acid,” or “noncoding nucleic acid” is a region located within the expression sequence. This sequence, but itself does not encode for a protein or polypeptide, but may have other regulator ' functions, including but not limited, allow the overall polynucleotide to act as a biomarker or adjuvant to a specific cell.
[0071] The terms “nucleic acid'’ and “polynucleotide” are used interchangeably herein to describe a polymer of any length, e.g, greater than about 2 bases, greater than about 10 bases, greater than about 100 bases, greater than about 500 bases, greater than 1000 bases, or up to about 10,000 or more bases, composed of nucleotides, e.g., deoxy ribonucleotides or ribonucleotides, and may be produced enzy matically or synthetically (e.g., as described in U.S. Pat. No. 5.948,902 and the references cited therein), which can hybridize with naturally occurring nucleic acids in a sequence specific manner analogous to that of two naturally occurring nucleic acids, e.g., can participate in Watson-Crick base pairing interactions. An “oligonucleotide” is a polynucleotide comprising fewer than 1000 nucleotides, such as a polynucleotide comprising fewer than 500 nucleotides or fewer than 100 nucleotides. Naturally- occurring nucleic acids are comprised of nucleotides, including guanine, cytosine, adenine, thymine, and uracil containing nucleotides (G, C, A, T, and U respectively). As used herein, “poly A” means a polynucleotide or a portion of a polynucleotide consisting of nucleotides comprising adenine. As used herein, “polyT” means a polynucleotide or a portion of a polynucleotide consisting of nucleotides comprising thymine. As used herein, “poly AC” means a polynucleotide or a portion of a polynucleotide consisting of nucleotides comprising adenine or cytosine. Unless otherwise indicated, a particular polynucleotide sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions) and complementary sequences as well as the sequence explicitly indicated. Specifically, degenerate codon substitutions may be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues.
[0072] As used herein, the term “expression sequence” refers to a nucleic acid sequence that encodes a product, e.g, a peptide or polypeptide, regulatory nucleic acid, or non-coding nucleic acid. An exemplary expression sequence that codes for a peptide or polypeptide can comprise a plurality of nucleotide triads, each of which can code for an amino acid and is termed as a “codon.”
[0073] Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof, alleles, orthologs, SNPs, and complementary sequences as well as the sequence explicitly indicated.
[0074] The terms “polypeptide,” “protein.” and “peptide” are used interchangeably herein to refer to a polymer of amino acid residues in a single chain. The terms apply to amino acidpolymers in which one or more amino acid residue is an artificial chemical mimetic of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers and non-naturally occurring amino acid polymers. Amino acid polymers may comprise entirely L-amino acids, entirely D-amino acids, or a mixture of L and D amino acids. The term “protein"’ as used herein refers to either a polypeptide or a dimer (i.e., two) or multimer (i.e., three or more) of single chain polypeptides. The single chain polypeptides of a protein may be joined by a covalent bond, e.g., a disulfide bond, or non-covalent interactions. The terms “portion” and “fragment” are used interchangeably herein to refer to parts of a polypeptide, nucleic acid, or other molecular construct.
[0075] The amino acids in the polypeptides described herein can be any of the 20 naturally occurring amino acids, D-stereoisomers of the naturally occurring amino acids, unnatural amino acids and chemically modified amino acids. Unnatural amino acids (that is. those that are not naturally found in proteins) are also known in the art, as set forth in, for example, Zhang et al., 2013, Curr. Opin. Struct. Biol. 23(4): 581-87; Xie et al., 2005, Curr. Opin. Chem. Biol. 9(6): 548-54; and all references cited therein. Beta and gamma amino acids are known in the art and are also contemplated herein as unnatural amino acids.
[0076] As used herein, a chemically modified amino acid refers to an amino acid whose side chain has been chemically modified. For example, a side chain can be modified to comprise a signaling moiety, such as a fluorophore or a radiolabel. A side chain can also be modified to comprise a new functional group, such as a thiol, carboxylic acid, or amino group. Post- translationally modified amino acids are also included in the definition of chemically modified amino acids.
[0077] The term “sequence identity,” as used herein, refer to the extent that sequences are identical on a nucleotide-by -nucleotide basis or an amino acid-by-amino acid basis over a window' of comparison. Thus, a “percentage of sequence identity"’ may be calculated by comparing two optimally aligned sequences over the window of comparison, determining the number of positions at which the identical nucleic acid base (e.g.. A, T, C, G, I) or the identical amino acid residue (e.g., Ala, Pro, Ser, Thr, Gly, Vai, Leu, He, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gin, Cys and Met) occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison (i.e.. the window' size), and multiplying the result by 100 to yield the percentage of sequence identity. Included are nucleotides and polypeptides having at leastabout 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity’ to any of the reference sequences described herein, typically where the polypeptide variant maintains at least one biological activity7of the reference polypeptide.
[0078] For sequence comparison, typically one sequence acts as a reference sequence to which test sequences are compared. When using a sequence comparison algorithm, test and reference sequences are entered into a computer, subsequence coordinates are designated, if necessary, and sequence algorithm program parameters are designated. Default program parameters can be used, or alternative parameters can be designated. The sequence comparison algorithm then calculates the percent sequence identities for the test sequences relative to the reference sequence, based on the program parameters.
[0079] A “comparison window,” as used herein, includes reference to a segment of any one of the number of contiguous positions selected from the group consisting of from 20 to 600, usually about 50 to about 200, more usually about 100 to about 150 in which a sequence may be compared to a reference sequence of the same number of contiguous positions after the two sequences are optimally aligned. Methods of alignment of sequences for comparison are well- known in the art. Optimal alignment of sequences for comparison may be conducted by the local homology algorithm of Smith & Waterman. 1981, Add. APL. Math. 2:482, by the homology7alignment algorithm of Needleman & Wunsch, 1970, J. Mol. Biol. 48:443, by the search for similarity method of Pearson & Lipman, 1988, Proc. Natl. Acad. Sci. (U.S.A.) 85:2444, by computerized implementations of these algorithms (e.g., BLAST), or by manual alignment and visual inspection.
[0080] Algorithms that are suitable for determining percent sequence identity7and sequence similarity are the BLAST and BLAST 2.0 algorithms, which are described in Altschul et al., 1990, J. Mol. Biol. 215: 403-10 and Altschul et al., 1977, Nucleic Acids Res. 25: 3389-402, respectively. Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information (NCBI) web site. The algorithm involves first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query7sequence, which either match or satisfy7some positive-valued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as the neighborhood word score threshold (Altschul et al. (1977)). These initial neighborhood word hits acts as seeds for initiating searches to find longer HSPs containing them. The word hits are then extended in both directions along each sequence for as far as the cumulative alignmentscore can be increased. Cumulative scores are calculated using, for nucleotide sequences, the parameters M (reward score for a pair of matching residues; always >0) and N (penalty score for mismatching residues: always <0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. Extension of the word hits in each direction are halted when: the cumulative alignment score falls off by the quantity X from its maximum achieved value; the cumulative score goes to zero or below, due to the accumulation of one or more negativescoring residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a word size (W) of 28, an expectation (E) of 10, M=l, N=-2, and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a word size (W) of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff & Henikoff, 1989, Proc. Natl. Acad. Sci. USA 89:10915).
[0081] The BLAST algorithm also performs a statistical analysis of the similarity between two sequences (see, e.g., Karlin & Altschul, 1993, Proc. Nat'l. Acad. Sci. USA 90:5873-87). One measure of similarity provided by the BLAST algorithm is the smallest sum probability (P(N)), which provides an indication of the probability by which a match between two nucleotide or amino acid sequences would occur by chance. For example, a nucleic acid is considered similar to a reference sequence if the smallest sum probability in a comparison of the test nucleic acid to the reference nucleic acid is less than about 0.01, more preferably less than about 10-5, and most preferably less than about 10-20.
[0082] "‘Transcription” means the formation or synthesis of an RNA molecule by an RNA polymerase using a DNA molecule as a template. The invention is not limited with respect to the RNA polymerase that is used for transcription. For example, in some embodiments, a T7- type RNA polymerase can be used.
[0083] “Translation” means the formation of a polypeptide molecule by a ribosome based upon an RNA template. As used herein, the term “translation efficiency” refers to a rate or amount of protein or peptide production from a ribonucleotide transcript. In some embodiments, translation efficiency can be expressed as amount of protein or peptide produced per given amount of transcript that codes for the protein or peptide.
[0084] As used herein, the terms “upstream” and “downstream” refer to relative positions of genetic code, e.g., nucleotides, sequence elements, in polynucleotide sequences. In some embodiments, in an RNA polynucleotide, upstream is toward the 5’ end of the polynucleotideand downstream is toward the 3’ end. In some embodiments, in a DNA polynucleotide, upstream is toward the 5’ end of the coding strand for the gene in question and downstream is toward the 3’ end.
[0085] Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.III. Reporter Construct Polynucleotides
[0086] Disclosed herein are reporter construct polynucleotides which can produce a detectable signal when extrachromosomal (i.e., non-genomic) circular DNA molecules, e.g, ecDNA, are formed in a cell. Thus, the reporter construct polynucleotides can be used to detect the formation of circular DNA, and are, therefore, biosensors for circular DNA. As used herein, the term ‘‘reporter construct polynucleotide,” “reporter construct,” or “reporter element” are synonymous with the term “biosensor” and these terms are used interchangeably. In some embodiments, the reporter polypeptide is useful for detecting circularization or circularization efficiency of a reporter construct polynucleotide in a cell, in a cell line or in a population of cells.
[0087] One aspect of the present disclosure provides ecDNA reporter constructs that comprise, consist of, or consist essentially of a linear DNA having a reporter element followed by a promoter, wherein the reporter element and promoter become operably linked upon circularization of the liner reporter construct. Other aspects of the present disclosure provide for ecDNA reporter constructs that comprise, consist of, or consist essentially of linear DNA polynucleotides having (from 5’ to 3’): a first reporter element; a second promoter; a second reporter element; and a first promotor. The first reporter element and first promoter become operably linked upon circularization of the linear reporter construct, driving expression of the first reporter element in the cell. In contrast, expression of the second reporter element is driven by the second promoter regardless of whether the reporter construct is in linear or circular form.
[0088] In some embodiments, a reporter construct polynucleotide of the present disclosure comprises a nucleotide sequence encoding a reporter polypeptide and a promoter. The reporter construct polynucleotide is a linear DNA molecule with the reporter polypeptide coding sequence on the DNA’s 5 ’-end and the promoter on the DNA’s 3 ’-end. See, e.g., “Version 1 biosensor” in FIG. 1A. While the reporter construct polynucleotide is in its linear form, the promoter remains downstream of the reporter polypeptide coding sequence and cannot driver reporter expression. Thus, the reporter polypeptide coding sequence and the promoter are notoperably linked and the reporter polypeptide is not expressed. When the free ends of a reporter construct polynucleotide are connected (i.e. ligated; circularized), the reporter construct becomes a circular DNA molecule (like ecDNA), thereby bringing the 3 ’-end of the promoter upstream of the reporter polypeptide coding sequence. This circularization operably links the promotor with the reporter element, allowing the reporter polypeptide to be expressed.
[0089] In some embodiments, the reporter construct polynucleotide comprises, from its 5‘- end to its 3 ’ -end, ( 1 ) a coding sequence for a first reporter polypeptide, (2) an internal promoter, (3) a coding sequence for a second reporter polypeptide operably linked to the internal promoter (thereby expressing the second reporter polypeptide when the reporter construct polynucleotide is in a linear form), and (4) a 3’-end promoter. See, e.g, “Version 2 biosensor” in FIG. ID. In some embodiments, the first and the second reporter polypeptides are different (e.g., the first reporter polypeptide comprises the eGFP amino acid sequence and the second reporter polypeptide comprises a mCherry or dsRed amino acid sequence). The coding sequence for the second reporter polypeptide is operably linked to the internal promoter. Thus, when the reporter construct polynucleotide is in its linear form, the second reporter polypeptide can be expressed. In contrast, when the reporter construct polynucleotide is in its linear form, the first reporter polypeptide is not expressed. However, when the reporter construct polynucleotide is circularized into a circular DNA molecule (like ecDNA), the first reporter polypeptide can be expressed the 3 ’-end promoter is positioned upstream of the first reporter polypeptide coding sequence, and the 3’-end promoter and the first reporter polypeptide coding sequence become operably linked. Thus, the first reporter polypeptide is expressed only when the reporter construct polynucleotide is in a circular DNA molecule, e.g., ecDNA.
[0090] In some embodiments where the reporter construct polynucleotide comprises coding sequences for two reporter polypeptides, the first reporter polypeptide is useful for detecting the presence ( / . e. formation) of the circular form of the reporter construct, which reflects a level of ecDNA biogenesis activity. In some embodiments, the first reporter polypeptide is useful for detecting circularization and / or circularization efficiency of the reporter construct polynucleotide in a cell, a cell line, or in a population (i.e., a plurality ) of cells. In some embodiments, the second reporter polypeptide is useful for detecting uptake efficiency of the reporter construct polynucleotide into a cell (i.e.. transfection efficiency, electroporation efficiency, etc ), a cell line, or population of cells. In some embodiments, the second reporter polypeptide is useful for detecting (e.g., visualizing) cells that comprise the reporter construct polynucleotide. In some embodiments, the second reporter polypeptide is useful fordetermining uptake efficiency of the reporter construct polynucleotide in a cell, cell line, or population of cells.
[0091] In some embodiments, the reporter construct polynucleotide further comprises a coding sequence (e.g, gene or cassette) for a selectable marker (e.g., an antibiotic-resistance gene). Suitable selectable markers are described in further detail below. In some embodiments, the selectable marker is operably linked to a promoter. Without intending to be limiting, in some embodiments, the selectable marker can be positioned in the reporter construct between the coding sequences of the second reporter polypeptide and the 3 ’-end promoter. See, e.g.. “Version 2 biosensor” in FIG. ID.
[0092] In some embodiments, the reporter construct polynucleotide can be delivered to cells as a linear DNA molecule. In some embodiments, the reporter construct polynucleotide can be cloned into an expression vector for delivery’ of the reporter construct polynucleotide to cells. Many suitable vectors and related methods may be used for cloning a reporter construct polynucleotide as known in the art, for example, without limitations, an adenoviral (AAV) vector, a lentiviral vector, a piggy7Bac® transposon vector, a non-viral plasmid, or a Sleeping Beauty transposon vector. In some embodiments, the vector can be a transposon vector, for example the piggyBac® transposon-based vector (SBI System Biosciences; Cat. # PB210PA- 1). a. Reporter Polypeptides
[0093] As used herein, the term “reporter polypeptide” refers to any reporter protein or variant thereof that produces a detectable signal that can be used to indicate the presence of a circular form of a reporter construct polynucleotide in a cell. Many genes are suitable for use in a reporter construct polynucleotide to express a reporter polypeptide, including genes that can be used to track the physical location or a segment of DNA, to monitor gene or plasmid expression, or to monitor circularization of a reporter construct polynucleotide (which is reflective of ecDNA formation) in a cell.
[0094] Under standard cell culture conditions, and when a reporter construct polynucleotide is circularized, the reporter polypeptide is expressed from its coding sequence in the reporter construct polynucleotide. The amount of signal that is detected from the reporter polypeptide can be used indicate the presence and / or the amount of circular DNA, e.g., ecDNA, in a cell. The reporter polypeptide can also be used to monitor gene or plasmid expression in a cell. In some embodiments, the reporter polypeptide is a fluorescent protein and the detectable signalis a fluorescence signal. Suitable examples of reporter polypeptides include, but are not limited to, green fluorescent proteins (e.g, GFP, enhanced GFP (eGFP), and mGreenLantem), red fluorescent proteins (e.g., RFP, mCherry, mScarlet, DsRed), yellow fluorescent proteins e.g. YFP, Citrine, Venus, and Ypet), cyan fluorescent protein (e.g., ECFP, Cerulean, CyPet, mTurquoise2), photoactivatable fluorescent proteins (e.g., PAGFP, PSCFP, PSCFP2, Dendra, Dendra2, EosFP. tdEos, mEos2, mEos3, PamCherry, PAtagRFP. mMaple, mMaple2. and mMaple3), luciferase (e.g, include firefly luciferase (FLuc), Renilla luciferase (RLuc), Gaussia luciferase (GLuc), NanoLuc® luciferase (NLuc; N-Luc; Promega)), chloramphenicol acetyltransferase (cat), and any variant thereof. A list of suitable reporters can be found at FPbase (fpbase.org / table). The use of fluorescent proteins and their applications in imaging cells are discussed, for example, in Chudakov et al. Physiological Reviews 90(3): 1103-1163 (2010); and Specht et al. Annual Review of Physiology 79: 93-117 (2017). In some embodiments, the reporter polypeptide is eGFP. In some embodiments, the reporter polypeptide is DsRed. In embodiments, the reporter construct does not comprise, consist of, or consist essentially of a pCBH promoter that becomes operably coupled to an eGFP coding sequence (thereby driving eGFP expression) when a reporter construct is circularized.
[0095] In some embodiments, a reporter construct polynucleotide comprises coding sequences for two or more reporter polypeptides, wherein some or all of the reporter polypeptides can be different or the same from each other. In some embodiments, the reporter construct polynucleotide comprises coding sequences for two reporter polypeptides that are different from each other. In some embodiments, the reporter construct polynucleotide comprises coding sequences for two reporter polypeptides that are the same reporter. In some embodiments, the two reporter polypeptides are eGFP and DsRed. b. Promoters
[0096] A variety of promoters may be used to drive the expression of one or more coding sequences in a reporter polypeptide of the present disclosure. Promoters that can be used may be any appropriate promoter sequence suitable for a host cell, which is capable of driving transcriptional activity, including mutant, truncated, and hybrid promoters.
[0097] As used herein, the term “promoter” refers to specified segments of DNA that lead to the initiation of transcription of a specific gene, and are thus capable of directing or driving expression of a coding sequence in a host cell. Promoters are involved in recognizing and binding of RNA polymerase and other proteins that initiate transcription. A promoter can belocated upstream and / or downstream from the start of transcription that is involved in recognition and binding of RNA polymerase and other proteins to initiate transcription. Promoters may be eukaryotic or prokaryotic. Further, promoters may also be constitutive, spatiotemporal, tissue-specific, and / or inducible. A promoter may be an inducible promoter or a constitutive promoter. An “inducible promoter” is a promoter that is active under environmental or developmental regulation, for example, regulated by the presence or absence of an induction signal. A “constitutive promoter” is a promoter that is active under most environmental and developmental conditions. Suitable promoters include, but are not limited to, CMV, EFla, SV40, PGK1, Ubc, human beta chain, CAG, TRE, UAS, Ac5, POlyhedrin, CaMKIIa, GALI, GAL10, TEF1, GDS, ADH1, CaMV35S, Ubi, Hl, U6, T7, T71ac, Sp6, araBAD, trp, lac. Ptac, pL, T3. and combinations thereof.
[0098] In some embodiments, the promoter is a eukaryotic promoter, for example, without limitations, a cytomegalovirus (CMV) enhancer / chicken (Lactin promoter (CAG) promoter, an EF-la promoter, a cytomegalovirus (CMV) promoter, a PGK promoter, a U6 promoter, or an UAS promoter, a TetOn / Off promoter. In some embodiments, the promoter is a CAG promoter. In some embodiments, the promoter is an EF-la promoter. In some embodiments, the promoter is a prokaryotic promoter, for example, a T7 promoter, an Sp6 promoter, a lac promoter, a pBAR promoter, a trp promoter, or a Ptac promoter.
[0099] In some embodiments, the promoter can be a constitutive promoter, for example, without limitations, a CAG promoter, an EF-la promoter, an MTL promoter, a CMV promoter, a U6 promoter, a PGK promoter, or an SV40 promoter. In some other embodiments, the promoter can be an inducible promoter, for example, without limitations, a pL promoter (induced by an increase in temperature), a BAD promoter (AraBAD promoter (pBAD); induced by the addition of arabinose to the growth medium), the tetracycline-controlled transcriptional activation system (TRE promoter; Tet-On / Tet-Off: Bujard and Gossen, PNAS, 89(12):5547-5551 (1992)). the Lac switch inducible system (Wyborski et al. Environ Mol Mutagen 28(4):447-58 (1996)), the ecdysone-inducible gene expression system (No et al. PNAS, 93(8):3346-3351 (1996)), the cumate gene-switch system (Mullick et al. BMC Biotechnology, 6:43 (2006)), the tamoxifen-inducible gene expression (Zhang et al. Nucleic Acids Research, 24:543-548 (1996)).
[0100] In some other embodiments, the promoter is a spatiotemporal or a tissue-specific promoter, for example, without limitations, a B29 promoter, a CD14 promoter, a CD43promoter, a CD45 promoter, a CD68 promoter, Desmin promoter, elastase-1 promoter, endoglin promoter, fibronectin promoter, Fit- 1 promoter. GFAP promoter, GPIIb. promoter, ICAM-2 promoter, mIFN-p promoter, Mb promoter, NphsI promoter, OG-2 promoter, SP-B promoter, Synl promoter, or WASP promoter.
[0101] In some embodiments where the reporter construct polynucleotide comprises two or more promoters (e.g., like the “Version 2 biosensor” shown in FIG. ID), the promoters may have different strengths in terms of the amount of gene expression each one can produce. Promoters can be a medium-strength promoter, a weak promoter, or a strong promoter. The strength of a promoter can be measured by comparing the level of transcription of a particular gene driven by that particular promoter or interest, relative to the level of transcription of that same gene driven by a suitable control promoter. For example, a promoter of a eukaryotic housekeeping gene could be used as a suitable control promoter.
[0102] In some embodiments, the reporter construct polynucleotide comprises one, two, or more eukaryotic promoters. In some embodiments, the reporter construct polynucleotide comprises one, two, or more constitutive promoters. In some embodiments, the reporter construct polynucleotide comprises a cytomegalovirus (CMV) enhancer / chicken [3-actin promoter (CAG) promoter. In some embodiments, the reporter construct polynucleotide comprises an EF-la promoter.
[0103] In some embodiments, the reporter construct polynucleotide comprises at least one promoter that is not operably linked to any coding sequence, and only becomes operably linked to a reporter element coding sequence upon circularization of the initial linear reporter construct. In some embodiments, the linear reporter construct polynucleotide comprises at least one promoter that is operably linked to a coding sequence, e.g., a coding sequence for a reporter polypeptide, such that the reporter polypeptide is expressed when the reporter construct polynucleotide is linear. See, for example, without limitations, the “Version 2 biosensor” shown in FIG. ID. In some embodiments, the reporter construct polynucleotide comprises one promoter that can drive expression of one reporter polypeptide. In some embodiments, the promoter is CAG. In some embodiments, the reporter element is EF-la. c. Selectable Marker[s]
[0104] The reporter construct polynucleotides disclosed herein can also comprise one or more selectable markers (i.e.. the polynucleotides can comprise one or more selection cassettes encoding selectable markers). The reporter construct polynucleotides can express a selectionmarker even when the reporter construct polynucleotides are still in linear form. As used herein, the term "‘selectable marker" refers to markers that help identify cells that have successfully transformed by, or have taken up, the reporter construct polynucleotide. In general, cells are transfected with a reporter construct polynucleotide comprising a selectable marker and are cultured in the presence of a corresponding selection molecule. Thus, only cells that have been successfully transfected with the reporter construct polynucleotide will survive selection pressure, e.g. culturing conditions that include the presence of a selection molecule such as a drug, and are thereby selected by the selection molecule. For example, in embodiments where the selectable marker is an antibiotic-resistance gene, the cells are cultured post-transfection in the presence of an antibiotic. For example, in some embodiments where the selectable marker is a puromycin-resistance gene, e.g., the pac gene encoding a puromycin N-acetyl-transferase, the cells are cultured post-transfection in the presence of puromycin. Selectable markers and their corresponding compounds are readily known to one of ordinary skill in the art.
[0105] Many suitable selectable markers can be used in a reporter construct polynucleotide. Suitable examples include, but are not limited to, metabolic selectable markers (e.g, dihydrofolate reductase (DHFR). glutamine synthase (GS), etc.), antibiotic selectable markers (e.g.. the pac gene encoding a puromycin N-acetyl-transferase. blasticidin deaminase, histidinal dehydrogenase, hygromycin phosphotransferase, zeocin resistance gene, bleomycin resistance gene, neomycin resistance gene, aminoglycoside phosphotransferase, etc.) and the like. In some embodiments, the selectable marker is puromycin N-acetyl-transferase. In some embodiments, the selectable marker is operably linked to its own promoter (e.g. a promotor as discussed above). In some embodiments, the selectable marker is operably linked to an EF-la promoter. In some embodiments, cells comprising a reporter construct polypeptide that comprise a selectable marker are cultured in the presence of a compound that corresponds to the selectable marker and puts selection pressure on a population of cells into which the reporter constructs are delivered. Selectable markers and their corresponding compounds are readily known to one of ordinary' skill in the art. For example, in some embodiments where the selectable marker is puromycin N-acetyl-transferase, the cells are cultured in the presence of puromycin.
[0106] In some embodiments, the reporter construct polynucleotide comprises a coding sequence for a selectable marker and a coding sequence for a reporter polypeptide, and the two coding sequences are adjacent on the reporter construct. In some embodiments, these two coding sequences are operably linked to the same promoter. In some embodiments, the reporterconstruct polynucleotide can also comprise a self-cleavage site sequence that separates the two coding sequences. The self-cleavage site sequence encodes for a "‘self-cleaving peptide" such that ribosomal skipping of a peptide bond between two adjacent amino acid residues in the sequence results in expression of the selectable marker and the reporter polypeptide as discrete and separate polypeptides from each other. Exemplary self-cleaving peptides include, without limitations, E2A, T2A, P2A, and F2A.
[0107] Accordingly, another aspect of the present disclosure provides for a reporter construct polynucleotide, e.g. an ecDNA reporter construct, comprising, consisting of, or consisting essentially of a linear DNA molecule having (i) a reporter element (a reporter construct polynucleotide); (ii) followed by a selection cassette; (iii) followed by a promoter element. In some embodiments, the reporter construct polynucleotide also comprises a cleavage site sequence encoding a self-cleaving peptide. d. Exemplary Reporter Construct Polynucleotides
[0108] In some embodiments, the reporter construct polynucleotide is a linear polynucleotide comprising a reporter polypeptide coding sequence and a promoter, wherein the reporter polypeptide coding sequence and the promoter are not operably linked in the linear reporter construct polynucleotide. In some embodiments, the reporter polypeptide coding sequence is on the 5 ’-end of the reporter construct polynucleotide, and the promoter is on the 3 ’-end of the reporter construct polynucleotide. In some embodiments, the reporter construct polypeptide also comprises a selectable marker. In some embodiments, the selectable marker is positioned between the reporter polypeptide coding sequence and the first promoter. In some embodiments, the selectable marker is operably linked to a second promoter, i.e., an internal promoter of the reporter construct polynucleotide that is positioned downstream of the reporter polypeptide coding sequence and upstream of the reporter polypeptide coding sequence and the promoter on the 3’-end of the reporter construct polynucleotide (3' end promoter). In some embodiments, the reporter polypeptide is eGFP and the 3’-end promoter is a CAG promoter. In some embodiments, the selectable marker is a puromycin-resistance gene, e.g., the pac gene encoding a puromycin N-acetyl-transferase. In some embodiments, the internal promoter for the selectable marker is an EF-la promoter.
[0109] In some embodiments, the reporter construct polynucleotide is a linear polynucleotide comprising, from its 5 ’-end to its 3 ’-end. a first reporter polypeptide coding sequence followed by a first promoter, wherein the linear polynucleotide comprises a nucleotide sequence havingat least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to SEQ ID NO: 1.
[0110] In some embodiments, the reporter construct polynucleotide is a linear polynucleotide comprising, from its 5’-end to its 3’-end, a first reporter polypeptide coding sequence on the 5'-end of the reporter construct and a first promoter on the 3’-end of the reporter construct (3‘- end promoter), wherein the reporter polypeptide coding sequence and the 3 '-end promoter are not operably linked in the linear reporter construct polynucleotide, and wherein the linear polynucleotide further comprises positioned in between the first reporter polypeptide coding sequence and the first promoter, a second promoter (i.e. internal promoter) operably linked to at least one of a second reporter polypeptide coding sequence or a selectable marker. In some embodiments, the linear polynucleotide further comprises a second promoter, a second reporter polypeptide coding sequence, and a selectable marker, and the second promoter drives expression of the second reporter polypeptide coding sequence and the selectable marker. In some embodiments, the second promoter is adjacent to the second reporter polypeptide coding sequence. In some embodiments, the second promoter is adjacent to the selectable marker. In some embodiments, the first reporter polypeptide and the second reporter polypeptide are different. In some embodiments, the first reporter polypeptide is eGFP. In some embodiments, the first promoter (i.e., the 3-end promoter) is a CAG promoter. In some embodiments, the second promoter is an EF-la promoter. In some embodiments, the second reporter polypeptide is DsRed. In some embodiments, the selectable marker is a puromycin-resistance gene, e.g, the pac gene encoding a puromycin N-acetyl-transferase. In some embodiments, a cleavage site sequence that encodes for a self-cleaving peptide lies in between the second reporter polypeptide coding sequence and the selectable marker. In some embodiments, the cleavage site sequence encodes for T2A.
[0111] In some embodiments, the reporter construct polynucleotide is a linear polynucleotide comprising, from its 5 ’-end to its 3 ’-end, (i) a first reporter polypeptide coding sequence, (ii) a promoter, (iii) a second reporter polypeptide coding sequence, (iv) a selection marker, and (v) a 3 '-end promoter, wherein the linear polynucleotide comprises a nucleotide sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to SEQ ID NO: 2.
[0112] Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g, degenerate codon substitutions), alleles, orthologs, SNPs, and complementary sequences as well as the sequence explicitly indicated. Specifically, degenerate codon substitutions may be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues. See Batzer et al. Nucleic Acid Res. 19:5081 (1991); Ohtsuka et al. J. Biol. Chem. 260:2605-2608 (1985); and Rossolini et al. Mol. Cell. Probes 8:91-98 (1994).IV. Methods of Use
[0113] The reporter construct polynucleotides provided herein have many applications, including, for example, the detection of ecDNA formation. Disclosed herein are methods of using the reporter construct polynucleotides, including, but not limited to, methods of detecting or monitoring the formation of circular DNA, e.g. , ecDNA, in a cell (for example, a mammalian cell, a plant cell, or an insect cell); methods of identifying compounds that can impair or enhance circular DNA. e.g, ecDNA, formation, for use in the treatment of diseases, such as cancer or other infectious diseases; and methods of identifying regulators, e.g., promoters or suppressors, of circular DNA, e.g., ecDNA, formation.
[0114] In general, the reporter construct polynucleotides disclosed herein are contacted with cells to produce cells that comprise the reporter construct polynucleotides. In some embodiments, a delivery system is used for introducing the reporter constructs into the cells (i.e., "‘contacting’7the cells) to produce cells that take up the reporter construct polynucleotides. Non-limiting examples of deliver}' systems include viral-based deliver}' systems (e.g. adenoviral (AAV) or lenti viral (LV) transduction) and nonviral-based (or physical or chemical based) delivery systems (e.g, electroporation, nucleofection, or transfection-based methods using, for example, Lipofectamine transfection reagents). Such methods are well known in the art and readily adaptable for use with the reporter construct polynucleotides, compositions, and methods described herein. See, for example, Chong, Zhi Xiong et al. PeerJ vol. 9 el 1165. 21 Apr. 2021, doi:10.7717 / peerj. l l 165; and Fus-Kujawa, Agnieszka et al. Frontiers in bioengineering and biotechnology vol. 9 701031. 20 Jul. 2021, doi: 10.3389 / fbioe.2021.701031.
[0115] In some embodiments, the reporter construct polynucleotides are used in a method of detecting or monitoring circular DNA, e.g.. ecDNA, presence in a cell (for example, an animalcell, a plant cell, or an insect cell). In some embodiments, cells comprising reporter construct polynucleotides as described herein are cultured and then analyzed for the presence of a detectable signal from a reporter polypeptide that is encoded by the reporter construct polynucleotides (thereby indicating the formation of ecDNA). In some embodiments, detection of a signal from a reporter polypeptide encoded by the reporter construct polypeptides indicates that circular DNA, e.g, ecDNA, is formed in the cells. In some embodiments, the method comprises detecting one or more fluorescence signals. In some embodiments, the reporter construct polynucleotide comprises the sequence as set forth in SEQ ID NO: 1 or SEQ ID NO: 2. In some embodiments, the detected signal is used to calculate a percentage of cells (e.g., number of fluorescent cells out of total number of cells) that contain circular DNA or ecDNA, where the calculated percentage is referred to as ecDNA efficiency or DNA circularization efficiency.
[0116] In some embodiments, the reporter construct polynucleotides can be used in a method of identifying compounds that can impair or enhance circular DNA, e.g., ecDNA, formation (i.e., regulators of ecDNA biogenesis, such as ecDNA biogenesis inhibitors that inhibit or block formation of ecDNA or ecDNA biogenesis enhancers that aid or improve in the formation of ecDNA). In some embodiments, cells (for example, animal, plant, or insect cells) comprising reporter construct polynucleotides are cultured and then analyzed for the presence of a first detectable signal from a reporter polypeptide that is encoded by the reporter construct polynucleotides. In some embodiments, detection of a signal from a reporter polypeptide encoded by the reporter construct polypeptides indicates that circular DNA. e.g, ecDNA, is present in the cells. In some embodiments, the cells comprising reporter construct polynucleotides are then contacted with one or more drugs or compounds of interest, and the cells are cultured and analyzed for the presence of a second detectable signal from the reporter polypeptide. In some embodiments, the compound(s) of interest are cultured with the cell prior to the addition of the reporting construct polynucleotide. In other embodiments, the drug(s) or compound(s) of interest are cultured with the cell concurrently with the addition of the reporting construct polynucleotide. In other embodiments, the compound(s) of interest are administered after the addition of the reporting construct polynucleotide. In some embodiments, the method comprises detecting one or more fluorescence signals. In some embodiments, the reporter construct polynucleotide comprises the sequence as set forth in SEQ ID NO: 1 or SEQ ID NO: 2.
[0117] In some embodiments, the method comprises comparing the second detectable signal to the first detectable signal. In some embodiments, the second detectable signal is recorded after the cells (for example, animal, plant, or insect cells) are contacted with compound(s) of interest, and when the second detectable signal is less than the first detectable signal (that is recorded before the cells are treated with compound(s)), the compound(s) decrease circular DNA presence or formation (e.g., ecDNA) in the cells. In contrast, when the second detectable signal is greater than the first detectable signal, the compound(s) increase circular DNA presence or formation (e.g., ecDNA) in the cells. In some embodiments, the first and second detectable signals are fluorescence signals. In some embodiments, the reporter construct polynucleotide comprises the sequence as set forth in SEQ ID NO: 1 or SEQ ID NO: 2.
[0118] In some embodiments, the percentage of cells that contain circular DNA or produced ecDNA is determined based on the detection of detectable signal for that the tested population of cells. In some embodiments, a first percentage of cells that contain circular DNA or ecDNA is determined based on a first detectable signal that is recorded before the cells are treated with compound(s). In some embodiments, a second percentage of cells that contain circular DNA or ecDNA is determined based on a second detectable signal that is recorded after the cells are treated with compound(s). In some embodiments, the second percentage of cells is less than the first percentage of cells, which indicates that the compound(s) decrease circular DNA presence or ecDNA formation. In some embodiments, the second percentage of cells is greater than the first percentage of cells, which indicates that the compound(s) increase circular DNA presence or ecDNA formation. In some embodiments, student’s t-test is used to determine the significance threshold of the fluorescence signals and / or the percentage of cells that contain circular DNA or produced ecDNA.
[0119] In some embodiments, the reporter construct polynucleotides are used in a method of identifying regulators, e.g, promoters or suppressors, of circular DNA, e.g, ecDNA, formation. In some embodiments, the method comprises a system for modifying the genome of a cell. In some embodiments, one or more genes in a cell may be disrupted or modified and the reporter construct polynucleotides disclosed herein are used to indicate the impact of the gene disruption / modification on circular DNA, e.g., ecDNA, presence or formation. In some embodiments, the reporter construct polynucleotide comprises the sequence as set forth in SEQ ID NO: 1 or SEQ ID NO: 2.a. Detection of Circular DNA Using the Reporter Construct Polynucleotides
[0120] Disclosed herein are methods for detecting the formation of circular DNA, particularly, ecDNA, in a cell using a linear reporter construct polynucleotide as described in this Section III of this disclosure. Circularization of the reporter construct polynucleotide into a circular DNA molecule can be detected in these methods and reflect a level of ecDNA formation in the cell. In some embodiments, the methods comprise introducing a linear reporter construct polynucleotide as described herein into a cell and assessing the cell to detect whether the reporter construct polynucleotide has become circularized. In some embodiments, as described in more detail below, the cell may be treated with a drug or exposed to particular conditions to assess the impact on circularization of the reporter construct polynucleotide and, thus, ecDNA formation (biogenesis).
[0121] Methods for detecting circularization of the reporter construct polynucleotide may be selected based on the nature of the signal that is produced by the reporter polypeptide(s) of the reporter construct polynucleotide. In some embodiments, the detection method comprises the detection of a fluorescence signal from the reporter polypeptide(s). In some embodiments, the detection method comprises fluorescence microscopy. In some embodiments, the detection method combines cell sorting with the detection of fluorescence signals, for example, by fluorescence-activated cell sorting (FACS). In some embodiments, the method comprises detecting one, two, or more different fluorescence signals produced by the reporter polypeptide(s) expressed from the reporter construct polynucleotides. In some embodiments, the method comprises detecting eGFP and / or DsRed fluorescence. Methods and systems related to fluorescence microscopy and imaging in cells and / or visualization of DNA are readily available to one of ordinary' skill in the art, and are discussed, for example, in Ettinger, Andreas, and Torsten Wittmann. Methods in Cell Ciology vol. 123 (2014): 77-94. doi: 10. 1016 / B978-0-12-420138-5.00005-7; Zhao. Yiheng et al. eLife vol. 11 e81412. 18 Oct. 2022, doi:10.7554 / eLife.8I4I2; Lichtman, J., Conchello, JA. Nat Methods 2, 910-919 (2005). doi: 10.1038 / nmeth817.
[0122] In some embodiments, the detection method comprises performing polymerase chain reaction (PCR). In some embodiments, the PCR method comprises using a pair of primers that spans the junction where the reporter polypeptide coding sequence and the promoter come together only when they are in a circular DNA (i.e., the end-to-end junction) and not in a linear DNA (FIG. 1A. ID). In some embodiments, the primers are divergent primers which face away from each other on linear DNA, so that the PCR method can only amplify circular DNA, andnot linear DNA. Said another way, in some instances, the primers selected for the PCR method may comprise a forward primer that binds to a portion of the reporter construct polynucleotide near its 3’ end and a reverse primer that binds to a portion of the reporter construct polynucleotide near its 5’ end. When the reporter construct polynucleotide is circularized, a PCR method using such divergent primers will produce a PCR product. When the reporter construct polynucleotide remains in its linear form, no PCR products are produced. PCR products can be visualized using methods available to one of ordinary skill in the art. such as, without limitations, an agarose gel. PCR methods related to detecting the end-to-end junction of a reporter construct polynucleotide are discussed in Brown, D D, and I B Dawid. Science (New York, N.Y.) vol. 160.3825 (1968): 272-80. doi: 10.1126 / science.l60.3825.272. Additional PCR methods are described below in Section V.
[0123] In some embodiments, the detection method comprises droplet digital PCR (ddPCR). DdPCR is a method based on water-oil emulsion droplet technology that can be used to quantify the number of circular DNA, e.g., ecDNA, molecules, that are produced from reporter construct polynucleotides disclosed herein. Methods of using ddPCR to detect circular DNA are discussed in Lange. Joshua T et al. Nature genetics vol. 54,10 (2022): 1527-1533. doi: 10. 1038 / s41588-022-01177-x.
[0124] In some embodiments, detection of the formation of circular DNA, e.g., ecDNA, can be used to quantitate the percentage of genomic DNA that is being converted into ecDNA as reflected by the number of circularized reporter constructs in the treated cells. Circular DNA concentrations determined by ddPCR may be used to determine circular DNA, e.g., ecDNA, frequency according to the formula below: ecDNA concentration circular DNA (%) = -GAPDH concentration
[0125] In some embodiments, the detection method comprises a sequencing method. Many sequencing methods are readily available to one of ordinary skill in the arts. Sequencing methods are discussed in detail, for example, in Goodwin, S., McPherson. J. & McCombie, W. Nat Rev Genet 17, 333-351 (2016). https: / / doi.org / 10.1038 / nrg.2016.49; Pervez, Muhammad Tariq et al. BioMed research international vol. 2022 3457806. 29 Sep. 2022, doi: 10. 1155 / 2022 / 3457806; Slatko, Barton E et al. Current protocols in molecular biology vol. 122,1 (2018): e59. doi: 10.1002 / cpmb.59.
[0126] The methods disclosed herein may be used with a variety of cells including eukaryotic cells, prokaryotic cells, and plant cells. In some embodiments, the cells are animal cells (e.g., mammalian, mouse, primate, mouse, etc.) or human cells. The cells may be established research cell lines or isolated and / or derived from an animal (e.g.. mammalian, mouse, primate, mouse, etc.) or human subject. In some embodiments, the cells are cervical cancer HeLa cells, non-small cell lung cancer PC9 cells, colon cancer HCT116 cells, or HEK293T cells. In some embodiments, the cells are cancer cells isolated from human subjects. As used herein the terms ■‘cancer” and '‘tumor” are used to indicate malignant tissue. The term “cancer” is also used to refer to the disease associated with the presence of malignant tumor cells in an individual, and the term “tumor” is used herein to refer to a plurality of cancer cells that are physically- associated with each other. Cancer cells are malignant cells that give rise to cancer, and tumor cells are malignant cells that can form a tumor and thereby give rise to cancer. The term “cancer,” as used herein, may be used to describe a solid tumor, metastatic cancer, or non- metastatic cancer. The term also encompasses a circulating tumor cell. In certain embodiments, the cancer may originate in the pancreas, colon, rectum, lung, bladder, blood, bone, bone marrow, brain, breast, esophagus, duodenum, small intestine, large intestine, gum, head, kidney, liver, nasopharynx, neck, ovary, pancreas, prostate, skin, stomach, testis, tongue, or uterus. In some embodiments, the cancer cell is a glioblastoma cell, a sarcoma cell, a esophageal cancer cell, a colon cancer cell, a cervical cancer cell, or a lung cancer cell.
[0127] In embodiments, animal cells as described herein may be derived from mammalian immortalized cell lines or other pnmary cell lines derived from a subject (for example, a human subject). Mammalian cells, for example, may be derived from any cellular germ layer, for example, mesoderm, endoderm, or ectoderm, or from any organ of the body (for example, liver hepatocytes or kidney cells, such as human embryonic kidney). In embodiments, mammalian cells may be placental or embryonic. Cells may be stem cells (for example, pluripotent cells such as human embryonic cells, or multipotent cells such as hematopoietic stem cells (HSCs) or bone-marrow derived mesenchymal stem cells (BM-MSCs)); bone cells (for example, osteoblasts or osteoclasts): blood cells (for example, white blood cells); muscle cells (also known as myocytes); sperm cells; a female egg; skin cells; endothelial cells; epithelial cells; fat cells; cells of the central or peripheral nervous system (for example, neurons, glia, and pericytes); or cells from an organ in the body, such as kidney or liver.
[0128] Without intending to be limiting, in embodiments, the cell can be: a neuron; a glial cell (i.e., an astrocyte, oligodendrocyte, or Schwann cell); a pericy te; a fibroblast; an intestinalepithelial cell; a mesenchymal cell; a T cell; a cancer cell; a stem cell; a chondrocyte; an osteoblast; an osteoclast; an osteocyte; a HSC; a BM-MSC; an induced pluripotent stem cells; an embryonic stem cells; a granulocyte; an agranulocyte; a skeletal, cardiac, or smooth muscle myocyte; or an adipocyte (i.e., a white or brown adipocyte).
[0129] In other embodiments, the cells are plant cells. In embodiments where the cell is a plant cell, a plant cell can be a parenchymal, collenchymal, sclerenchymal, xylem, phloem, meristematic, or epidermal plant cell. Generally, plant cells according to the present disclosure may include eukaryotic cells with large central vacuoles, cell walls containing cellulose, and plastids such as chloroplasts and chromoplasts. Additionally, plant cells: may be non-motile; may make their own food (i.e., are autotrophic); may reproduce asexually by vegetative propagation or sexually; may contain an outer cell wall and a large central vacuole; may contain photosynthetic pigments (i.e., a chlorophyll) that can be present in the plastids; and may have different organelles for anchorage, reproduction, support and photosynthesis.
[0130] Plant cells can be cultured according to the methods known in the art. For example, and without intending to be limiting, plant cells can be cultured by culture processes such as seed culture, meristem culture, callus culture, and bud culture. In another culture method, plant tissues can be placed on a gel substrate such as Murashige and Skoog (often called MS media, MSO, or MSO) or Gamborg B5 medium. Plant tissues may also be placed into a liquid medium, as is the case with cell suspension culture. The plant culture media formulation may include macronutrients, micronutrients, vitamins and organic supplements, amino acids and nitrogen supplements, plant growth hormones and plant growth regulators (PGRs), and will vary depending on specific plant needs.
[0131] In some embodiments, the method further comprises culturing the cell in the presence of a compound corresponding to a selectable marker that is present in the reporter construct polynucleotide. A variety of selectable markers and corresponding markers may be used and are readily available to one of ordinary skill in the art. Selectable markers are discussed above in detail.
[0132] In some embodiments, the method further comprises detecting circularization of the reporter construct polynucleotide as a way to detect circular DNA, e.g., ecDNA, presence in cells. A variety of detection methods are readily available to one of ordinary' skill in the art. Non-limiting examples of methods of detection include fluorescence microscopy, fluorescence-activated cell sorting (FACS), polymerase chain reaction (PCR), droplet digitalPCR (ddPCR), and sequencing. Methods of detecting reporter construct polynucleotides are discussed in detail above.
[0133] Accordingly, an aspect of the present disclosure provides a method for evaluating the presence of ecDNA in a cell, the method comprising, consisting of, or consisting essentially of contacting a cell with a reporter construct polynucleotide as provided herein, culturing the cell, and assessing the fluorescence of the reporting construct in the cell to determine the presence of a circularized form of the reporter construct polynucleotide, which reflects a level of ecDNA biogenesis in the cell.
[0134] In other embodiments, the method comprises culturing the cells in the presence of a compound of interest. In some embodiments, the compound of interest is cultured with the cell prior to the addition of the reporting construct. In other embodiments, the compound of interest is cultured with the cell concurrently with the addition of the reporting construct. In other embodiments, the compound of interest is administered after the addition of the reporting construct. b. Treatment Resistance Studies
[0135] In embodiments, described herein are methods of screening for cellular mechanisms of ecDNA formation in cells, in addition to screening compounds that may prevent or inhibit ecDNA formation in the cells. The cells can be isolated cells or cells in a subject or plant. As used herein, the term “subject” refers to both human and nonhuman animals. The term “nonhuman animals” of the disclosure includes all vertebrates, e.g., mammals and nonmammals, such as nonhuman primates, sheep, dog, cat, horse, cow, chickens, amphibians, reptiles, and the like. The methods and compositions disclosed herein can be used on a sample either in vitro (for example, on isolated cells or tissues) or in vivo in a subject (i.e. living organism, such as a patient). The compositions and methods provided herein may be used in medical (i.e., used to treat a human subject) and veterinary (i.e., used to treat non-human subjects) settings. The term subject also includes insects and plants.
[0136] Methods as described herein can comprise contacting one or more cells with a linear reporter construct as described herein, thereby introducing the linear reporter construct into one or more cells, wherein the one or more cells are isolated cells or cells in an organism. Cells that contain the reporter can then be subjected to genetic manipulation (i.e., knockout, knockdown, etc.) of proteins of interest. For example, in animal cells, a protein of interest can be Ligase 4 (Lig4), XCCR4, or a protein of the BRCA1-A complex (e.g., BABAM2,ABRAXAS 1. BRCC3, or Rap80). For example, in plant cells, a protein of interest can be 5- enolpyruvylshikimate-3-phosphate synthase (EPSPS). Following manipulation, a signal from the reporter polypeptide can then be read, indicating the formation of ecDNA (or lack thereof) in the cell.
[0137] Such methods can further involve contacting cells comprising a reporter as described herein with a candidate ecDNA inhibitor compound, and optionally in conjunction with a known active compound (e.g., a cancer drug or a herbicide (e.g., glyphosate), for animal cells and plant cells, respectively). In some embodiments, the candidate ecDNA inhibitor compound can inhibit or otherwise suppress ecDNA formation in the cell. Therapeutic agents can be directed at cellular machinery that drives ecDNA formation in the plant cells, such as those targets identified in the screens in the preceding paragraph. c. Exemplary Methods
[0138] In some embodiments, the method of detecting circular DNA presence, e.g., ecDNA, in a cell comprises (i) contacting a plurality of cells with any reporter construct polynucleotide disclosed herein, (ii) culturing the plurality of cells comprising the reporter construct polynucleotide, and (iii) detecting a first signal from the first reporter polypeptide. In some embodiments, detection of the first signal indicates circular DNA presence in the plurality of cells.
[0139] In some embodiments, the method of determining whether a compound impairs circular DNA formation in a cell comprises: (i) contacting a plurality' of cells with any reporter construct polynucleotide disclosed herein, (ii) culturing the plurality of cells comprising the reporter construct polynucleotide, (iii) detecting a first signal from a first reporter polypeptide expressed from the reporter construct polynucleotide, (iv) contacting the plurality of cells with a compound of interest, (v) detecting a second signal from the first reporter polypeptide, and (iv) comparing the second signal from the first reporter polypeptide to the first signal from the first reporter polypeptide. In some embodiments, the compound impairs circular DNA formation if the second signal is less than the first signal.
[0140] In some embodiments, the method comprises detecting or monitoring the uptake of the reporter construct polynucleotides into cultured cells thereby allowing for the quantification of the number of cells that can produce circular DNA, e.g., ecDNA from the reporter construct polynucleotides. In some embodiments, detection of a signal from a reporter polypeptideencoded by the reporter construct polypeptides labels the cells that comprise the reporter construct polynucleotides.
[0141] In some embodiments, the method further comprises culturing the cells with a selectable marker. Culturing of cells in the presence of a corresponding selectable marker compound can be helpful in selecting for cells that comprise reporter construct polynucleotides.
[0142] In some embodiments, the cells are cancer cells.
[0143] In some embodiments, the reporter construct polynucleotide comprises the sequence as set forth in SEQ ID NO: 1 or SEQ ID NO: 2.V. Therapeutic Agents Targeting ecDNA Biogenesis
[0144] Also described herein are therapeutic agents that can reduce, suppress, or otherwise eliminate intracellular ecDNA formation in one or more cells (referred to as “ecDNA inhibitors). ecDNA inhibitors as described herein can be, for example, a small molecule, inhibitory nucleic acid, genetic modification, antibody, or peptide, that can target a member of a ecDNA biogenesis pathway. ecDNA inhibitors as described herein can target the ecDNA biogenesis pathway, for example, by inhibiting or suppressing one or more proteins that are involved in ecDNA biogenesis.
[0145] In embodiments, ecDNA inhibitors target one or more ecDNA biogenesis pathwayproteins such as ligase 4 (i.e., gene product of the LIG4 gene. NCBI Gene ID: 3981, the disclosure of which is incorporated herein by reference). X-ray repair cross complementing 4 (XRCC4; gene product of the X7?CC4 gene, NCBI Gene ID: 7518, the disclosure of which is incorporated herein by reference), or the Lig4 / XRCC4 complex. ecDNA inhibitors can target the Li 4 ligase domain (1-654 aa), and either of the two BRCT domains (654-911 aa; or both). Lig4 only can be stabilized and functional by directly interacting with XRCC4 via its two BRCT domains, and ecDNA inhibitors can target XRCC4 as well to prevent Lig4 stabilization.
[0146] In embodiments, ecDNA inhibitors target one or more ecDNA biogenesis pathway proteins such as those that form the BRCA1-A complex, such as, BABAM2 (encoded by the BABAM2 gene, NCBI Gene ID: 9577. the disclosure of which is incorporated herein by reference), ABRAXAS1 (encoded by the ABRAXAS1 gene, NCBI Gene ID: 84142, , the disclosure of which is incorporated herein by reference), and RAP80 encoded by gene UIMC1, NCBI Gene ID: 51720, the disclosure of which is incorporated herein by reference).
[0147] Any agent that reduces, decreases, counteracts, attenuates, inhibits, blocks, downregulates, or eliminates in any way the expression, stability, or activity (e.g.. ceDNA formation) of a protein of an ecDNA biogenesis pathway (e.g., Lig4, XRCC4, or a protein of the BRCA1-A complex) can be used in the present methods as an ecDNA inhibitor. ecDNA inhibitors can be small molecule compounds, peptides, polypeptides, nucleic acids, antibodies, e.g., blocking antibodies or antibody fragments, or any other molecule that reduces, decreases, counteracts, attenuates, inhibits, blocks, downregulates, or eliminates in any way the expression, stability and / or activity of a pathway of an ecDNA biogenesis pathway (e.g., Lig4, XRCC4, or a factor of the BRCA1-A complex) . In particular embodiments, a ecDNA biogenesis pathway is inhibited using a small molecule inhibitor such as using gene silencing, or by modifying the a gene of a protein of an ecDNA biogenesis pathway (such as those described above) using a CRISPR-Cas system so as to reduce or eliminate its expression.
[0148] In some embodiments, the ecDNA inhibitor decreases the activity (e.g., phosphatase activity), stability' or expression of protein of an ecDNA biogenesis pathway by at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%. 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or more relative to a control level, e.g., a level determined in the absence of the inhibitor, in vivo or in vitro.
[0149] The efficacy of inhibitors can be assessed in any of a variety of ways, including in vitro and in vivo methods. For example, the ecDNA formation activity' can be assessed using a method as described herein, in particular, a method comprising a reporter as described herein.
[0150] In some embodiments, the ecDNA inhibitor is considered effective if the ecDNA formation as described herein is decreased by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90% or more as compared to the reference value, e.g., the value in the absence of the inhibitor, in vitro or in vivo. In some embodiments, an ecDNA inhibitor (e.g. , an RNAi molecule) is considered effective if the level of expression of a protein of an ecDNA biogenesis pathway is decreased by at least 1.5-fold, at least 2-fold, at least 3-fold, at least 4-fold, at least 5-fold, at least 6-fold, at least 7-fold, at least 8-fold, at least 9-fold, at least 10-fold or more as compared to the reference value.
[0151] The efficacy of inhibitors can also be assessed, e.g. , by detection of decreased reporter polynucleotide (e.g., eGFP) expression, which can be analyzed using routine techniques such as fluorescent microscopy. RT-PCR, Real-Time RT-PCR. semi-quantitative RT-PCR, quantitative polymerase chain reaction (qPCR), quantitative RT-PCR (qRT-PCR), multiplexedbranched DNA (bDNA) assay, microarray hybridization, or sequence analysis (e.g., RNA sequencing (“RNA-Seq”)). Methods of quantifying polynucleotide expression are described, e.g., in Fassbinder-Orth, Integrative and Comparative Biology, 2014, 54:396-406; Thellin et al., Biotechnology Advances , 2009, 27:323-333; and Zheng et al., Clinical Chemistry, 2006, 52:7 (doi: 10 / 1373 / clinchem.2005.065078). In some embodiments, real-time or quantitative PCR or RT-PCR is used to measure the level of a polynucleotide (e.g., mRNA) in a biological sample. See, e.g.. Nolan et al.. Nat. Protoc, 2006. 1 : 1559-1582; Wong et al., BioTechniques, 2005, 39:75-75. Quantitative PCR and RT-PCR assays for measuring gene expression are also commercially available (e.g, TaqMan® Gene Expression Assays, ThermoFisher Scientific).
[0152] In some embodiments, the ecDNA inhibitor is considered effective if the level of expression of an ecDNA biogenesis pathway protein is decreased by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%. at least 60%. at least 70%. at least 80%. at least 90% or more as compared to the reference value, e.g., the value in the absence of the inhibitor, in vitro or in vivo. In some embodiments, a ecDNA inhibitor is considered effective if the level of expression of an ecDNA biogenesis pathway protein is decreased by at least 1.5-fold, at least 2-fold, at least 3-fold, at least 4-fold, at least 5-fold, at least 6-fold, at least 7-fold, at least 8- fold, at least 9-fold, at least 10-fold or more as compared to the reference value.
[0153] The effectiveness of a ecDNA inhibitor can also be assessed by detecting protein expression or stability, e.g., using routine techniques such as immunoassays, two-dimensional gel electrophoresis, western blot, and quantitative mass spectrometry, all of which are known to those skilled in the art. Protein quantification techniques are generally described in ■‘Strategies for Protein Quantitation,” Principles of Proteomics, 2nd Edition, R. Twyman, ed.. Garland Science, 2013. In some embodiments, protein expression or stability is detected by immunoassay, such as but not limited to enzyme immunoassays (EIA) such as enzyme multiplied immunoassay technique (EMIT), enzyme-linked immunosorbent assay (ELISA), IgM antibody capture ELISA (MAC ELISA), and microparticle enzyme immunoassay (MEIA); capillary electrophoresis immunoassays (CEIA); radioimmunoassays (RIA); immunoradiometric assays (IRMA); immunofluorescence (IF); fluorescence polarization immunoassays (FPIA); and chemiluminescence assay s (CL). If desired, such immunoassays can be automated. Immunoassays can also be used in conjunction with laser induced fluorescence (see, e.g., Schmalzing et al.. Electrophoresis, 18:2184-93 (1997); Bao, J. Chromatogr. B. Biomed. Sci., 699:463-80 (1997)).
[0154] For determining whether levels of an ecDNA biogenesis pathway protein are decreased in the presence of an ecDNA inhibitor, the method comprises comparing the level of the protein in the presence of the inhibitor to a reference value, e.g. , the level in the absence of the inhibitor. In some embodiments, a an ecDNA biogenesis pathway protein is decreased in the presence of an inhibitor if the level of the an ecDNA biogenesis pathway protein is decreased by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%. at least 90% or more as compared to the reference value. In some embodiments, an ecDNA biogenesis pathway protein is decreased in the presence of an inhibitor if the level of the an ecDNA biogenesis pathway protein is decreased by at least 1.5- fold, at least 2-fold, at least 3-fold, at least 4-fold, at least 5-fold, at least 6-fold, at least 7-fold, at least 8-fold, at least 9-fold, at least 10-fold or more as compared to the reference value. a. Small molecules
[0155] In some embodiments of the invention, ecDNA formation is inhibited by the administration of a small molecule inhibitor (for example, one that targets a protein of an ecDNA biogenesis pathway). Any small molecule inhibitor can be used that reduces, e.g., by 10%, 15%, 20%, 25%, 30%. 35%. 40%. 45%. 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or more, the expression, stability or activity of ecDNA formation relative to a control, e.g., the expression, stability or activity in the absence of the inhibitor. In particular embodiments, small molecule inhibitors are used that can bind to Lig4 (for example on that targets the ligase or BRCT domains), XCCR4. or a BRCA1-A complex protein.
[0156] In some embodiments, the small molecule binds to Ligase 4 (Lig4). XCCR4. or a protein of the BRCA1-A complex (e.g., BABAM2, ABRAXAS1, BRCC3, or Rap80) and disrupts the interaction of one of these proteins with one of its binding partners. For example, Lig4 interacts with XRCC4 as well as other proteins. For example, a complex of Lig4-XRCC4 interacts with DNA protein kinase (DNA-PK) on DNA ends. The various proteins in the BRCA1-A complex interact with each other and with other proteins as well. In some embodiment, with respect to Lig4, in some embodiments, the small molecule binds in the catalytic domain of Lig4. In some embodiments, the small molecule binds to a region of Lig4 that interacts with XCCR4. In some embodiments, this region is betw een amino acids 600- 800, and in particular embodiments, between amino acids 767-783 (as described in Grawunder, U„ et al. Curr Biol. 1998; 8(15): 873-876; doi: 10.1016 / s0960-9822(07)00349-l, which is incorporated by reference herein). In some embodiments, the small molecule binds to XCCR4 in a region that interacts with this region of Lig4. In particular embodiments, with respect toLig4, the small molecule can be SCR7 (ApexBio Technology Catalog #: A8705), SCR130, SCR116, or SCR132. In some embodiments, the small molecule binds to the BRCT domain of BRCA1. In some embodiments, the small molecule disrupts the interactions between one or more of BRCA1, BABAM2, ABRAXAS1, BRCC3, or Rap80, as described, for example, in Her, J., et al., Acta Biochimica et Biophysica Sinica, 2016; 48(7):658-664; doi: 10. 1093 / abbs / gmw047).
[0157] Useful small molecule ecDNA inhibitors can be identified using the screening methods described above in this disclosure. Candidate small molecule (and other) inhibitors can be identified and / or assessed using molecular docking analysis. For example, using predictive software such as AlphaFold© or other artificial intelligent algorithms to provide protein structure and conformational information, and optionally, followed by medicinal chemistry and synthesis of candidate molecules. Alternatively, shotgun screening from available compound libraries or randomly synthesize molecules for inhibition of targets. b. Inhibitory nucleic acids
[0158] In some embodiments, the ecDNA inhibitor comprises an inhibitory nucleic acid, e.g, antisense DNA or RNA, small interfering RNA (siRNA), microRNA (miRNA), or short hairpin RNA (shRNA). In some embodiments, the inhibitory RNA targets a sequence that is identical or substantially identical (e.g, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical) to a target sequence of an ecDNA biogenesis pathway polynucleotide (e.g, a portion comprising at least 20, at least 30, at least 40, at least 50. at least 60, at least 70. at least 80, at least 90, or at least 100 contiguous nucleotides, e.g. from 20-500. 20-250, 20-100. 50-500, or 50-250 contiguous nucleotides of an ecDNA biogenesis pathway protein-encoding gene or gene product thereof (e.g. , the human LIG4 gene, XRCC4 gene, a protein opr gene product of a gene encoding a protein of the BRCA1-A complex, such as those described above).
[0159] In some embodiments, the methods described herein comprise silencing the I.IG4. XRCC4, BABAM2, ABRAXAS1, or UIMC1 gene using an shRNA or siRNA. A shRNA is an artificial RNA molecule with a hairpin turn that can be used to silence target gene expression via the siRNA it produces in cells. See, e.g, Fire et. al., Nature 391 :806-811, 1998; Elbashir et al., Nature 411 :494-498, 2001; Chakraborty et al., Mol Ther Nucleic Acids 8: 132-143, 2017; and Bouard et al., Br. J. Pharmacol. 157: 153-165, 2009. In some embodiments, a cell iscontacted with a modified RNA or a vector comprising a polynucleotide that encodes an shRNA or siRNA capable of hybridizing to a portion of an mRNA transcription product of a I.IG4. XRCC4, BABAM2, ABRAXAS 1. or UIMC1 gene. In some embodiments, the vector further comprises appropriate expression control elements known in the art, including, e.g., promoters (e.g. , inducible promoters or tissue specific promoters), enhancers, and transcription terminators.
[0160] In some embodiments, the inhibitor is an ecDNA biogenesis pathway protein-specific microRNA (miRNA or miR). A microRNA is a small non-coding RNA molecule that functions in RNA silencing and post-transcriptional regulation of gene expression. miRNAs base pair with complementary' sequences within the mRNA transcript. As a result, the mRNA transcript may be silenced by one or more of the mechanisms such as cleavage of the mRNA strand, destabilization of the mRNA through shortening of its poly(A) tail, and decrease in the translation efficiency of the mRNA transcript into proteins by ribosomes. In embodiments, miRNA can target Lig4, XCCR4, BABAM2, ABRAXAS 1, or Rap80 transcripts (i.e., mRNA).
[0161] In some embodiments, the inhibitor is an antisense oligonucleotide, e.g., an RNase H-dependent antisense oligonucleotide (ASO). ASOs are single-stranded, chemically modified oligonucleotides that bind to complementary sequences in target mRNAs and reduce gene expression both by RNase H-mediated cleavage of the target RNA and by inhibition of translation by steric blockade of ribosomes. In some embodiments, the oligonucleotide is capable of hybridizing to a portion of a SHP1 mRNA. In some embodiments, the oligonucleotide has a length of about 10-30 nucleotides (e.g., 10, 12, 14. 16. 18. 20, 22, 24, 26, 28, or 30 nucleotides). In some embodiments, the oligonucleotide has 100% complementarity to the portion of the mRNA transcript it binds. In other embodiments, the DNA oligonucleotide has less than 100% complementarity' (e.g., 95%, 90%, 85%, 80%, 75%, or 70% complementarity) to the portion of the mRNA transcript it binds, but can still form a stable RNA:DNA duplex for the RNase H to cleave the mRNA transcript.
[0162] Suitable antisense molecules, siRNA, miRNA, and shRNA can be produced by standard methods of oligonucleotide synthesis or by ordering such molecules from a contract research organization or supplier by providing the polynucleotide sequence being targeted. The manufacture and deployment of such antisense molecules in general terms may be accomplished using standard techniques descnbed in contemporary' reference texts: for example, Gene arid Cell Therapy: Therapeutic Mechanisms and Strategies, 4thedition by' N.S.Templeton; Translating Gene Therapy to the Clinic: Techniques and Approaches, 1stedition by J. Laurence and M. Franklin: High-Throughput RNAi Screening: Methods and Protocols (Methods in Molecular Biology) by D.O. Azorsa and S. Arora; and Oligonucleotide-Based Drugs and Therapeutics: Preclinical and Clinical Considerations by N. Ferrari and R. Segui.
[0163] Inhibitory' nucleic acids can also include RNA aptamers, which are short, synthetic oligonucleotide sequences that bind to proteins (see, e.g., Li et al., Nuc. Acids Res. (2006), 34:6416-24). They are notable for both high affinity and specificity for the targeted molecule, and have the additional advantage of being smaller than antibodies (usually less than 6 kD). RNA aptamers with a desired specificity are generally selected from a combinatorial library, and can be modified to reduce vulnerability to ribonucleases, using methods known in the art.
[0164] In some embodiments, endoribonuclease-prepared siRNAs (esiRNAs) are used to inhibit an ecDNA biogenesis pathway protein. esiRNAs are a mixture of siRNA oligos resulting from cleavage of long double-stranded RNA (dsRNA) with an endoribonuclease such as Escherichia coli RNase III or Dicer. esiRNAs are a heterogeneous mixture of siRNAs that all target the same mRNA sequence (e.g, Lig4, XCCR4, BABAM2, ABRAXAS 1, or Rap80 )• c. Genetic modification
[0165] In some embodiments, the genes encoding Lig4, XCCR4, BABAM2, ABRAXAS 1, or Rap80 are inhibited by genomic modification, e.g., by deleting the gene encoding Lig4, XCCR4, BABAM2, ABRAXAS 1, or Rap80 in a cell or by introducing a mutation in the gene that, e.g., decreases or abolishes its expression, activity, or stability. Such methods can be carried out using any suitable method known in the art, e.g.. using a CRISPR-Cas system. The CRISPR-Cas system comprises at least one guide RNA (typically a single guide RNA, or sgRNA), an RNA-guided nuclease (such as Cas9 or Cpfl), and optionally a homologous donor template. A homologous donor template can be used to introduce specific modifications into the genome by homologous recombination. In the absence of a homologous donor template, however, cleavage of the gRNA target sequence can still inactivate a gene through the introduction of small insertions or deletions (indels).
[0166] Many genome modifying systems and corresponding nucleases may be used, for example, without limitations, a homing nuclease polypeptide; a FokI polypeptide; a transcription activator-like effector nuclease (TALEN) polypeptide; a MegaTAL polypeptide; a meganuclease polypeptide; a zinc finger nuclease (ZFN); an ARCUS nuclease; and the like.The meganuclease can be engineered from an LADLIDADG homing endonuclease (LHE). A megaTAL polypeptide can comprise a TALE DNA binding domain and an engineered meganuclease. See, e.g.. International Patent Application Publication No. WO 2004 / 067736 (homing endonuclease); Umov et al. (2005) Nature 435:646 (ZFN); Mussolino et al. (2011) Nude. Acids Res. 39:9283 (TALE nuclease); Boissel et al. (2013) Nucl. Acids Res. 42:2591 (MegaTAL). In some embodiments, the genome modifying system is a CRISPR-Cas system comprising a CRISPR-Cas nuclease and guide polynucleotides, e.g, guide RNAs (gRNAs). In some embodiments, for example, the CRISPR-Cas nuclease is CRISPR-Cas9. See, e.g., Moller, et al. Nucleic Acids Research (2018).
[0167] The CRISPR-Cas nuclease can be any of a variety of CRISPR-Cas nucleases. CRISPR-Cas nucleases can be derived from a variety of bacterial species including, but not limited to, Veillonella atypical, Fusobacterium nucleatum, Filifactor alocis, Solobacterium moorei. Coprococcus cams. Treponema deniicola. Peptoniphilus duerdenii. Catenibacterium mitsuokai, Streptococcus mutans, Listeria innocua, Staphylococcus pseudintermedius, Acidaminococcus intestine, Olsenella uli, Oenococcus kitaharae, Bifidobacterium bifidum, Lactobacillus rhamnosus, Lactobacillus gasseri. Finegoldia magna, Mycoplasma mobile, Mycoplasma gallisepticum. Mycoplasma ovipneumoniae. Mycoplasma canis. Mycoplasma synoviae, Eubacterium rectale, Streptococcus thermophilus, Eubacterium dolichum, Lactobacillus coryniformis subsp. Torquens, Ilyobacter polytropus, Ruminococcus albus, Akkermansia muciniphila, Acidothermus cellulolyticus , Bifidobacterium longum, Bifidobacterium dentium, Corynebacterium diphtheria, Elusimicrobium minutum, Nitratifractor salsuginis, Sphaerochaeta globus, Fibrobacter succinogenes subsp. Succinogenes, Bacteroides fragilis, Capnocytophaga ochracea, Rhodopseudomonas palustris, Prevotella micans, Prevotella ruminicola. Flavobaclerium columnare, Aminomonas paucivorans, Rhodospirillum rubrum. Candidatus Puniceispirillum marinum, Verminephrobacter eiseniae, Ralstonia syzygii. Dinoroseobacter shibae, Azospirillum. Nitrobacter hamburgensis, Bradyrhizobium. Wolinella succinogenes, Campylobacter jejuni subsp. Jejuni, Helicobacter mustelae, Bacillus cereus, Acidovorax ebreus, Clostridium perfringens. Parvibaculum lavamentivorans, Roseburia intestinalis , Neisseria meningitidis, Pasteurella multocida subsp. Multocida, Sutterella wadsworthensis. proteobacterium. Legionella pneumophila. Parasutterella excrementihominis . Wolinella succinogenes, and Francisella novicida. Suitable CRISPR-Cas nucleases are described in detail below. Examples of CRISPR-Cas nucleases are CRISPR-Cas endonucleases (e.g.. class 2 CRISPR-Casnucleases such as a type II, pe V, or type VI CRISPR-Cas nuclease). In some embodiments, the CRISPR-Cas nuclease is a type II CRISPR-Cas nuclease. In some embodiments, the ty pe II CRISPR-Cas nuclease is a Cas9 polypeptide. In some embodiments, the CRISPR-Cas nuclease is a type V CRISPR-Cas nuclease, e.g., a Casl2a, a Casl2b, a Casl2c, a Casl2d, a Casl2e, a Cpfl, a C2cl, or a C2c3 polypeptide. In some embodiments, the CRISPR-Cas nuclease is a type VI CRISPR-Cas nuclease, e.g., a Casl3a, a Casl3b, a Casl3c. a Casl3d, a C2c2 (also referred to as Casl3a) polypeptide. In some embodiments, the CRISPR-Cas nuclease is a Casl4 polypeptide. In some embodiments, the CRISPR-Cas nuclease is a Casl4a polypeptide, a Casl4b polypeptide, or a Casl4c polypeptide. In some embodiments, a suitable CRISPR-Cas nuclease is a CasX or a CasY polypeptide. CasX and CasY polypeptides are described in Burstein et al. (2017) Nature 542:237.
[0168] Also suitable for use is a variant CRISPR-Cas nuclease, where the variant is a high- fidelity or enhanced specificity CRISPR-Cas nuclease with reduced off-target effects and robust on-target cleavage. Non-limiting examples of CRISPR-Cas nuclease variants yvith improved on-target specificity include the SpCas9 (K855A), SpCas9 (K810A / K1003A / R1060A) (also referred to as eSpCas9(1.0)), and SpCas9 (K848A / K1003A / R1060A) (also referred to as eSpCas9(l. 1)) variants described in Slaymaker et al. Science, 351(6268): 84-8 (2016), and the SpCas9 variants described in KI einstiver et al. Nature, 529(7587):490-5 (2016) containing one, two, three, or four of the following mutations: N497A, R661A, Q695A, and Q926A (e.g.. SpCas9-HFl contains all four mutations).
[0169] Also suitable for use. e.g, when fused with a second enzyme with nicking of DNA cleaving activity, is a variant CRISPR-Cas nuclease, yvhere the variant CRISPR-Cas nuclease has reduced or no nucleic acid cleavage activity. For example, the dCas9 variant (Jinek et al. Science, 2012, 337:816-821; Qi et al. Cell, 152(5): 1173- 1183) contains tyvo silencing mutations of the RuvCl and HNH nuclease domains (D10A and H840A). In some embodiments, the dCas9 polypeptide from Streptococcus pyogenes comprises at least one mutation at position D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, A987 or any combination thereof. Descriptions of such dCas9 polypeptides and variants thereof are provided in, for example, International Patent Application Publication No. WO 2013 / 176772. The dCas9 enzyme can contain a mutation at D10. E762, H983. or D986, as well as a mutation at H840 or N863. In some instances, the dCas9 enzyme can contain a D10A or DION mutation. Also, the dCas9 enzyme can contain a H840A, H840Y, or H840N. In some embodiments, the dCas9 enzy me can contain D 10 A and H840 A; D 10A and H840Y ; D 10A andH840N; DION and H840A; DION and H840Y; or DION and H840N substitutions. The substitutions can be conservative or non-conservative substitutions to render the Cas9 polypeptide catalytically inactive and able to bind to target nucleic acid.
[0170] In many embodiments, guide polynucleotides, e.g., gRNAs, are used with a CRISPR- Cas nuclease. In some embodiments, a guide polynucleotide, e.g., gRNA, comprises a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% identity to a sequence at a target site. In some embodiments, the target site is on the genomic DNA of a host cell. In some embodiments, a gRNA library is used with the CRISPR-Cas nuclease. A non-limiting example of a gRNA library is the CRISPR-Cas9 MinLib plasmid library (MinLibCas9 Library , Addgene #164896). sgRNAs
[0171] In some embodiments, the single guide RNAs (sgRNAs) used to inhibit Lig4, XCCR4, B ABAM2, ABRAXAS 1 , or Rap80 is designed to target the loci of the genes encoding Lig4, XCCR4, BABAM2, ABRAXAS 1, or Rap80. sgRNAs interact with a site-directed nuclease such as Cas9 and specifically bind to or hybridize to a target nucleic acid within the genome of a cell, such that the sgRNA and the site-directed nuclease co-localize to the target nucleic acid in the genome of the cell. The sgRNAs as used herein comprise a targeting sequence that has homology (or complementarity) to a target DNA sequence at the loci of a gene encoding Lig4, XCCR4, BABAM2, ABRAXAS 1, or Rap80 , and a constant region that mediates binding to Cas9 or another RNA-guided nuclease. The sgRNA can target any sequence within the gene encoding Lig4, XCCR4, BABAM2, ABRAXAS 1, or Rap80 adjacent to a PAM sequence
[0172] The targeting sequence of the sgRNAs may be, e.g., 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides in length, or, e.g., 15-25, 18-22, or 19-21 nucleotides in length, and shares homology with a targeted genomic sequence, in particular at a position adjacent to a CRISPR PAM sequence. The sgDNA targeting sequence is designed to be homologous to the target DNA, i.e., to share the same sequence with the non-bound strand of the DNA template or to be complementary to the strand of the template DNA that is bound by the sgRNA. The homology or complementarity of the targeting sequence can be perfect (i.e., sharing 100% homology’ or 100% complementarity to the target DNA sequence) or thetargeting sequence can be substantially homologous (i.e., having less than 100% homology or complementarity, e.g, with 1-4 mismatches with the target DNA sequence).
[0173] Each sgRNA also includes a constant region that interacts with or binds to the site- directed nuclease, e.g. , Cas9. In the nucleic acid constructs provided herein, the constant region of an sgRNA can be from about 70 to 250 nucleotides in length, or about 75-100 nucleotides in length, 75-85 nucleotides in length, or about 80-90 nucleotides in length, or 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86. 87. 88. 89. 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more nucleotides in length. The overall length of the sgRNA can be, e.g, from about 80-300 nucleotides in length, or about 80-150 nucleotides in length, or about 80-120 nucleotides in length, or about 90-110 nucleotides in length, or, e.g, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100. 101, 102, 103, 104, 105. 106, 107, 108, 109, or 110 nucleotides in length.
[0174] It will be appreciated that it is also possible to use two-piece gRNAs (crtracrRNAs) in the present methods, i.e., with separate crRNA and tracrRNA molecules in which the target sequence is defined by the crispr RNA (crRNA), and the trans-activating crispr RNA (tracrRNA) provides a binding scaffold for the Cas nuclease.
[0175] In some embodiments, the sgRNAs comprise one or more modified nucleotides. For example, the polynucleotide sequences of the sgRNAs may also comprise RNA analogs, derivatives, or combinations thereof. For example, the probes can be modified at the base moiety, at the sugar moiety, or at the phosphate backbone (e.g., phosphorothioates). In some embodiments, the sgRNAs comprise 3’ phosphorothioate intemucleotide linkages, 2’-O- methyl-3'-phosphoacetate modifications, 2’ -fluoro-pyrimidines, S-constrained ethyl sugar modifications, or others, at one or more nucleotides.
[0176] The sgRNAs can be obtained in any of a number of ways. For sgRNAs, primers can be synthesized in the laboratory using an oligo synthesizer, e.g., as sold by Applied Biosystems, Biolytic Lab Performance, Sierra Biosystems, or others. Alternatively, primers and probes with any desired sequence and / or modification can be readily ordered from any of a large number of suppliers, e.g.. ThermoFisher. Biolytic, IDT, Sigma- Aldritch, GeneScript, etc.RNA-guided nuclease
[0177] Any CRISPR-Cas nuclease can be used in the method, i.e., a CRISPR-Cas nuclease capable of interacting with a guide RNA and cleaving the DNA at the target site as defined by the guide RNA. In some embodiments, the nuclease is Cas9 or Cpfl. In particularembodiments, the nuclease is Cas9. The Cas9 or other nuclease used in the present methods can be from any source, so long that it is capable of binding to an sgRNA of the invention and being guided to and cleaving the specific sequence targeted by the targeting sequence of the sgRNA. In particular embodiments, Cas9 is from Streptococcus pyogenes.
[0178] In addition to the CRISPR / Cas9 platform (which is a type II CRISPR / Cas system), alternative systems exist including type I CRISPR / Cas systems, type III CRISPR / Cas systems, and type V CRISPR / Cas systems. Various CRISPR / Cas9 systems have been disclosed, including Streptococcus pyogenes Cas9 (SpCas9), Streptococcus thermophilus Cas9 (StCas9), Campylobacter jejuni Cas9 (CjCas9) and Neisseria cinerea Cas9 (NcCas9) to name a few. Alternatives to the Cas system include the Francisella novicida Cpfl (FnCpfl), Acidaminococcus sp. Cpfl (AsCpfl), and Lachnospiraceae bacterium ND2006 Cpfl (LbCpfl) systems. Any of the above CRISPR systems may be used to induce a single or double stranded break at the locus of interest to carry out the methods disclosed herein.Introducing the sgRNA and Cas protein into cells
[0179] The sgRNA and nuclease can be introduced into a cell using any suitable method, e.g., by introducing one or more polynucleotides encoding the sgRNA and the nuclease into the cell, e.g., using a vector such as a viral vector or delivered as naked DNA or RNA, such that the sgRNA and nuclease are expressed in the cell. In particular embodiments, the sgRNA and nuclease are assembled into ribonucleoproteins (RNPs) prior to delivery to the cells, and the RNPs are introduced into the cell by, e.g., electroporation. RNPs are complexes of RNA and RNA-binding proteins. In the context of the present methods, the RNPs comprise the RNA- binding nuclease (e.g. , Cas9) assembled with the guide RNA (e.g. , sgRNA), such that the RNPs are capable of binding to the target DNA (through the gRNA component of the RNP) and cleaving it (via the protein nuclease component of the RNP).Homologous Repair Templates
[0180] The CRISPR-Cas system ty pically includes a homologous repair template, or homologous donor template. The template includes a sequence that will be integrated into the genome in the place of a corresponding sequence in the genome, e.g., the sequence present between homologous regions in the template will replace a corresponding sequence present between the corresponding homologous regions in the genome. For example, in some embodiments the sequence in the template will introduce a deletion or an inactivating mutation into the genomic sequence of the genes encoding Lig4. XCCR4, BABAM2, ABRAXAS I, orRap80, thereby eliminating or reducing Lig4, XCCR4, BABAM2, ABRAXAS 1. or Rap80 expression and / or activity’ in the cell. In particular embodiments, the sequence to be introduced is flanked in the template by homology regions, e.g., sequences of from, e.g, 100, 200, 300, 400, 500 or more nucleotides comprising homology to the genomic sequence on either side of the gRNA target sequence. d. Antibodies
[0181] In particular embodiments, the inhibitor is an anti-Lig4, anti-XCCR4, anti-BABAM2, anti-ABRAXASl, or anti-Rap80 antibody or an antigen-binding fragment thereof. In some embodiments, the antibody is a blocking antibody (i.e., an antibody that binds to a target and directly interferes with the target's function, e.g., activity’ relating to ecDNA formation). In some embodiments, the antibody is a neutralizing antibody i.e., an antibody that binds to a target and negates the downstream cellular effects of the target). In particular embodiments, the antibody binds to mammalian Lig4, XCCR4, BABAM2, ABRAXAS 1, or Rap80 (e.g, human Lig4, XCCR4, BABAM2, ABRAXAS 1, or Rap80).
[0182] In some embodiments, the antibody is a monoclonal antibody. In some embodiments, the antibody is a polyclonal antibody. In some embodiments, the antibody is a chimeric antibody. In some embodiments, the antibody is a humanized antibody. In some embodiments, the antibody is a human antibody. In some embodiments, the antibody is an antigen-binding fragment, such as a F(ab’)2, Fab’, Fab, scFv, and the like. The term "antibody or antigenbinding fragment" can also encompass multi-specific and hybrid antibodies, with dual or multiple antigen or epitope specificities.
[0183] In some embodiments, an anti-Lig4, anti-XCCR4, anti-BABAM2, anti-ABRAXASl, or anti-Rap80 antibody comprises a heavy chain sequence or a portion thereof, and / or a light chain sequence or a portion thereof, of an antibody sequence disclosed herein. In some embodiments, an anti-Lig4, anti-XCCR4, anti-BABAM2, anti-ABRAXAS l, or anti-Rap80 antibody comprises one or more complementarity determining regions (CDRs) of an anti-Lig4, anti-XCCR4, anti-BABAM2, anti-ABRAXASl, or anti-Rap80 antibody. In some embodiments, an anti-Lig4, anti-XCCR4, anti-BABAM2, anti-ABRAXASl, or anti-Rap80 antibody is a nanobody, or single-domain antibody (sdAb), comprising a single monomeric variable antibody domain, e.g. a single VHH domain.
[0184] For preparing an antibody that binds to Lig4, XCCR4, BABAM2, ABRAXAS 1. or Rap80, many techniques known in the art can be used. See, e.g, Kohler & Milstein, Nature256:495-497 (1975); Kozbor et al.. Immunology Today 4: 72 (1983); Cole et al., pp. 77-96 in Monoclonal Antibodies and Cancer Therapy, Alan R. Liss, Inc. (1985); Coligan. Current Protocols in Immunology (1991); Harlow & Lane, Antibodies, A Laboratory Manual (1988); and Goding, Monoclonal Antibodies: Principles and Practice (2nd ed. 1986)). In some embodiments, antibodies are prepared by immunizing an animal or animals (such as mice, rabbits, or rats) with an antigen for the induction of an antibody response. In some embodiments, the antigen is administered in conjugation with an adjuvant (e.g., Freund's adjuvant). In some embodiments, after the initial immunization, one or more subsequent booster injections of the antigen can be administered to improve antibody production. Following immunization, antigen-specific B cells are harvested, e.g., from the spleen and / or lymphoid tissue. For generating monoclonal antibodies, the B cells are fused with myeloma cells, which are subsequently screened for antigen specificity.
[0185] The genes encoding the heavy and light chains of an antibody of interest can be cloned from a cell, e.g., the genes encoding a monoclonal antibody can be cloned from a hybridoma and used to produce a recombinant monoclonal antibody. Gene libraries encoding heavy and light chains of monoclonal antibodies can also be made from hybridoma or plasma cells. Additionally, phage or yeast display technology can be used to identify antibodies and heteromeric Fab fragments that specifically bind to selected antigens (see, e.g., McCafferty et al., Nature 348:552-554 (1990); Marks et al., Biotechnology710:779-783 (1992); Lou et al.m PEDS 23:311 (2010); and Chao et al., Nature Protocols, 1:755-768 (2006)). Alternatively, antibodies and antibody sequences may be isolated and / or identified using a yeast-based antibody presentation system, such as that disclosed in, e.g, Xu et al., Protein Eng Des Sei, 2013, 26:663-670; WO 2009 / 036379; WO 2010 / 105256; and WO 2012 / 009568. Random combinations of the heavy and light chain gene products generate a large pool of antibodies with different antigenic specificity (see, e.g., Kuby, Immunology (3rd ed. 1997)). Techniques for the production of single chain antibodies or recombinant antibodies (U.S. Patent 4,946,778, U.S. Patent No. 4,816,567) can also be adapted to produce antibodies.
[0186] Antibodies can be produced using any number of expression systems, including prokaryotic and eukaryotic expression systems. In some embodiments, the expression system is a mammalian cell, such as a hybridoma. or a CHO cell. Many such systems are widely available from commercial suppliers. In embodiments in which an antibody comprises both a VH and VL region, the VH and VL regions may be expressed using a single vector, e.g., in adi-cistronic expression unit, or be under the control of different promoters. In other embodiments, the VH and VL region may be expressed using separate vectors.
[0187] In some embodiments, an anti-Lig4, anti-XCCR4, anti-BABAM2, anti- ABRAXAS 1, or anti-Rap80 antibody comprises one or more CDR, heavy chain, and / or light chain sequences that are affinity matured. For chimeric antibodies, methods of making chimeric antibodies are known in the art. For example, chimeric antibodies can be made in which the antigen binding region (heavy chain variable region and light chain variable region) from one species, such as a mouse, is fused to the effector region (constant domain) of another species, such as a human. As another example, “class switched” chimeric antibodies can be made in which the effector region of an antibody is substituted with an effector region of a different immunoglobulin class or subclass.
[0188] In some embodiments, an anti-Lig4, anti-XCCR4, anti-BABAM2, anti- ABRAXAS 1, or anti-Rap80 antibody comprises one or more CDR, heavy chain, and / or light chain sequences that are humanized. For humanized antibodies, methods of making humanized antibodies are known in the art. See, e.g, US 8,095,890. Generally, a humanized antibody has one or more amino acid residues introduced into it from a source which is non-human. As an alternative to humanization, human antibodies can be generated. As a non-limiting example, transgenic animals (e.g., mice) can be produced that are capable, upon immunization, of producing a full repertoire of human antibodies in the absence of endogenous immunoglobulin production. For example, it has been described that the homozygous deletion of the antibody heavy-chain joining region (JH) gene in chimeric and germ-line mutant mice results in complete inhibition of endogenous antibody production. Transfer of the human germ-line immunoglobulin gene array in such germ-line mutant mice will result in the production of human antibodies upon antigen challenge. See, e.g, Jakobovits et al., Proc. Natl. Acad. Sci. USA, 90:2551 (1993); Jakobovits et al.. Nature, 362:255-258 (1993); Bruggermann et al., Year in Immun., 7:33 (1993); and U.S. Patent Nos. 5,591,669, 5.589,369, and 5.545,807.
[0189] In some embodiments, antibody fragments (such as a Fab, a Fab’, a F(ab’)2, a scFv, nanobody, or a diabody) are generated. Various techniques have been developed for the production of antibody fragments, such as proteolytic digestion of intact antibodies (see, e.g., Morimoto et al., J. Biochem. Biophys. Meth., 24: 107-117 (1992); and Brennan et al., Science, 229:81 (1985)) and the use of recombinant host cells to produce the fragments. For example, antibody fragments can be isolated from antibody phage libraries. Alternatively, Fab’-SHfragments can be directly recovered from E. coli cells and chemically coupled to form F(ab’)2 fragments (see. e.g, Carter et al., BioTechnology, 10: 163-167 (1992)). According to another approach, F(ab’)2 fragments can be isolated directly from recombinant host cell culture. Other techniques for the production of antibody fragments will be apparent to those skilled in the art.
[0190] Methods for measuring binding affinity and binding kinetics are known in the art. These methods include, but are not limited to, solid-phase binding assays (e.g, ELISA assay), immunoprecipitation, surface plasmon resonance (e.g.. Biacore™ (GE Healthcare, Piscataway, NJ)), kinetic exclusion assays (e.g, KinExA®), flow cytometry, fluorescence-activated cell sorting (FACS), BioLayer interferometry (e.g., Octet™ (ForteBio, Inc., Menlo Park, CA)), and western blot analysis. e. Peptides
[0191] In some embodiments, the inhibitor is a peptide, e.g,. a peptide that binds to and / or inhibits the activity or stability of an ecDNA biogenesis pathway protein. In some embodiments, the inhibitor is a peptide that decreases Lig4, XCCR4, BABAM2, ABRAXAS 1, or Rap80 activity . In some embodiments, the inhibitor is a peptide aptamer. Peptide aptamers are artificial proteins that are selected or engineered to bind to specific target molecules. Typically, the peptides include one or more peptide loops of variable sequence displayed by the protein scaffold. Peptide aptamer selection can be made using different systems, including the yeast two-hybrid system. Peptide aptamers can also be selected from combinatorial peptide libraries constructed by phage display and other surface display technologies such as mRNA display, ribosome display, bacterial display and yeast display. See. e.g, Reverdatto et al., 2015, Curr. Top. Med. Chem. 15: 1082-1101.
[0192] In some embodiments, the agent is an affimer. Affimers are small, highly stable proteins, typically having a molecular weight of about 12-14 kDa, that bind their target molecules with specificity and affinity similar to that of antibodies. Generally, an affimer displays two peptide loops and an N-terminal sequence that can be randomized to bind different target proteins with high affinity and specificity in a similar manner to monoclonal antibodies. Stabilization of the two peptide loops by the protein scaffold constrains the possible conformations that the peptides can take, which increases the binding affinity' and specificity' compared to libraries of free peptides. Affimers and methods of making affimers are described in the art. See, e.g.. Tiede et al.. eLife. 2017, 6:e24903. Affimers are also commercially available, e.g., from Avacta Life Sciences.f. Vectors and modified RNA
[0193] In some embodiments, polynucleotides providing ecDNA biogenesis inhibiting activity, e.g., anucleic acid inhibitor such as an siRNA or shRNA, or a polynucleotide encoding a polypeptide that inhibits ecDNA formation (i.e., by targeting a protein of an ecDNA biogenesis pathway) such as a blocking antibody fragment, are introduced into cells, e.g., dendritic cells, using an appropriate vector. Examples of delivery’ vectors that may be used with the present disclosure are viral vectors, plasmids, exosomes, liposomes, bacterial vectors, or nanoparticles. In some embodiments, any of the herein-described ecDNA inhibitors, e.g., a nucleic acid inhibitor or a polynucleotide encoding a polypeptide inhibitor, are introduced into cells, e.g., muscle cells, using vectors such as viral vectors. Suitable viral vectors include but not limited to adeno-associated viruses (AAVs), adenoviruses, and lenti viruses. In some embodiments, a ecDNA inhibitor, e.g., a nucleic acid inhibitor or a polynucleotide encoding a polypeptide inhibitor, is provided in the form of an expression cassette, typically recombinantly produced, having a promoter operably linked to the polynucleotide sequence encoding the inhibitor. In some cases, the promoter is a universal promoter that directs gene expression in all or most tissue types. In other embodiments, a promoter that is specific for or whose performance is optimized for a given cell type can be used according to the cell type that is to be used.VI. Pharmaceutical Compositions
[0194] In some embodiments, the herein-described ecDNA inhibitors are present within a pharmaceutical composition or formulation. The pharmaceutical compositions of the ecDNA inhibitors of the present invention may comprise a pharmaceutically acceptable carrier. In certain aspects, pharmaceutically acceptable carriers are determined in part by the particular composition being administered, as well as by the particular method used to administer the composition. Accordingly, there is a wide variety of suitable formulations of pharmaceutical compositions of the present invention (see, e.g., REMINGTON'S PHARMACEUTICAL SCIENCES, 18TH ED., Mack Publishing Co., Easton, PA (1990)).
[0195] As used herein, ‘'pharmaceutically acceptable carrier” comprises any of standard pharmaceutically accepted carriers known to those of ordinary’ skill in the art in formulating pharmaceutical compositions. Thus, the compounds, by themselves, such as being present as pharmaceutically acceptable salts, or as conjugates, may be prepared as formulations in pharmaceutically acceptable diluents; for example, saline, phosphate buffer saline (PBS),aqueous ethanol, or solutions of glucose, mannitol, dextran, propylene glycol, oils (e.g., vegetable oils, animal oils, synthetic oils, etc.), microcrystalline cellulose, carboxymethyl cellulose, hydroxylpropyl methyl cellulose, magnesium stearate, calcium phosphate, gelatin, polysorbate 80 or the like, or as solid formulations in appropriate excipients.
[0196] The pharmaceutical compositions will often further comprise one or more buffers (e.g., neutral buffered saline or phosphate buffered saline), carbohydrates (e.g, glucose, mannose, sucrose or dextrans), mannitol, proteins, polypeptides or amino acids such as glycine, antioxidants (e.g, ascorbic acid, sodium metabisulfite, butylated hydroxy toluene, butylated hydroxyanisole, etc.), bacteriostats, chelating agents such as EDTA or glutathione, solutes that render the formulation isotonic, hypotonic or weakly hypertonic with the blood of a recipient, suspending agents, thickening agents, preservatives, flavoring agents, sweetening agents, and coloring compounds as appropriate.
[0197] The pharmaceutical compositions of the invention are administered in a manner compatible with the dosage formulation, and in such amount as will be therapeutically effective. The quantity to be administered depends on a variety of factors including, e.g., the age, body weight, physical activity, hereditary characteristics, general health, sex and diet of the individual, the condition or disease to be treated, the mode and time of administration, rate of excretion, drug combination, the stage or severity of the condition or disease, etc. In certain embodiments, the size of the dose may also be determined by the existence, nature, and extent of any adverse side effects that accompany the administration of a therapeutic agent(s) in a particular individual.
[0198] In certain embodiments, the dose of the compound may take the form of solid, semisolid, lyophilized powder, or liquid dosage forms, such as, for example, tablets, pills, pellets, capsules, powders, solutions, suspensions, emulsions, suppositories, retention enemas, creams, ointments, lotions, gels, aerosols, foams, or the like, preferably in unit dosage forms suitable for simple administration of precise dosages.
[0199] As used herein, the term “unit dosage form’7refers to physically discrete units suitable as unitary dosages for humans and other mammals, each unit containing a predetermined quantity of a therapeutic agent calculated to produce the desired onset, tolerability, and / or therapeutic effects, in association with a suitable pharmaceutical excipient (e.g., an ampoule). In addition, more concentrated dosage forms may be prepared, from which the more dilute unit dosage forms may then be produced. The more concentrated dosage forms thus will containsubstantially more than, e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more times the amount of the therapeutic compound. Kits as described herein can comprise unit dosage forms of ecDNA inhibitors as described herein.
[0200] Methods for preparing such dosage forms are known to those skilled in the art (see, e.g., REMINGTON’S PHARMACEUTICAL SCIENCES, supra). The dosage forms typically include a conventional pharmaceutical carrier or excipient and may additionally include other medicinal agents, carriers, adjuvants, diluents, tissue permeation enhancers, solubilizers, and the like. Appropriate excipients can be tailored to the particular dosage form and route of administration by methods well known in the art (see, e.g., REMINGTON’S PHARMACEUTICAL SCIENCES, supra).
[0201] Examples of suitable excipients include, but are not limited to, lactose, dextrose, sucrose, sorbitol, mannitol, starches, gum acacia, calcium phosphate, alginates, tragacanth, gelatin, calcium silicate, microcrystalline cellulose, polyvinylpyrrolidone, cellulose, water, saline, syrup, methylcellulose, ethylcellulose, hydroxypropylmethylcellulose, and polyacrylic acids such as Carbopols, e.g., Carbopol 941, Carbopol 980, Carbopol 981, etc. The dosage forms can additionally include lubricating agents such as talc, magnesium stearate, and mineral oil; wetting agents; emulsifying agents; suspending agents; preserving agents such as methyl-, ethyl-, and propyl-hydroxy-benzoates (i.e., the parabens); pH adjusting agents such as inorganic and organic acids and bases; sweetening agents; and flavoring agents. The dosage forms may also comprise biodegradable polymer beads, dextran, and cyclodextrin inclusion complexes.
[0202] For oral administration, the therapeutically effective dose can be in the form of tablets, capsules, emulsions, suspensions, solutions, syrups, sprays, lozenges, powders, and sustained-release formulations. Suitable excipients for oral administration include pharmaceutical grades of mannitol, lactose, starch, magnesium stearate, sodium saccharine, talcum, cellulose, glucose, gelatin, sucrose, magnesium carbonate, and the like.
[0203] The therapeutically effective dose can also be provided in a lyophilized form. Such dosage forms may include a buffer, e.g, bicarbonate, for reconstitution prior to administration, or the buffer may be included in the lyophilized dosage form for reconstitution with, e.g., water. The lyophilized dosage form may further comprise a suitable vasoconstrictor, e.g., epinephrine. The lyophilized dosage form can be provided in a syringe, optionally packaged in combination with the buffer for reconstitution, such that the reconstituted dosage form can be immediately administered to an individual.VII. Treatment
[0204] The present methods and compositions can be used to treat diseases, for example, cancer, infectious disease (i.e., those that are caused by bacterial or viral infections), and other immune-related diseases, i.e.. diseases for which an enhanced immune reaction can be beneficial in a subject need thereof. Tn various embodiments, the subject (also referred to herein as “patient”) for these methods cay be an adult of any age, a child, or an adolescent. The subject may be male or female. In particular embodiments, the subject is a human. In some embodiments, the cancer is a cancer of an immune-privileged organ, which can be referred to as an organ that is less subject to immune response compared to other areas of the body, e.g. the central nervous system (CNS), brain, eyes, and testes. In some embodiments, the cancer is a cancer of the CNS, brain, eye, or testis. In some embodiments, the cancer is a glioblastoma. In some embodiments, the cancer is non-small cell lung carcinoma (NSCLC). In some embodiments, the cancer is a gastric cancer.
[0205] Thus, in one aspect, provided herein is a method treating a disease in a subject in need thereof (e.g., a cancer or infectious disease whose etiology lies in intracellular circular DNA formation), the method comprising administering to the subject a therapeutic agent as described herein (also referred to herein as an “ecDNA inhibitor”). In embodiments, methods of treatment as described herein can comprise administering a therapeutic agent (or pharmaceutical composition comprising a therapeutic agent as described herein) to a subject in need thereof. In embodiments, methods of treatment as described herein can comprise administering an effective amount of a therapeutic agent (or pharmaceutical composition comprising a therapeutic agent as described herein) to a subject in need thereof (i.e., a subject having or suspected of having a cancer or infectious disease), wherein the effective amount is an amount or concentration of a therapeutic agent sufficient to reduce the effects of one or more symptoms of the disease. In some embodiments, the therapeutic agent is an ecDNA inhibitor. In some embodiments, the ecDNA inhibitor is administered at a dose of 0.01 nM, 0.05 nM, 0.1 nM, 0.2 nM, 0.3 nM, 0.4 nM, 0.5 nM, 0.6 nM, 0.7 nM, 0.8 nM, 0.9 nM, 1 nM, 5 nM, 10 nM, 20 nM, 30 nM, 40 nM, 50 nM, 60 nM, 70 nM, 80 nM, 90 nM, 100 nM, 150 nM, 200 nM, 250 nM, 300 nM, 350 nM, 400 nM, 450 nM, 500 nM, 550 nM, 600 nM, 650 nM, 700 nM, 750 nM, 800 nM, 850 nM, 900 nM, 950 nM, 1000 nM, 1050 nM, 1100 nM, 1150 nM, 1200 nM. 1250 nM, 1300 nM, 1350 nM, 1400 nM. 1450 nM, 1500 nM, 0.01-1 nM, 0.1- 10 nM, 1-100 nM, 20-250 nM, 50-500 nM, 100-500 nM, 100-1000 nM, 200-1000 nM, 500- 1000 nM, 500-1500 nM, or 750-1500 nM.
[0206] In some embodiments, a therapeutic agent as described herein can be administered to the subject in conjunction with another treatment such as immunotherapy and / or an anti-cancer or anti-infection agent.
[0207] In some embodiments, the subject has an infection and the therapeutic agent is administered to the subject in conjunction with another appropriate therapy such as an antiviral or antibiotic (anti-bacterial) compound.
[0208] In some embodiments, the subject has cancer, and the therapeutic agent is administered to the subject as a combination therapy in conjunction with another anti-cancer therapy such as chemotherapy, radiation treatment, and / or surgical treatment. In some embodiments, the subject is administered one or more of a tyrosine kinase inhibitor, costimulatory mAb, epigenetic modulator, chemotherapeutic agent, radiation therapeutic, vaccine, adoptive T-cell therapeutic, or oncolytic virus.
[0209] In some embodiments, the subject receives surgical treatment for the cancer in addition to the therapeutic agent. For example, the patient may receive surgical resection (removal of the tumor with surgery). Small tumors may also be treated with other types of treatment such as ablation or radiation. Ablation is treatment that destroys tumors without removing them. These techniques can be used in patients with a few small tumors and when surgery is not a good option. They are less likely to cure the cancer than surgery, but they can still be very helpful for some people. Ablation is best used for tumors no larger than 3 cm across. For slightly larger tumors (1 to 2 inches, or 3 to 5 cm across), it may be used along with embolization. Because ablation often destroys some of the normal tissue around the tumor, it might not be a good choice for treating tumors near major blood vessels, the diaphragm, or major bile ducts. In some embodiments, the ablation is radiofrequency ablation (RFA). In some embodiments, the ablation is microwave ablation (MW A). In some embodiments, the ablation is cryoablation (cryotherapy). In some embodiments, the ablation is ethanol (alcohol) ablation, e.g., percutaneous ethanol injection (PEI).
[0210] In some embodiments, a patient with cancer is treated using radiation therapy in conjunction with a therapeutic agent as described herein. Radiation therapy uses high-energy rays, or particles to destroy cancer cells. Radiation can be helpful, e.g., in treating cancer that cannot be removed by surgery, cancer that cannot be treated with ablation or did not respond well to such treatment; cancer that has spread to areas such as the brain or bones; patients experiencing severe pain due to large cancers; and patients having a tumor thrombus.
[0211] In some embodiments, a patient with cancer is treated using drug therapy, e.g, targeted drug therapy, immunotherapy, or chemotherapy in combination with a therapeutic agent as described herein. Targeted drugs work differently from standard chemotherapy drugs and include, e.g., kinase inhibitors; Sorafenib (Nexavar), lenvatinib (Lenvima), Regorafenib (Stivarga), and cabozantinib (Cabometyx). Immunotherapy can comprise the administration of monoclonal antibodies. Monoclonal antibodies are designed to attach to a specific target. The monoclonal antibodies used to treat liver cancer affect a tumor’s ability to form new blood vessels, also known as angiogenesis. These therapeutics are often referred to angiogenesis inhibitors and include: Bevacizumab (Avastin), which can be used in conjunction with the immunotherapy drug atezolizumab (Tecentriq); Ramucirumab (Cyramza).
[0212] Common chemotherapy drugs for treating cancer include, for example: Gemcitabine (Gemzar); Oxaliplatin (Eloxatin); Cisplatin; Doxorubicin (pegylated liposomal doxorubicin); 5-fluorouracil (5-FU); Capecitabine (Xeloda); Mitoxantrone (Novantrone), or combinations thereof. Chemotherapy can be regional when drugs are inserted into an artery that leads to the part of the body with the tumor, thereby focusing the chemotherapy on the cancer cells in that area of the body and reducing side effects by limiting the amount of drug reaching the rest of the body. For example, hepatic artery infusion (HAI), or chemo given directly into the hepatic artery, is an example of a regional chemotherapy that can be used for liver cancer.
[0213] In some embodiments, the subject receives immunotherapy in conjunction with the administration of a therapeutic agent as described herein for the treatment of cancer, infection, or other immune-related condition. For example, in some embodiments the immunotherapy comprises administering an adoptive T-cell therapeutic (e.g. , CAR-T cell, CAR-NK cell, C AR- Macrophage), co-stimulatory mAb, epigenetic modulator, vaccine against an infectious agent, tumor vaccine, oncolytic virus vaccine, TLR3 / 7 / 8 / 9 agonist, anti-CD47. or IL-2 receptor agonist to the subject. In some embodiments, the subject receives an immune checkpoint therapy (ICT). An important part of the immune system is its ability to keep itself from attacking normal cells in the body. To do this, it uses “checkpoints” - proteins on immune cells that need to be switched on or off to start an immune response. Cancer cells sometimes use these checkpoints to avoid being attacked by the immune system. Newer drugs that target these checkpoints hold a lot of promise as cancer (or infectious disease, or other immune-related condition) treatments and include, for example: PD-1 and PD-L1 inhibitors; Atezolizumab (Tecentriq), which can be used in conjunction with the targeted drug bevacizumab (Avastin); Pembrolizumab (Keytruda) and nivolumab (Opdivo), alone or in combination with ipilimumab(described below) may also be an option. In some embodiments, the immune checkpoint blockade binding agent is an anti-CTLA4, anti-PDl, anti-PD-Ll. anti-LAG-3, anti-TIM-3, anti-TIGIT, anti-CD47 or anti-VISTA antibody.VIII. Kits and Systems
[0214] Provided herein are kits that comprise reporter construct polynucleotides as described herein. For example, in embodiments, kits comprise a Version 1 reporter construct and / or a Version 2 reporter construct, or variations thereof (e.g, the VI or V2 reporters with different promoters or reporters than what is shown in FIG. 1). Polynucleotides as described herein can be present in a kit in a lyophilized form or other non-aqueous form for ease of transport. The kits may also comprise any additional reagents for carrying out any of the methods described herein. Exemplary reagents, without limitations, include any nucleic acid, DNA construct, vector, polypeptide, host cell, cell line, and / or cell population of the present disclosure.
[0215] In some embodiments, a kit may further include other components, for example, commercially available reagents and tools that are familiar to one of ordinary skill in the art. Such components may be provided individually or in combinations, and may be provided in any suitable container such as a vial, a bottle, or a tube. Examples of such components include, but are not limited to. (i) one or more additional reagents, such as one or more dilution buffers; one or more reconstitution solutions; one or more wash buffers; one or more storage buffers, one or more control reagents and the like, (ii) one or more control expression vectors; (iii) one or more reagents for in vitro production and / or maintenance of the of the constructs, cells, deli ven’ systems etc. provided herein; and the like. Components (e.g, reagents) may also be provided in a form that is usable in a particular assay, or in a form that requires addition of one or more other components before use (e.g. in concentrate or lyophilized form). Suitable buffers include, but are not limited to, phosphate buffered saline, sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer. HEPES buffer, and combinations thereof.
[0216] In some embodiments, a kit may further include delivery systems for contacting the reporter constructs of the present disclosure to cells. Non-limiting examples of delivery systems (viral and non-viral) are discussed above.
[0217] In some embodiments, for example, a kit may comprise, consist of, or consist essentially of one or more of the following: (i) a Version 1 reporter construct as provided herein, or a variation thereof that allows for expression of the reporter upon circularization ofthe reporter; (ii) a Version 2 reporter construct as provided herein, or a variation thereof that allows for expression of the reporter upon circularization of the reporter; (iii) an ecDNA reporting system as provided herein; (iii) delivery systems comprising a Version 1 reporter construct, a Version 2 construct (or variations thereof), and / or an ecDNA reporting system as provided herein; and / or (iv) cells comprising a Version 1 reporter construct (or a variation thereol) that allows for expression of the reporter upon circularization of the reporter, aversion 2 reporter construct (or a variation thereof that allows for expression of the reporter upon circularization of the reporter), an ecDNA reporting system, and / or a delivery system comprising a Version 1 reporter construct (or a variation thereof that allows for expression of the reporter upon circularization of the reporter), a Version 2 reporter construct (or a variation thereof that allows for expression of the reporter upon circularization of the reporter), and / or an ecDNA reporting system as provided herein.
[0218] In some embodiments, the kit can be used to detect the formation of circular DNA, e.g., ecDNA, in a cell. In some embodiments, the kit is used for identifying regulators, e.g., promoters or suppressors, of circular DNA (e.g., ecDNA) formation. In some embodiments, the kit is used for identifying agents (e.g.. small molecules less than 2500 daltons, other nucleic acid or polypeptide agents) that can block or otherwise suppress circular DNA (e.g., ecDNA) formation. Such agents are useful, for example, for use in the treatment of diseases, such as cancer. In some embodiments, a kit may further comprise instructions on how to use the kit to perform any of the methods described herein. The instructions for practicing the disclosed methods are generally recorded on a suitable recording medium. For example, the instructions may be printed on a substrate, such as paper or plastic, etc. As such, the instructions may be present in the kits as a package insert, in the labeling of the container of the kit or components thereof (i.e., associated with the packaging or subpackaging), etc. In other embodiments, the instructions are present as an electronic storage data file present on a suitable computer readable storage medium, e.g. CD-ROM, diskette, flash drive, etc. In yet other embodiments, the actual instructions are not present in the kit, but means for obtaining the instructions from a remote source, e.g. via the internet, are provided. An example of this embodiment is a kit that includes a QR code or a web address where the instructions can be viewed and / or from which the instructions can be downloaded. As with the instructions, the means for obtaining the instructions is recorded on a suitable substrate.
[0219] Also provided herein are systems for studying circular DNA, e.g., ecDNA, formation. For example, a system may include a kit disclosed herein and instruments for performing anymethod disclosed herein. In some cases, the system includes a chamber (i.e., one or more bioreactors) for maintaining the physiological conditions of cells that comprise a reporter construct polynucleotide disclosed herein. In some cases, the system includes a microscope, a camera, and / or a display for observing the samples and / or a means for recording images of the samples. In some cases, the system includes a component capable of detecting and / or measuring fluorescence signals. In some cases, the system may comprise an instrument that is capable of sorting cells based on the presence, absence, or magnitude of fluorescence signals emitted by the cells..EXAMPLES
[0220] The following Examples are provided by way of illustration and not by way of limitation.Overview
[0221] The following Examples relate to the engineering of a more versatile and robust system. The Examples discuss designing reporter constructs for use in mammalian cells and using such constructs to identify regulators of ecDNA formation.Example 1 - Materials and Methods1.1 - Cell Culture of HEK293T, HeLa, PC9, and HCT116 Cells
[0222] HEK293T, HeLa, and HCT116 cells were cultured under the following conditions: Gibco DMEM (High Glucose, GlutaMAX Supplement, Pyruvate; ThermoFisher Scientific, Cat # 10569044), supplemented with 10% FBS (Cytvia, SH30396.03) and 1% penicillinstreptomycin (ThermoFisher Scientific, Cat # 15140122). PC9 cells were cultured in RPMI 1640 Medium (Sterile, pH 7.0 to 7.6, with L-glutamine and sodium bicarbonate, suitable for cell culture; Sigma-Aldrich, Cat # R8758-500ML), supplemented similarly with 10% FBS (Cytvia, SH30396.03) and 1% penicillin-streptomycin (Thermo Fisher Scientific, Cat # 15140122). HEK293T biosensor cells were a generous gift from Dr. Alun Luo1and were reengineered with lentivirus containing vector lentiCas9-Blast (addgene # 52962) to express multiple copies of Cas9 protein. These cells were cultured in similar conditions as normal HEK293T cells. The incubation conditions were maintained at 37 °C and 5% CO2.1.2 - Design, Cloning, and Transfection of Version 1 & 2 EGFP ReporterConstructs
[0223] To construct the Version 1 biosensor, the pCAG-GFP plasmid (Addgene #11150) was first digested with EcoRI (NEB. Cat # R3101S) and Hindlll (NEB, Cat # R3104S),followed by re-ligation. GFP-polyA was then inserted into theNotl (NEB, Cat # R3189S) site. For Version 2. DsRed-T2A-PuroR fragment was amplified from a U6-sgRNA-DsRed-2A- PuroR vector (gifted by the Kris Wood lab) using CloneAMP HiFi PCR Premix (Takara, Cat # 639298), and subsequently integrated into Version 1 biosensor pre-digested with Spel (NEB, Cat # R3133S). Both plasmids were validated via Sanger sequencing. For linearization, 10 pg of Version 1 or 2 biosensor (50 pl volume) was treated with 1 pl EcoRV (NEB, Cat # R3195T) and 5 pl CutSmart buffer, incubated at 37 °C for 4 hours. The products were then purified by using 0.8% agarose gel. Bands of expected size were extracted using E.Z.N.A. Gel Extraction Kit (Omega Bio-Tek, Cat # D2500). HEK293T, HeLa, PC9, and HCT116 cells (IxlO6cells each) were transfected with 500 ng (HEK293T) or 2000 ng (HeLa, HCT116, and PC9) of linearized Version 1 & 2 biosensor using Lipofectamine 3000 Transfection Reagent (Thermo Scientific, Cat # L3000008). The culture medium was refreshed 4 hours post-transfection. Images were acquired at 24 hours post-transfection using a Zeiss AxioObserver microscope. Image assembly and processing were conducted using Adobe Photoshop and Illustrator.1.3 - Homology End Design and Amplification
[0224] To prepare the Version 2 biosensor with homologous ends (i.e., the 5’-end and the 3’-end of the biosensor polynucleotide have the same sequence), 1 ng of the Version 2 biosensor was amplified using CloneAMP HiFi PCR Premix (Takara, Cat # 639298). The amplification conditions were as follows: initial denaturation at 95°C for 3 min; 30 cycles of 95°C for 15 s, 64°C for 15 s. and 72°C for 2 min; final extension at 72°C for 5 min (Primers listed in Table 2). The PCR products underwent electrophoresis and subsequent gel purification as described above.1.4 - PCR Amplification of Junctions
[0225] The PCR amplification of junctions was carried out as previously described1. Cells transfected with the Version 1 biosensor were harvested 24 hours post-transfection. Total DNA was extracted using the Quick-DNA Miniprep Kit (Zymo Research, Cat # D3025). For each sample, 100 ng total DNA was prepared for exonuclease treatment. This involved mixing with 1 pl of 10X Plasmid-Safe DNase buffer, 1 pl Plasmid-Safe DNase, and 1 pl of 25 mM ATP. The mixture was incubated in a thermocycler at 37 °C for 16 hours, followed by a 30-minute incubation at 70 °C. For control samples (without exonuclease treatment), 1 pl of water replaced the Plasmid-Safe DNase. The amplification of the treated DNA was performed using either CloneAMP HiFi PCR Premix (Takara Bio, Cat # 639398) or GoTaq Green Master Mix (Promega, Cat # M7123). The amplification conditions were: an initial denaturation at 95 °Cfor 3 minutes; followed by 30 cycles of 95 °C for 15 seconds, 58 °C for 15 seconds, and 72 °C for 30 seconds; and a final extension at 72 °C for 3 minutes. Primer sequences used are detailed in Table 2.1.5 - Lentivirus Production and Titering
[0226] Transfection reagents were prepared in Opti-MEM reduced serum medium (Gibco) with appropriate scaling for culture surface area, according to manufacturer instructions. For the lentiviral preparation of the genome-wide sgRNA lentiviral library, 80% confluent HEK293T cells were transfected in a T-225 Flask with 42.4 pg of the MinLib plasmid library (MinLibCas9 Library was a gift from Dr. Mathew Garnett, Addgene #164896), 32.4 pg of psPAX2, and 21.2 pg of pMD2.G, using 374 pl of Lipofectamine 2000 supplemented with 17.54 pl of PLUS Reagent. After 6 hours following lipofection, the transfection media was replaced with fresh pre-warmed harvest media (HEK293T media supplemented with 25% FBS). After 48 hours, the viral supernatant was collected, filtered using a 0.45 pm PES filter, and either used directly or stored at -80 °C for future use. Large batches of lentivirus containing vectors encoding biosensor CRISPR-C sgRNAs that direct cutting to the left or right flank of the eGFP biosensor (eGFP ORF and CAG promotor) were prepared using the Gibco LV-MAX Lentiviral Production system (Thermo Fisher Scientific Cat# A35684) as per manufacturer instructions. Equal portions of left cut and right cut virus were pooled together to make CRISPR-C biosensor cutting virus. Titers of the lentiviral sgRNA library were determined by flow cytometry’. Aliquots of 3xl06HEK293T biosensor cells were seeded with varying volumes of library lentivirus in 15 cm dishes containing a final concentration of 8 pg / ml polybrene (EMD Millipore, Cat# TR-1003-G). Transduction media was replaced with fresh media 24 hours later. Cells were harvested 96 hours post-transduction and the levels of BFP in each sample were measured to determine the percent transduction.1.6 - Lentiviral SgRNA Fluorescent Activated Cell Sorting (FACS) Screening
[0227] HEK293T biosensor cells were seeded into 15 cm tissue culture dishes with 8 pg / ml polybrene and transduced with the titered lentiviral sgRNA library at a low multiplicity of infection ~0.4 (statistically ensuring that most cells harbor no more than 1 sgRNA) to achieve greater than l,000x coverage of the sgRNA library, in biologic triplicate. Twenty -four hours post-transduction, media was replaced with fresh media and cells were incubated for 72 hours. Cells were then selected for BFP expression by FACS. Sorted cells containing sgRNA were allowed to recover for 24 hours in fresh media with 30% FBS. Cells were then infected with infected with optimal CRISPR-C biosensor cutting vims with 8 pg / ml polybrene for 24 hoursto allow for excision of the DNA biosensor. Virus media was replaced with fresh media and cells were passaged and maintained above 1000X coverage for 72 hours post-transduction. Replicates were harvested and cells were then sorted by FACS based on GFP expression, collecting the top and bottom 10%; the cells were collected in FBS-coated tubes, maintaining approximately 1000x coverage per high and low population. An ungated control sample was also collected for each replicate.1.7 - Screen Processing and Data Analysis
[0228] Immediately following FACS collection, replicate samples were individually subjected to genomic DNA extraction using the Quick-DNA Miniprep Kit (ZymoResearch, Mfr.# D3025). sgRNA libraries were recovered from gDNA via PCR amplification using NEBNext Ultra II Q5 Master Mix (NEB, Cat# M0544L) according to manufacturer instructions and using custom primers, as previously described (Klann et al., Nat. Biotech.. 35(6) 2017; primer sequences of which incorporated by reference as if fully set forth herein):
[0229] Amplified libraries were purified using SPRIselect beads (Beckman Coulter, Cat# B23317) employing right-sided selection of 0.8x then to 1.2x the original volume. Each DNA sample was quantified using the Qubit dsDNA Broad Range Assay Kit (ThermoFisher Scientific, Cat# Q32850) and quality checked using an Agilent 4150 TapeStation System. DNA was analyzed using D5000 ScreenTape (Agilent Cat# 5067-5588). Samples were pooled and sequenced on a NextSeq 500 (Illumina) with 20-bp single-end sequencing using custom read and index primers.
[0230] Raw sequencing read counts were processed, and analysis of enrichment and depletion metrics comparing the eGFP+ and eGFP- populations was performed using the MAGeCK software analysis pipeline under default settings. Volcano plot and snake plot were generated using R studio.1.8 - Gene-Network Construction (Co-Essentiality and Gene Ontology Analysis)
[0231] Project Achilles4gene essentiality data was first obtained from the Broad Institute’s DepMap portal (20ql release). The locus-adjusted gene coessentiality was then determined with FIREWORKS5 (https: / / github.com / mendillolab / fireworks / ) as described. Essentiality scores for each gene were adjusted by subtracting half of the median value of its 40 nearest neighbor genes’ (20 upstream, 20 downstream) while excluding those located within duplicate gene clusters, followed by pairwise Pearson correlations conducted at a genome scale. Protein- Protein Interactions (PPIs) were added to enhance the gene network. PPIs of genes were extracted from Metascape6 (3.5) which utilizes physical PPIs captured in BioGrid7and STRINGS database (version 12.0) as the main datasource. PPIs from top 60 drivers for ecDNA biogenesis with coessentiality correlations score above 0.2 were integrated into gene pairs. The gene network was constructed in Cytoscape 3.10.19. For suppressors with beta scores above 0.375 and p-value less than 0.05, the network was constructed based on their co-essentiality with PPI information added.
[0232] Gene ontology analysis was conducted using enrichGO in the clusterProfiler R package (https: / / rdrr.io / bioc / clusterProfiler / ). Top 30 ecDNA drivers with beta scores below - 0.350 and top 80 ecDNA suppressors with beta scores above 0.375 were analyzed for enrichment of biological processes (BP). Bar plots for GO analysis were generated using RStudio.1.9 - Homology Modeling
[0233] Random genome fragments were generated by the 'randomBed' function from BEDToolslO (v2.26.0) tool for the hg38 genome. Different size intervals were obtained by setting the parameter ‘-1’ while generating 1 million intervals each time. To eliminate redundancy, only unique intervals were retained. Both strand sequences at the tw o ends of each interval were extracted using 'gelFaslaFromBed' from BEDTools. When calculating the percentage of identical nucleotides, the sequences containing ambiguous ’N’ bases were excluded. The same process w as iterated 201 times and the mean value of these simulations was calculated and visualized using ggplot2 package in RStudio. Error bars w ere added using the ■stat suniniary ’ function and extended to 5 times the standard deviation (SD) from the mean to provide visualization of the variability within the random simulations.1.10 - Gene Mutation by CRISPR-Cas9 in HEK293T, PC9, and HeLa Cells
[0234] The sgRNAs were derived from the MinLib CRISPR guide RNA library. To ensure efficiency, an additional guanine was appended to the sgRNAs that do not start with a guanine. Each sgRNA was cloned into pU6-(BbsI)-CBh-Cas9-T2A-BFP plasmid (Addgene, Cat # 64323) and verified by Sanger sequencing. To generate mutant gene mutations, cells were seeded at 300,000 cells per well in 6-well plates with complete media and allowed to grow overnight. The next day, 2 ml of culture media was replaced prior to transfection. Transfection was carried out by delivering 3 pg of plasmid to cells using Lipofectamine™ 3000 Transfection Reagent (ThermoFisher, Cat # L3000001) according to manufacturer’s recommended protocol. Cells were incubated with transfection mixture for 24 hours and then incubated with fresh media for an additional 48 hours. Then cells were sorted for BFP signal using a Beckman Coulter Astrios EQ High-Speed Cell Sorter. The collected cells were plated into 6-cm dishes and allowed to recover for 48 hours post-sorting. Next, the cells were dissociated and diluted to 30 cells / ml. One hundred pl of the cell suspension was distributed into 96-well plates per well. Single clones were allowed to grow into stable colonies over a period of approximately 10-14 days for HEK293T cells and 30-40 days for HeLa and PC9 cells. Finally, the colonies were validated for mutation by Sanger sequencing the site of sgRNA directed mutation or protein depletion by western blot. sgRNA sequences are listed in Table 2.
[0235] Wild-type and mutant HEK293T cell clones were plated at 300,000 cells in 6-well with 2 ml of media (DMEM, 10% FBS) and incubated overnight. Ten pM Mirin was added with media. The next day, CRISPR-C left and right sgRNA lentivirus was added with 8 pg / ml polybrene to the cells and allowed to grow for 72 hours. Cells were observed for eGFP with Zeiss Axio Observer microscope.1.11 - Cloning and expressing of Lig4 wild-type and mutants for rescue
[0236] The cloning of the LIG4 wild-type and mutant alleles were amplified by using CloneAMP HiFi PCR Premix (Takara, Cat # 639298) from the pDR119 plasmid (Addgene, Cat # 13332). The specific primers used for this purpose are listed in Table 2. Following amplification, the PCR products were assembled into a pcDNA3.1 plasmid (Thermo Fisher Scientific, Cat # V79020), which was pre-digested with EcoRI (NEB, Cat # R3101S) and BamHI (NEB, Cat # R3136S). This assembly was facilitated by the CloneExpress Ultra One Step Cloning Kit (Cellagen Technology, Cat # Cl 15-02). For plasmid propagation, Mix & Go Competent Cells-Zymo strain 10B (Zymo Research, Cat # T3020), a chemically competent E. coli strain, was utilized. The propagated plasmids were then extracted using the PureYieldPlasmid Miniprep System (Promega, Cat # Al 222). To confirm the accuracy of the constructed vectors. Sanger sequencing was employed for verification.
[0237] HEK293T biosensor cells (WT and / . / GV- / -) were seeded at 300,000 cells / well in 6- well plates with 2 ml of culture media and incubated overnight. The following day, plasmids were transfected into cells using Lipofectamine 3000 Transfection Reagent (Thermo Scientific, Cat # L3000008) as per manufacturer's instruction. The culture medium was refreshed 16 hours post-transfection. Images were acquired at 48 hours post-transfection using a Zeiss AxioObserver microscope. For western blotting, cells were harvested at 48 hours after transfection.1.12 - Quantitative Droplet Digital PCR (ddPCR) of CRISPR-C Junction Sites
[0238] Cells were washed with 2 ml PBS, followed by adding 100 pl 0.25% Trypsin + EDTA and incubated in 37 °C for 5-10 minutes. Trypsin was quenched by adding 1 ml of culture medium (10% FBS). Cells were pelleted at 300 g for 5 minutes. The supernatant was aspirated, and the cell pellets were stored at -80°C.
[0239] Genomic DNA was isolated using DNeasy Blood & Tissue Kit (QIAGEN, Cat. 69594) according to the manufacturer’s instructions. Briefly, the cell pellets were resuspended in 200 pl PBS, and then mixed with 200 pl lysis buffer with 20 pl of Proteinase K. The protein was digested at 56 °C for 10 min. The cell lysate was transferred to the binding columns and washed 2 times. gDNA was eluted with 80 pl RNase free water. Amplicons for the circular DNA junctions and glyceraldehyde 3-phosphate dehydrogenase (GAPDH) were designed using the IDT PrimerQuest™ Tool (https: / / www.idtdna.com / pages / tools / primerquest). Dualquenched probes (IDT) were used. Probes for ecDNA junction amplicons were labeled with FAM, and probe for GAPDH amplicon was labeled with HEX to facilitate the multiplexing. The sequences of probes and primers are listed in Table 2.
[0240] ddPCR was performed on samples using the QX200TM ddPCR system (Bio-Rad Laboratories). The amplification reactions were set up according to the manufacturer's specifications (Bio-Rad Laboratories). Briefly, approximately 50 ng of gDNA was used in a 20 pl reaction with 10 pl ddPCR Supermix for probes (no dUTP) (Bio-Rad Laboratories), 900 nM for primers (225 nM for each primer), 125 nM of FAM probe, and 125 nM of HEX probe. Then, the droplets were created using droplet-generating oil for probes, DG8 cartridges, DG8 gaskets and the QX200 Droplet generator (Bio-Rad Laboratories). The 96-well plate with droplets was sealed with a PX1 PCR plate sealer (Bio-Rad Laboratories). PCR of droplets wasperformed using the following setting: 95 °C for 10 minutes, 40 cycles (95 °C for 30 seconds, 55.8 °C for 60 seconds), 98 °C for 10 minutes. After PCR amplification, the plate was analyzed with a QX200 Droplet Reader (Bio-Rad Laboratories). The positive and negative droplets were quantified by QuantaSoft Software. The ecDNA frequency was calculated as: ecDNA concentration ecDNA frequency (%) = — — — — - -GAPDH concentration1.13 - Fly strains, housing, and husbandry conditions
[0241] All flies were grown on standard agar-com medium and maintained at 25°C. Flies carrying the Lig4
[0057] mutation were generously gifted by Dr. Jefferey Sekelsky. Female flies aged 3-6 days were selected for experiments. Ovaries were carefully dissected and washed in lx Phosphate-buffered saline (PBS) prior to experiments.1.14 - Divergent PCR of chorion-ecDNA
[0242] DNA from Drosophila ovary and carcass was purified using the Quick-DNA Microprep Kit (ZymoResearch, Mfr.# D3020). One hundred ng total DNA (in 10 pl volume) was mixed with 1 pl 10X Plasmid-safe DNase buffer, 0.5 pl Plasmid-safe DNase (Biosearch Tech., Cat# E3101K), and 0.5 pl 100 mM ATP. Controls did not contain Plasmid-safe DNase. The mixture was incubated at 37 °C for 16 hours on thermocycler and followed by 70 °C for 30 minutes. One pl of the digested DNA was used for divergent PCR using CloneAMP HiFi PCR Premix (Takara Bio, Cat # 639398). The primer sequences are listed in Table 2. The amplified product was assayed by electrophoresis through a 0.8% agarose gel.1.15- 5-ethynyl-2'-deoxyuridine (EdU) staining of Drosophila follicle cells
[0243] Dissected fly ovaries were stained with the Click-iT EdU Cell Proliferation Kit (Thermo Fisher Scientific, Cat# C10337). The tissue was first washed twice in lx Phosphate- buffered saline (PBS) and incubated with 20 pM EdU for 60 minutes at room temperature. The tissue was then washed twice with lx PBS followed by fixation in 4% Paraformaldehyde for 15 minutes. Fixed tissues were washed twice with lx PBS and detected with Click-iT reaction as per manufacturer's instructions. Tissues were washed twice with lx PBS and wholemounted on microscope slides in Vectashield Antifade Mounting Medium with DAPI (Vector labs, Cat# H-1200-10). Stained tissues were imaged using a Leica SP5 Inverted Confocal Microscope was used for divergent PCR using CloneAMP HiFi PCR Premix (Takara Bio, Cat # 639398). The primer sequences are listed in Table 2. The amplified product was assayed by electrophoresis through a 0.8% agarose gel.1.16- ecDNA-sequencing and genome-sequencing of Drosophila ovaries
[0244] For ecDNA-sequencing, total DNA from ovaries was extracted by using Quick- gDNA MicroPrep Kit (Zymo Research, Cat # D3021). After removing linear DNA, rolling circle amplification, and debranching as described previously1, the library was prepared with the Ligation Sequencing Kit (Oxford Nanopore, Cat #SQK-LSK109). Briefly, 2 pg of total DNA was mixed with 2 pl Plasmid-safe DNase (Biosearch Tech., Cat # E3110K), 5 pl lOx Plasmid-safe buffer, 2 pl 100 mM ATP (Thermo Fisher Scientific. Cat # R0441), and ultrapure water (Thermo Fisher Scientific, Cat #10977023) to 50 pl. On a thermocycler machine, the mixture was incubated at 37 °C for 3 hours. Then 2 pl Plasmid-safe DNase and 1 pl ATP were added to the mixture. The mixture was further incubated at 37°C for 16 hours and 70°C for 30 minutes on a thermocycler machine. Then 50 pl AMPure XP beads (Beckman Coulter, Cat # A63881) was used to purify DNA. The concentration of the purified circular DNA was measured by Qubit dsDNA HS Assay kit (Thermo Fisher Scientific, Cat # Q33231). The RCA reaction was conducted as follows: 2 ng circular DNA, 5 pl 1 Ox Phi29 DNA Polymerase buffer, 1 pl Phi29 DNA Polymerase (NEB, Cat # M0269L), 2.5 pl 10 mM dNTP (Qiagen. Cat # 201901 ), 2.5 pl Exo Resistant Random Primer (Thermo Scientific. Cat # SO 181 ), and ultrapure water to 50 pl. The mixture was incubated at 30 °C for 16 hours and 65 °C for 10 minutes on a thermocycler machine. The RCA product was purified by isopropanol precipitation and debranched by T7 Endonuclease I (NEB, Cat # 0302L). The short fragments were eliminated by Short-Read Eliminator XL kit (Circulomics. Cat # SS-100-111-01). Then the library was constructed following the Oxford Nanopore SQK- LSK109 protocol. All libraries were sequenced in R9.4 flow cells on a GridlON instrument according to the manufacturer’s instructions.
[0245] For genome sequencing, fly ovarian gDNA libraries were generated using the Nextera XT DNA Library Preparation Kit (illumina, Cat# FC-131-1024). Briefly, fifty ng of gDNA was mixed with 10 pL of 2x Tagmentation buffer, 2.5 pl of Tagmentase, and up to 20 pl of nuclease-free water. The mixture was then incubated at 55 °C for 10 minutes. Thirty7pl of nuclease-free water was then added to the mix. and bead purified with 90 pl (1.8x) of AMPure XP beads (Beckman Coulter. Cat # A63881). Tagmented DNA was eluted with 10 pl of water. PCR amplification reaction was composed of the following: 10 pl of tagmented DNA was mixed with 5 pl of i7 index primer, 5 pl of i5 index primer, 5 pl of Universal primer mix / PPC, and 25 pl of 2X Ultra II Q5 Hotstart Master Mix (NEB, Cat# M0544L). On a thermocycler, the mixture was incubated at 72°C for 3 minutes, 98°C for 30 seconds, 10 cycles of thefollowing steps: 98°C for 15 seconds, 63°C for 30 seconds, 72°C for 3 minutes, and 72°C for 5 minutes. Libraries were size selected using AMPure XP beads, first at 0.6x to remove large fragments, then 1.2X beads to remove small fragments. Beads were washed with fresh 80% ethanol. DNA was eluted from beads with 28 pl of water. Library concentrations and quality were quantified using Qubit dsDNA HS Assay kit and an Agilent 4150 TapeStation System, respectively. Final libraries were sequenced at 150 bp paired-ends on an Illumina NextSeq 550 Instrument.1.17- Nanopore sequencing reads pre-processing and mapping
[0246] The fast5 files generated by the Nanopore GridlON machine were used as input in MinKNOW version 21.05.25 (MinKNOW core 4.3.12). Guppy 5.0.16 is integrated into MinKNOW. The basic data preprocessing parameters are the following: Basecall model = High-accuracy base-calling; Read filtering = 9; The passed fastq files produced by MinKNOW were used for further quality control. Adapter sequences were detected and trimmed by porechop (0.2.4) with parameters: — extra end trim 0 — discard_middle. This setting only removes the adapter sequencing detected at the beginning and the end of the reads, and if the adapter sequence is detected in the middle of the reads, the reads were filtered out. Output files of porechop were used for further analysis. Reads were mapped to the reference genome of Drosophila melanogaster version dm6 (GCA_000001215.4). Read mapping was performed using the minimap2 (2.17-r941)41 software with parameter settings -ax map-ont -Y -t 16 to keep the soft clipping sequences for all supplementary alignments in the SAM output. Mapped reads were converted to bam format, sorted by reference coordinates, and indexed by samtools (1.12)42. Data visualization was achieved by R (4.1.2) and Python (3.9.12). IGV (2.12.0)43 was used to visualize mapping results.1.18 - Western Blots
[0247] Cells were lysed in RIPA buffer (Thermo Scientific. Cat # 89900) with lx complete protease inhibitor cocktail (Roche. Cat # 4693159001). The lysate was resolved by SDS-PAGE gels, transferred with an iBlot2, and analyzed by immunoblotting with indicated primary antibodies. The following primary7antibodies were used in overnight incubation at 4°C: anti- DNA ligase 4 (Proteintech, Cat# 66705-1-Ig; 1: 1000), anti-XRCC4 (Proteintech, Cat # 15817- 1-AP; 1: 1000), anti-RAP80 (Cell Signaling Tech., Cat# 14466S; 1 :500), anti-a-Tubulin (Sigma Aldrich, Cat # SAB4500087; 1 : 10000), and anti-f>-Actin (Proteintech, Cat # 66009-1-Ig; 1:10000). Secondary7antibodies include: anti-mouse and anti-rabbit IgG-HRP (Thermo Scientific, Cat # G-21040 and # G-21234; 1:5000) or anti-mouse IRDye680RD (Licor, Cat#926-68070; 1: 10000) and anti-rabbit IRDye800CW (Licor, Cat# 926-32211; 1 :10000). The membrane bound with IgG-HRP was developed by SuperSignal West Pico PLUS Chemiluminescent Substrate Kit (Thermo Scientific, Cat # 34577) according to the manufacturer’s instructions. All other membranes bound were washed with lx TBST (Tris- Buffered Saline, 0.1% Tween 20) and incubated with IRDye secondary' antibodies for 1 hour at room temperature, followed by lx TBST washes. Membranes were imaged with a Licor Odyssey CLx.1.19 - CRISPR-C approach to generate DHFR and EGFR ecDNA[024S] sgRNAs to induce DHFR-containing ecDNA were obtained from a previous report1 1. which generate a 1.81 Mb size ecDNA. sgRNAs to induce EGFR-containing ecDNA were designed by the Integrated DNA Technologies (IDT) Custom Alt-R CRISPR-Cas9 guide RNA software (https: / / www.idtdna.com / site / order / designtool / index / CRISPR_CUSTOM). These sgRNAs were designed to target both flanks of EGFR and generate a 1.83 Mb size ecDNA. All the sgRNAs were purchased from Synthego (https: / / www.synthego.com / ). The sequences of the sgRNAs were listed in Table 2.
[0249] To generate ecDNA. HeLa or PC9 cells (either wild-type or LIG4-I- cells) were trypsinized. quenched with DMEM or RPMI-1640 (with 10% FBS). Cells were counted and 5xl05HeLa wild-ty pe W&LIG4 -I- cells, 8x105 PC9 wild-type m LIG4-l- cells were collected and centrifuged at 90 g for 10 minutes. The Cas9-sgRNA Ribonucleoprotein (RNP) complexes were introduced into cells by electroporation with Lonza 4D-Nucleofector™ system (X Unit). Briefly, the RNP complexes were assembled by mixing 2 pl of SpCas9 (IDT, Cat# 1081059, 61 pM), 1 pl of left-sgRNA and 1 pl of right-sgRNA (100 pM, in TE buffer, pH 8.0). The RNP mixtures were incubated at room temperature for 15 minutes. Then, 4 pl of RNP, 5 pl of electroporation enhancer (IDT, Cat# 1075916, 100 pM), 91 pl of cells (in SE solution, Lonza, Cat# V4XC-1024) were mixed, transferred to the Nucleocuvett (Lonza) and electroporated with Lonza 4D-Nucleofector™ system with the preset code: CN 114 for HeLa cells. EN-138 for PC9 cells. After electroporation, cells were transferred to 24-well plates with pre-warmed culture medium. For ecDNAjunction detection by ddPCR, HeLa cells were collected at 3 hours post electroporation; PC9 cells were collected at 12 hours post electroporation. Cells were washed, trypsinized, pelleted and stored at -80°C.1.20 - Drug treatments after CRIPSR-C induced ecDNA production
[0250] Seventy-two hours after CRIPSR-C to induce DHFR ecDNA formation, HeLa cells (either wild-type or LIG4 - / -) were counted, and 2x105 cells were seeded. Cells were allowed to recover for 24 hours, then treated with 100 nM of Methotrexate (Sigma- Aldrich, Cat# 454126) for one month. Medium (with 100 nM methotrexate) was replaced twice per week, meanwhile the live cells were counted under microscope.1.21 - Drug treatments to induce natural ecDNA production
[0251] To induce natural ecDNA formation in HeLa or PC9 cells: 8x105 of HeLa (wild-type or LIG4 - / -) cells, 5x105 of PC9 (wild-type or LIG4 - / -) cells were seeded into 10 cm dishes. After 24 hours of recovering, HeLa cells were subjected to treatment with 100 nM Methotrexate for up to one and half months. PC9 cells were treated with 20 nM Osimertinib (Selleck Chemicals, Cat# S7297) for up to two months. Culture medium supplemented with Methotrexate or Osimertinib were replaced twice per week. Cell numbers were monitored with microscope when replacing the medium.
[0252] Additionally, HeLa cells were treated with Mirin and Methotrexate: 8x105 of HeLa wild-type cells were seeded in 10 cm dish. Twenty-four hours later, the cells were treated with either Methotrexate (100 nM) alone or in combination with 1 pM Mirin. The treatment lasted for up to one and half months. Then the cells were subjected to ascending Methotrexate concentrations (200 nM, 400nM, lOOOnM), increasing the dose once the cells acquired resistance to each concentration, with Mirin remaining at 1 pM. Culture medium supplemented with Methotrexate or Methotrexate with Mirin were replaced twice per week. Cell numbers were monitored with a microscope when replacing the medium.1.22 - Metaphase DNA-FISH
[0253] Cells were arrested in metaphase with KaryoM AX™ Colcemid™ Solution (Thermo Scientific, Cat# 15212012) treatment at 0.2 pg / ml for 6 hours (HeLa), or 16 hours (PC9). Cells were washed with IX PBS, trypsinized, and resuspended in 75 mM KC1 for 15 min. Then the cells were fixed with fresh ice-cold Camoy’s fixative (3: 1 methanol: glacial acetic acid, v / v) for 10 minutes, followed by three additional washes with the fixative. Finally, the cells suspended in the fixative were dropped onto humidified glass slides. The slides w ere air-dried and aged at room temperature overnight. The slides ith fixed cells were briefly equilibrated in 2X SSC buffer (Thermo Scientific. Cat# AM9763), followed by dehydration in ascending ethanol series (70%, 85%, 100%) for 2 minutes each. Then the slides were air dried completely.Oligopaint FISH probes with hybridization buffer were mixed well and applied onto the slides, covered by coverslips, and sealed with rubber cement. Then the probes and samples were codenatured at 75°C for 5 minutes, and hybridization was carried out at 37°C overnight. The coverslips were removed, and the slides were washed in 0.4x SSC with 0.3% IGEPAL (72°C), and 2* SSC with 0.1% IGEPAL, for 2 minutes each. DNA was counter-stained with DAPI, and the slides were mounted with VECTASHIELD® Antifade Mounting Medium with DAPI (Vector Laboratories, Cat# H-1200-10). FISH images were acquired by Zeiss Axio Observer microscope using a 20X objective.
[0254] The custom Oligopaint probes used in this study were prepared based on the method developed from the laboratory of T. Wul2. Briefly, the oligo library was designed from the algorithm developed from the Wu lab. Each individual oligo consisted of a unique set of primer pairs for PCR amplification, and a T7 promoter sequence was attached in the forward primer to enable in vitro transcription. The Oligopaint-covered genomic regions used in this study were as follows: DHFR (chr5: 80,626,226-80,654,983), EGFR (chr7:55, 016, 118-55, 213, 816). The oligo pools for each ecDNA were synthesized by Twist Bioscience. The oligos were amplified by PCR. followed by in vitro transcription, and then the RNA was converted to ssDNA by reverse transcription, in which the fluorophores were introduced to the probes.1.23 - Statistics and Reproducibility
[0255] Statistical tests were calculated by using Rstudio (version 2023.12.0+369) and Microsoft Excel (vl6.67). The significance test of different groups was determined by a two- tailed, two sample unequal variance t-test. The experiments in FIGs. 2, 3. 4B (ecDNA-seq), 5A, 8, 9A, and 10A-10B were repeated at least three times independently with similar results. Similarly, the experiments in FIG. 4B (genome-seq), along with FIG. 9B were independently repeated twice, yielding consistent results. Biological replicates were used for all the independent experiments.Example 2 - A Version 1 Reporter Construct for Monitoring EcDNA Formation
[0256] Current tools for identifying and observing ecDNA are limited. To address this limitation, a CRISPR-C approach, using CRISPR / Cas9 cleavage to generate DNA fragments with the promoter located downstream of eGFP at the opposite end, was previously invented23. Under this design, circularization of the fragments brings the promoter upstream of eGFP. thus giving rise to fluorescence that reflects ecDNA formation23. Despite its elegant design, this system necessitates pre-integration of the eGFP cassette into the genome and the introductionof the Cas9 and sgRNAs to activate it23. Moreover, the circularization efficiency is not optimal, with maximally only -34% of the reporter-containing cells (HEK293T cells) displaying eGFP expression (data not shown)23. To engineer a more versatile and robust system, the prior CRISPR-C system was adapted to directly introduce linear DNA molecules into a cell that mimic the fragments generated by CRISPR / Cas9 cleavage and tested it in mammalian cells. Under this design, the sequences of linear DNA can be easily modified, and these linear molecules can be introduced into any cells and utilized as non-genomic reporter without a genomic pre-integration of the reporter. As a proof-of-principal, the biosensor was first designed with a single promoter and single reporter (eGFP) (FIG. 1 A, Version 1 biosensor, also referred to herein as “VI”). Transfecting these DNA molecules into HEK293T cells gave eGFP positive cells, suggesting the formation of ecDNA (FIG. IB). The percentage of eGFP- positive cells increased as the dose of DNA molecules accrued, with the capability of achieving > 80% (data not shown), suggesting a high efficiency for ecDNA biogenesis from this “V ersion 1” (i.e., “VI”) biosensor.
[0257] To validate that eGFP expression reflects ecDNA production and is not due to the introduced eGFP sequences integrating into the host cell genome with a local promoter driving eGFP expression, two other control constructs were designed (FIG. IB). One is comprises the eGFP sequence without a promoter (FIG. IB). The other is a circle harboring the designed biosensor (FIG. IB), thus preventing the joining of the eGFP and promoter. Introducing either construct into cultured cells could not drive eGFP expression, suggesting eGFP expression from this VI biosensor reflects circular DNA formation. Meanwhile, given that the cells can take up multiple biosensor molecules, their oligomerization can form linear DNA and bring a promoter upstream of eGFP. To test this possibility7, DNA from the eGFP-positive cells was treated with exonuclease, which removes linear DNA but leaves ecDNA intact. By using a pair of primers that spans the junction site of eGFP and promoter (FIG. 1A), it was observed that the PCR signals were preserved upon exonuclease treatment (FIG. 1C), indicating that the V 1 biosensor mainly forms ecDNA.Example 3 - A Version 2 Reporter Construct for Monitoring EcDNA Formation
[0258] Given that the Version 1 biosensor could not reflect the percentage of the cells that uptook our biosensor, a DsRed cassette was introduced into a new version of the biosensor that is directly driven by its own promoter into the middle of the biosensor (FIGS. 1D-1E, Version 2 biosensor, also referred to herein as “V2”). As such, DsRed fluorescence signal can be used to monitor the cells harboring our biosensor. The version 2 biosensor was applied to multiplecancer cell lines: cervical cancer HeLa cells, non-small cell lung cancer PC9 cells, and colon cancer HCT116 cells. Although these cell lines showed varied efficiency on uptake of the biosensors, as reflected by DsRed expression, they all robustly expressed eGFP (FIG. IF), indicating their capability for ecDNA biogenesis. The VI and V2 biosensors described herein are versatile and robust biosensors that allows us to monitor the ecDNA formation process.Example 4 - CRISPR screens identify ecDNA biogenesis regulators
[0259] Given that eGFP expression can be used to reflect ecDNA biogenesis, genome-wide CRISPR screening was performed to identify factors regulating this process (i.e.. the formation of ecDNA). For this purpose, the original reporter from the CRISPR-C system was chosen (i.e., not the VI or V2 biosensors described above) that produces eGFP after being cut off of the genome by CRIPSR / Cas9 — rather than the VI and V2 engineered biosensors of FIG. 1 — for two reasons. First, all cells have the same copy number of the reporter pre-integrated into their genome (data not shown). Second, given that the process of creating DNA breaks may couple with ecDNA production, generating the reporter fragment from the chromosome allowed for the capture of all potential factors involved in this process. For the screen, HEK293T cells was used that harbored the original CRISPR-C reporter to mutate one gene / cell for 18.761 genes, by transducing the genome-wide MinLib CRISPR lentiviral library24. Upon inducing ecDNA production and eGFP expression, eGFP-positive and eGFP-negative cells were sorted to achieve 1,000-fold coverage of the sgRNA library. Under this design, sgRNAs targeting genes required for ecDNA biogenesis were enriched in the eGFP-negative population and excluded from the eGFP-positive population (FIG. 2A). Conversely, sgRNAs targeting genes encoding factors that suppress ecDNA formation would show the reverse pattern (FIG. 2A).
[0260] Given that ecDNA biogenesis is coupled with the DNA break formation, factors from DNA repair pathways w ere the most significantly enriched screen hits (FIG. 2B and FIG. 6A). Meanwhile, factors from other processes were also destected as potential regulators for ecDNA biogenesis (FIGS. 6A-6B). For example. GO term analysis revealed factors related with rRNA processing or ribosome biogenesis suppress ecDNA production under this screen design (FIGS. 6A-6B). While these findings point to directions for future investigations, according to this example, the primary focus on the function of DNA repair factors in ecDNA biogenesis.
[0261] Among the top hits for driving ecDNA production are Lig4 and its physical partner XRCC4, and DNA-PKcs (encoded by the PRKDC gene) (FIGS. 2B-2D). Notably, DNA-PKcs has previously been implicated in mediating ecDNA production when cancer cells undergochromothripsis16, giving credence to the screening approach of this example. Despite the Lig4 / XRCC4 complex and DNA-PKcs being classified into the non-homologous end joining (NHEJ) DNA repair process25, 26, the other factors from this pathway were not uncovered as significant hits from the screen of this example (FIG. 2D). The NHEJ core factor Ku80 and Ku70 are essential for the viability of the human cells27, 28, confounding the efforts to discern them as ecDNA biogenesis factors. Other NHEJ factors (ARTEMIS, PAXX, POLL, and POLM) are required when the ends possess non-complementary overhangs29, thus would be expected to be dispensable when the DNA fragments harbor blunt ends, such as under the present screening condition.
[0262] Instead, the 2nd top hits for driving ecDNA biogenesis are factors that form the BRCA1-A complex (BABAM2. ABRAXAS 1, and RAP80 encoded by gene UIMC1, FIG. 2B), which acts at the upstream of DNA break signaling and is linked to regulate the Homologous Recombination (HR) pathway choice30'33. The BRCA1-A complex has been characterized for its function in preventing end-resection from the DNA break sites30'33. In accordance with the finding that preventing end-resection promotes ecDNA production, the screen of this example identified factors that drive end-resection as suppressors of ecDNA biogenesis (FIGS. 2B, 2D and FIGS. 6A-6B). These factors are MRE11, NBN1, and RAD50 (FIGS. 2B, 2D and FIG. 6B), which form the MRN complex that resects the DNA break ends and recruits ATM34. Consistently, ATM is screened as a negative regulator of ecDNA production (FIG, 2B and FIGS 6B). In summary, the present screen results suggest that upon DNA fragmentation, preventing end resection drives the ecDNA production process that is potentially catalyzed by the Lig4 complex.Example 5 - Validation of Screen Hits
[0263] To validate the screen hits that serve as drivers for ecDNA production, CRISPR / Cas9 was used on the reporter HEK293T cells to generate mutant cell lines for the following factors: Lig4. XRCC4, DNA-PKcs, RAP80 from the BRCA1-A complex (FIGS. 8A-8B). For each factor, at least tw o different sgRNAs were employed to generate distinct mutant clones (FIG. 8A), minimizing the possibility7of obtaining false-positive findings from the off-targeting mutation by an individual sgRNA. Using eGFP expression as the proxy for ecDNA production from the reporter, findings consistent with the screen results from the prior example were observed. Lig4 mutants generated by three independent sgRNAs showed no eGFP expression (FIG. 8B), suggesting that the cells without Lig4 lost the capability of ecDNA formation. Similar findings were obtained for XRCC4 mutants, PRKDC mutants, UIMC1 mutants (datanot shown). Besides using eGFP to indicate ecDNA production, a digital droplet PCR (ddPCR)-based assay was designed to precisely quantify the number of ecDNA produced from the reporter utilized for this example (FIG. 3A, not the VI or V2 biosensor described in the examples above). Consistent with eGFP expression, mutating Lig4, XRCC4, PRKDC, UIMC1 abrogated ecDNA production, as measured by the ddPCR assay (FIG. 8B and additional data not shown).
[0264] On the other end, validation of the MRN complex as a suppressor of ecDNA production via two independent approaches was sought. First, two RAD50 mutant cell lines were generated using two distinct sgRNAs (data now shown). Upon mutating RAD50, the ecDNA-producing cells significantly increased from 21.2% to 37.4% (P-value = 0.030, sgRNA-1) and 36% (P-value = 0.046, sgRNA-2, additional data not shown). Next, validation of the antagonizing function of the MRN complex was performed by using a small molecule drug, Mirin, which serves as an MRN inhibitor to block the end-resection function of MRE1 I35,36Applying this inhibitor significantly increased the eGFP-positive population by 2.7^-fold (P-value = 0.004. additional data not shown). Altogether, these results robustly validated the screen findings and support a model that suppressing end-resection promotes the Lig4 complex-mediated ecDNA biogenesis.Example 6 - DNA Ligase 4 (Lig4) Catalyzes EcDNA Formation
[0265] Given its apparent ligase activity and position as the strongest screen hit, Lig4 likely serves as the key enzyme that catalyzes the end-end joining events to form ecDNA. Alternatively, Lig4 may mediate other cellular processes that indirectly promote ecDNA biogenesis. It is worth noting that mutating Lig4 does not give rise to any apparent cell growth defects (data not shown). As such, it was next tested whether Lig4 directly drives ecDNA production, and the rescue experiments] w ere performed by re-introducing either wild-ty pe or catalytically-dead Lig4 into its corresponding mutant cells (data not shown). Meanwhile, since XRCC4 stabilizes Lig4 by directly interacting with its two BRCT domains37, 38, a mutant version of Lig4 w as also included by removing the BRCT domains — termed as Lig4-ABRCT- -for the rescue experiment (data not shown). As expected, re-introducing wild-type Lig4 restored ecDNA production in the mutant cells, as evidenced by eGFP expression from the reporter (data not shown). However, either mutating Lig4 ligase activity or suppressing its interaction with XRCC4 abolished the rescue outcome (data not shown). These findings suggest that Lig4 is the ligase that directly mediates ecDNA formation.
[0266] Besides the HEK293T cells, it was next tested whether Lig4 is essential for ecDNA biogenesis in other cell lines. Given that the V2 engineered biosensor described above can be easily applied to any cell lines, the Version 2 biosensor was used for this purpose. Mutating Lig4 abrogated ecDNA production in colon cancer HCT116 cells (data not shown). It is worth noting that a previous study suggested a function of LIG3 (the other ligase encoded in the human genome) in mediating circular DNA formation when cells undergo apoptosis39. However, loss of LIG3 did not affect ecDNA biogenesis in HCT116 cells (data not shown), suggesting that in cycling HCT116 cells, it is Lig4, not LIG3, that catalyzes the ecDNA biogenesis. Besides HCT116, the function of Lig4 in HeLa and PC9 cells was also tested. Similarly, mutating LIG4 lead to the abolishment of ecDNA production in these cells (data not shown), suggesting a generic function of Lig4 in mediating ecDNA biogenesis.
[0267] Next, it was reasoned that having homology sequences at the two ends may impact the essentiality of Lig4 on ecDNA formation. To this end, the possibility of having identical sequences with different lengths at the two ends was calculated if the human genome was radonmly fragmented into different sizes: 1 kb, 2 kb. 5 kb, 10 kb, 50 kb, 100 kb, 500 kb, 1,000 kb, and 5,000 kb (data not shown). Regardless of the fragment size, the chance of having 2 identical nucleotides at the two ends is 6.71%, and this frequency decreases as the length of the homology' sequence increased (1.86% for 3 bp homology7, 0.52% for 4 bp homology, and 0. 15% for 5 bp homology, data not shown). Although the chances of having identical sequences at the two ends as homology are low. it was still tested whether Lig4 is required when the homology sequences are present (data not shown). HeLa cells without Lig4 failed to produce ecDNA even when the two ends have 2 bp overlap (data not shown). As the length of homology increased to 3-5 bp, a few LIG4 mutant HeLa cells produced ecDNA (data not shown). Although the frequency was still dramatically lower than the wild-type cells (data not shown), this finding suggests that a compensatory program can produce ecDNA when HeLa cells lost Lig4 and meanwhile having > 3 bp homology at the two ends of a given DNA fragment (with a chance < 3%). In HCT116 and PC9 cells, mutating LIG4 completely abolished circle formation even when the tw o ends have 5 bp homology (data not shown), indicating that these two cell lines strongly depend on Lig4 for ecDNA biogenesis even when up to 5 bp homology presents. Interestingly, it was previously found that Lig4 drives 2-LTR-ecDNA production in both Drosophila and mammalian cells for the replicated retrotransposon DNA19. These molecules have the identical long terminal repeat (LTR) sequences that are hundreds of nucleotides in length at the two ends, but still count on Lig4 for the circularization process19.
[0268] Lastly, while these findings suggest Lig4 as the key factor driving ecDNA formation, these findings were based on two reporter systems that produced DNA fragments with fixed DNA sequences at the two ends. Is Lig4 still essential if the DNA sequences at the two ends are altered? To address this question, the versatility of the present biosensors were harnessed by including 6 random nucleotides at the two ends of the Version 2 biosensor (data not shown). In all cells tested (HCT116. PC9, and HeLa cells), mutating LIG4 lead to abrogation of ecDNA formation from the Version 2 biosensor with random end sequences (data not shown). Altogether, these findings suggest that when cells experience DNA fragmentation, Lig4 drives ecDNA biogenesis regardless of the nucleotide composition at the two ends.Example 7 - Lig4 mediates natural ecDNA formation in vivo
[0269] During Drosophila egg development, two chorion genomic clusters undergo massive amplification in the oocyte-surrounding follicle cells to immensely produce chorion protein for egg encapsulation40. Previous studies showed that the uneven amplification of chorion clusters often leads to DNA breaks with variable length (FIG. 4A)41. All circular DNA in Drosophila ovaries was previously sequenced to examine retrotransposon-derived ecDNA19. Mining this dataset, sequencing reads were found with head-to-tail junctions that support ecDNA generated from the two chorion clusters (FIG. 9A). To investigate whether Lig4 is required for the formation of chorion ecDNA, ovarian ecDNA from LIG4 homozygotes (LIG4~ ~) and heterozygotes (I.IG4 ) was sequenced. While the LIG4 heterozygotes still produced ecDNA from both chorion regions, loss of Lig4 resulted in a nearly 100% abolishment of ecDNA production (FIG. 4B and FIG. 9B). To further validate these findings, a pair of primers was designed that can only yield PCR products when ecDNA is generated (FIG. 4A, 4C). Similarly, while ecDNA with variable length can be detected from LIG4' / +ovaries, LIG4 ~ flies lost chorion-ecDNA production (FIG. 4C).
[0270] While these findings suggest Lig4 directly drives chorion ecDNA production in vivo, it is possible that Lig4 is essential for chorion region amplification, thus its depletion would abolish ecDNA biogenesis. To test this possibility, the ovarian genomic DNA from both LIG4~4and LIG4^ flies was sequenced. Regardless of the genotype, the two chorion clusters amplified to the same magnitude (FIG. 4B and FIG. 9B). suggesting that Lig4 is dispensable for the genomic amplification of these two regions. The amplification process of these two regions can incorporate a large amount of EdU into the replicated DNA and be visualized by EdU staining. Consistently, bright EdU foci were observed in the follicle cells from both LIG4~ ’ and LIG4~ / +flies (FIG. 4D), further indicating that mutating LIG4 does not alter chorioncluster amplifation but does abolish ecDNA production. In summary, these data suggest that Lig4 plays a conserved function in driving natural ecDNA production during Drosophila egg development.Example 8 - Lig4 drives ecDNA-mediated cancer cell adaptation
[0271] Oncogenic ecDNA have been extensively show n to contribute to tumor heterogeneity7and cancer cell adaptation3'10. It was next sought to test if Lig4 is required for the formation of oncogenic ecDNA. Previous research established the CRISPR-C approach to generate megabase-sized ecDNA species containing the DHFR gene, which allows cancer cells to adapt to the chemotherapy drug, methotrexate11. The same approach was employed in the present example and used ddPCR to quantify DHFR ecDNA biogenesis, a method established previously11. While wild-type HeLa cells generated 1.81 mega-bases DHFR ecDNA and survived from the methotrexate treatment by accumulating a massive amount of DHFR ecDNA (FIGS. 5A-5C, additional data not shown), mutating LIG4 abrogated DHFR ecDNA production (FIG. 5A and additional data not shown. P = 0.0002). Correspondingly, loss of Lig4 in HeLa cells resulted in complete cell-death upon methotrexate exposure (FIG. 5B). Moreover, the same CRISPR-C approach was used to generate ecDNA harboring the EGFR gene in PC9 cells. Similarly, while wild-type PC9 produced EGFR ecDNA upon introducing CRISPR / Cas9 to generate a 1.83 mega-base segment, the amount of ecDNA generated m I G4 mutant PC9 cells significantly dropped to 17% of the wild-type level (FIG. 5A and additional data not shown, P < 0.0001).
[0272] The CRISPR-C approach relies on CRISPR / Cas9 to generate pre-defined DNA fragments for ecDNA biogenesis. Without CRISPR / Cas9, is Lig4 is still required for natural ecDNA biogenesis in cancer cells? To answer this question, the function of Lig4 in two systems was tested: (1) methotrexate (MTX) induced spontaneous DHFR ecDNA production in HeLa cells; and (2) EGFR inhibitor induced natural EGFR ecDNA biogenesis in PC9 cells. Pioneering research in 1970s established the system that upon methotrexate treatment, HeLa cells rely on ecDNA — at that time termed “double minutes” — harboring the DHFR gene to evolve resistance for this chemotherapy drug42'45. This system was adapted for the present example and tested the function of Lig4 during this natural ecDNA-mediated evolution process (FIG. 5D). While wild-type cells can adapt to methotrexate by producing DHFR ecDNA, mutating Lig4 led to the abolishment of the evolution (FIGS. 5E, 5F). To examine whether Lig4 is only essential for this one cell-ecDNA pair or is broadly required for natural ecDNA production, PC9 lung cancer cells were established as another system to study spontaneousEGFR ecDNA production. After optimizing the dose of Osimertinib, the 3rd-generation EGFR inhibitor, the formation of natural EGFR ecDNA to facilitate the adaptation process for PC9 cells was observed (FIG. 5E). Similar to HeLa cells, while wild-type cells evolved resistance to Osimertinib, mutating LIG4 resulted in the abrogation of adaptation (FIG. 5F). Given that mutating LIG4 in both PC9 and HeLa cells led to cell death upon drug treatment (FIG. 5F), the direct function of Lig4 in mediating natural ecDNA production in these two systems could not be tesetd. However, these data provide strong evidence to suggest the essentiality of Lig4 for the process of ecDNA-mediated cancer cell adaptation.
[0273] Lastly, since the CRISPR screen of the examples uncovered the MRN complex as a suppressor for ecDNA production from the eGFP reporter (FIGS. 2B, 2D and FIG. 6B), it was also tested whether the MRN complex inhibits natural ecDNA-mediated cancer cell adaptation. To this end, Mirin was applied to suppress the function of the MRN complex and monitor the process of HeLa cells adapting to methotrexate treatment (FIGS. 10A, 10B). Suppressing MRN complex function promoted the adaptation process, as evidenced by the expedited ecDNA production and quicker adaptation to drug treatment (FIGS. 10A, 10B). These findings suggest that the regulators uncovered from this screen not only control ecDNA production from a kilobases reporter, but also mediate the natural formation process of large ecDNA (on the order of megabase pairs) in cultured cancer cells.Discussion
[0274] ecDNA biogenesis is a fundamental biological process that frequently occurs in genomes across species. Previous studies aimed to delineate the function of specific DNA repair pathways in driving this process46, 47Largely relying on analyzing the nucleotide sequences at the end-end junction sites for deduction but lacking direct genetic data, models have been proposed separately favoring a role for NHEJ or MMEJ in this process3, 4’46, 47. According to the present examples, genome-wide CRIPSR screening was performed to identify factors that mediate ecDNA formation. These data suggest that the ecDNA biogenesis process is not solely controlled by a single DNA repair pathway. Rather, selective factors from different DNA repair steps orchestrate ecDNA generation. This work suggests that upon DNA fragmentation, the end-processing complexes with opposite end-resection functions antagonize each other to funnel the un-resected ends for a Lig4-catalyzed ecDNA production event. Besides frequently forming from the fragmented DNA, ecDNA could be produced via other mechanisms, such as direct recombination of the homology sequences from the chromosomes in yeasts or during the telomere lengthening process in cancer cells13, 48These findings notonly delineate a mechanism that likely frequently drives ecDNA production from genomic fragments, but also serve as a solid anchor point to compare the similarities or differences of distinct ecDNA formation processes.
[0275] Given that ecDNA is frequently generated in cancer cells to drive tumor evolution and adaptation, targeting ecDNA biogenesis represents a novel strategy for developing new cancer therapies. According to the present examples it is reported report Lig4 sen es as an important factor for catalyzing ecDNA formation. Notably, these data show that cancer cells without Lig4 lost their capability of initiating ecDNA-mediated adaptation. These findings highlight Lig4 as a potential target for cancer therapy. Humans appear to be able to tolerate the loss of Lig4 at a high level49. Mutations in LIG4 leads to Lig4 syndrome, a disease from which patients show developmental microcephaly and growth retardation but only have manageable immune deficiency in adult life49. Moreover, these data show that Lig4 normally is not required for cell viability, but only appears to be critical when cancer cells need to adapt to drug treatment. The above information suggests that the potential toxicity from targeting Lig4 should be minimal and controllable, suggesting a feasibility to drug Lig4. In fact, with the potential usage to sensitize cancer cells for radiotherapy, efforts were spent in the past to develop Lig4 drugs50'53, but failed to reach a fruition. Uncovering its essential function in driving ecDNA biogenesis and cancer cell adaptation, this work calls for innovative approaches to develop Lig4 inhibitors for cancer therapy.Example 9 - Mutating the BRCA1-A Complex Core Component UIMC1 Also Leads to Failure of Cancer Cells to Evolve Chemotherapy Resistance
[0276] Mutating the BRCA1-A complex core component UIMC1 also leads to failure of cancer cells to evolve resistance to methotrexate. As shown in FIG. 11, HeLa cells or UIMC1 - / - HeLa cells w ere treated with either DMSO or methotrexate. DMSO treatment does not lead to cell death regardless of genotype. The HeLa cells eventually developed resistance to methotrexate. In contrast, the absence of the U1MC1 gene (which encodes for the Rap80 protein) in the UIMC1 - / - HeLa cells did not.List of References
[0277] Below' is a list of references relevant to the Example 1:1. Yang. F. et al. Retrotransposons hijack alt-EJ for DNA replication and eccDNA biogenesis. Nature. doi: 10.1038 / s41586-023-06327-7 (2023).2. Moller, H. et al. CRISPR-C: circularization of genes and chromosome by CRISPR in human cells. Nucleic Acids Research, doi: 10.1093 / nar / gky767 (2018).3. Klann, T. S. et al. CRISPR-Cas9 epigenome editing enables high-throughput screening for functional regulatory elements in the human genome. Nature Biotechnology, doi: 35:6 35, 561 (2017)4. Dempster, J. M. et al. Extracting Biological Insights from the Project Achilles Genome-Scale CRISPR Screens in Cancer Cell Lines. http: / / biorxiv.org / lookup / doi / 10. 1101 / 720243 (2019) doi: 10. 1101 / 720243.5. Amici, D. R. et al. FIREWORKS: a bottom-up approach to integrative coessentiality network analysis. Life Sci. Alliance 4, e202000882 (2021).6. Zhou. Y. et al. Metascape provides a biologist-oriented resource for the analysis of systems-level datasets. Nat. Commun. 10, 1523 (2019).7. Chatr-aryamontri, A. et al. The BioGRID interaction database: 2017 update. Nucleic Acids Res. 45, D369-D379 (2017).8. Szklarczyk, D. et al. The STRING database in 2023: protein-protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Res. 51, D638-D646 (2023).9. Shannon, P. et al. Cytoscape: A Software Environment for Integrated Models of Biomolecular Interaction Networks. Genome Res. 13, 2498-2504 (2003).10. Quinlan, A. R. & Hall, I. M. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics 26, 841-842 (2010).11. Lange, J.T. et al. The evolutionary dynamics of extrachromosomal DNA in human cancers. Nat Genet 54, 1527-1533 (2022).12. Beliveau, B.J. et al. In Situ Super-Resolution Imaging of Genomic DNA with OligoSTORM and OligoDNA-PAINT. Methods in Molecular Biology, 1663: 231— 252. (2017).
[0278] Below is a list of other references relevant to this disclosure:13. Brown DD, Dawid IB. Specific gene amplification in oocytes. Oocyte nuclei contain extrachromosomal replicas of the genes for ribosomal RNA. Science.1968;160(3825):272-80. Epub 1968 / 04 / 19. doi: 10. 1126 / science. 160.3825.272.14. Hotta Y, Bassel A. Molecular Size and Circularity of DNA in Cells of Mammals and Higher Plants. Proc. Nat’l Acad. Sci. USA. 1965; 53:356-62. doi: 10.1073 / pnas.53.2.356.Wu S. et al. Extrachromosomal DNA: An Emerging Hallmark in Human Cancer. Annu Rev Pathol. 2022;17:367-86. Epub 2021 / 11 / 10. doi: 10. 1146 / annurev- pathmechdis-051821-114223. Noer JB, et al. Extrachromosomal circular DNA in cancer: history, current knowledge, and methods. Trends in Genetics. 2022; 38(7):766-781. Epub 2022 / 03 / 13. doi: 10.1016 / j.tig.2022.02.007. Zhu Y, et al. Oncogenic extrachromosomal DNA functions as mobile enhancers to globally amplify chromosomal transcription. Cancer Cell. 2021;39(5):694-707 e7. Epub 2021 / 04 / 10. doi: 10. 1016 / j.ccell.2021.03.006. PubMed PMID: 33836152. Hung KL. et al. ecDNA hubs drive cooperative intermolecular oncogene expression. Nature. 2021. Epub 2021 / 11 / 26. doi: 10.1038 / s41586-021-04116-8. Koche RP, et al. Extrachromosomal circular DNA drives oncogenic genome remodeling in neuroblastoma. Nature Genetics. 2020;52(l):29-34. Epub 2019 / 12 / 18. doi: 10.1038 / s41588-019-0547-z. Kim H, et al. Extrachromosomal DNA is associated with oncogene amplification and poor outcome across multiple cancers. Nature Genetics. 2020. Epub 2020 / 08 / 19. doi: 10. 1038 / s41588-020-0678-2. Wu S, et al. Circular ecDNA promotes accessible chromatin and high oncogene expression. Nature. 2019;575(7784):699-703. Epub 2019 / 11 / 22. doi: 10.1038 / s41586- 019-1763-5. Morton AR, et al. Functional Enhancers Shape Extrachromosomal Oncogene Amplifications. Cell. 2019;179(6): 1330-41 el3. Epub 2019 / 11 / 26. doi:10. 1016 / j.cell.2019.10.039. Lange JT. et al. The evolutionary dynamics of extrachromosomal DNA in human cancers. Nature genetics. 2022;54(10): 1527-33. Epub 2022 / 09 / 19. doi:10. 1038 / s41588-022-01177-x. Durkin K, et al. Serial translocation by means of circular intermediates underlies colour sidedness in cattle. Nature. 2012;482(7383):81-4. Epub 2012 / 02 / 01. doi: 10.1038 / naturel0757. Libuda DE, Winston F. Amplification of histone genes by circular chromosome formation in Saccharomyces cerevisiae. Nature. 2006;443(7114): 1003-7. Epub 2006 / 10 / 27. doi: 10.1038 / nature05205.Kalt MR, Gall JG. Observations on early germ cell development and premeiotic ribosomal DNA amplification in Xenopus laevis. The Journal of Cell Biology. 1974;62(2):460-72. Epub 1974 / 08 / 01. doi: 10.1083 / jcb.62.2.460. Ilic M, et al. Life of double minutes: generation, maintenance, and elimination. Chromosoma. 2022; 131(3): 107-25. Epub 20220430. doi: 10.1007 / s00412-022- 00773-4. Shoshani O, et al. Chromothripsis drives the evolution of gene amplification in cancer. Nature. 2021; 591 (7848): 137-41. Epub 2020 / 12 / 29. doi: 10. 1038 / s41586-020- 03064-z. Rosswog C, et al. Chromothripsis followed by circular recombination drives oncogene amplification in human cancer. Nature Genetics. 2021; 53(12): 1673-85. Epub 20211115. doi: 10.1038 / s41588-021 -00951 -7. Turner KM, et al. Extrachromosomal oncogene amplification drives tumour evolution and genetic heterogeneity. Nature. 2017; 543(7643): 122-5. Epub 20170208. doi: 10.1038 / nature21356. Yang F, et al. Retrotransposons hijack alt-EJ for DNA replication and eccDNA biogenesis. Nature. 2023;620:218-225. Epub 2023 / 07 / 12. doi: 10.1038 / s41586-023- 06327-7. Wells JN, Feschotte C. A Field Guide to Eukaryotic Transposable Elements. Annual Review of Genetics. 2020. Epub 2020 / 09 / 22. doi: 10.1146 / annurev-genet-040620- 022145. Kazazian HH, Jr., Moran JV. Mobile DNA in Health and Disease. The New England Journal of Medicine. 2017;377(4):361-70. doi: 10.1056 / NEJMral510092. Bums KH. Transposable elements in cancer. Nature reviews Cancer. 2017; 17(7):415- 24. Epub 2017 / 06 / 09. doi: 10.1038 / nrc.2017.35. Moller HD, et al. CRISPR-C: circularization of genes and chromosome by CRISPR in human cells. Nucleic Acids Research. 2018;46(22):el31. doi: 10.1093 / nar / gky767. Goncalves E, et al. Minimal genome-wide human CRISPR-Cas9 library. Genome Biology. 2021;22(l):40. Epub 2021 / 01 / 21. doi: 10.1186 / sl3059-021-02268-4. Lieber MR. The mechanism of double-strand DNA break repair by the nonhomologous DNA end-joining pathway. Annu Rev Biochem. 2010;79: 181-211. doi : 10.1146 / annure v . biochem.052308.093131.Scully R, et al. DNA double-strand break repair-pathway choice in somatic mammalian cells. Nature Reviews Molecular Cell Biology. 2019;20(l 1):698-714. Epub 2019 / 07 / 01. doi: 10. 1038 / s41580-019-0152-0. Wang T, et al. Identification and characterization of essential genes in the human genome. Science. 2015;350(6264):1096-101. Epub 2015 / 10 / 15. doi: 10.1126 / science.aac7041. Blomen VA, et al. Gene essentiality and synthetic lethal ity in haploid human cells. Science. 2015;350(6264): 1092-6. Epub 2015 / 10 / 15. doi: 10.1126 / science.aac7557. Zhao B, et al. The molecular basis and disease relevance of non-homologous DNA end joining. Nature reviews Molecular Cell Biology. 2020;21(12):765-81. Epub 20201019. doi: 10.1038 / s41580-020-00297-8. Hu Y, et al. RAP80-directed tuning of BRCA1 homologous recombination function at ionizing radiation-induced nuclear foci. Genes & Development. 2011;25(7):685-700. Epub 2011 / 03 / 15. doi: 10.1101 / gad.20H011. Vohhodina J, et al. RAP80 and BRCA1 PARsylation protect chromosome integrity by preventing retention of BRCA1-B / C complexes in DNA repair foci. Proc. Nat’l Acad. Sci. USA. 2020;l 17(4):2084-91. Epub 2020 / 01 / 13. doi: 10. 1073 / pnas. 1908003117. Sobhian B, et al. RAP80 targets BRCA1 to specific ubiquitin structures at DNA damage sites. Science. 2007;316(5828): 1198-202. doi: 10.1126 / science.l 139516. Her J. et al. Factors forming the BRCA1-A complex orchestrate BRCA1 recruitment to the sites of DNA damage. Acta Biochim Biophys Sin (Shanghai). 2016;48(7):658- 64. Epub 2016 / 06 / 20. doi: 10.1093 / abbs / gmw047. Blackford AN, Jackson SP. ATM, ATR, and DNA-PK: The Trinity at the Heart of the DNA Damage Response. Molecular Cell. 2017;66(6):801-17. doi: 10.1016 / j.molcel.2017.05.015. Dupre A, et al. A forward chemical genetic screen reveals an inhibitor of the Mrel 1- Rad50-Nbsl complex. Nat Chem Biol. 2008;4(2): 119-25. Epub 2008 / 01 / 06. doi: 10.1038 / nchembio.63. Gamer KM, et al. Corrected structure of mirin. a small-molecule inhibitor of the Mrel l-Rad50-Nbsl complex. Nat Chem Biol. 2009;5(3): 129-30; Author Reply 30. doi: 10. 1038 / nchembio0309-129. Wu PY, et al. Structural and functional interaction between the human DNA repair proteins DNA ligase IV and XRCC4. Molecular and Cellular Biology.2009;29(l 1 ):3163-72. Epub 2009 / 03 / 30. doi: 10.1128 / MCB.01895-08.Sibanda BL, et al. Crystal structure of an Xrcc4-DNA ligase IV complex. Nat Struct Biol. 2001;8(12): 1015-9. doi: 10.1038 / nsb725. Wang Y, et al. eccDNAs are apoptotic products with high innate immunostimulatory activity. Nature. 2021. Epub 2021 / 10 / 22. doi: 10.1038 / s41586-021-04009-w. Spradling AC, Mahowald AP. Amplification of genes for chorion proteins during oogenesis in Drosophila melanogaster. Proc. Nat’l Acad. Sci. USA. 1980:77(2): 1096- 100. doi: 10.1073 / pnas.77.2.1096. Yarosh W, Spradling AC. Incomplete replication generates somatic DNA alterations within polytene salivary' gland cells. Genes & Development. 2014;28(16):1840-55. doi: 10.1101 / gad.245811.114. Snapka RM, Varshavsky A. Loss of unstably amplified dihydrofolate reductase genes from mouse cells is greatly accelerated by hydroxyurea. Proc. Nat’l Acad. Sci. USA. 1983;80(24):7533-7. doi: 10.1073 / pnas.80.24.7533. Brown PC, et al. Relationship of amplified dihydrofolate reductase genes to double minute chromosomes in unstably resistant mouse fibroblast cell lines. Molecular and Cellular Biology. 1981; 1(12): 1077-83. doi: 10.1128 / mcb.l. 12.1077-1083.1981. Kaufman RJ, et al. Amplified dihydrofolate reductase genes in unstably methotrexateresistant cells are associated with double minute chromosomes. Proc. Nat’l Acad. Sci. USA. 1979:76(11):5669-73. doi: 10. 1073 / pnas.76. 11.5669. Balaban-Malenbaum G, Gilbert F. Double minute chromosomes and the homogeneously staining regions in chromosomes of a human neuroblastoma cell line. Science. 1977;198(4318):739-41. doi: 10.1126 / science.71759. Dutta A, et al. Microhomology -mediated end joining is activated in irradiated human cells due to phosphorylation-dependent formation of the XRCC1 repair complex. Nucleic Acids Research. 2017:45(5):2585-99. Epub 2016 / 12 / 21. doi: 10.1093 / nar / gkwl262. Paulsen T, et al. MicroDNA levels are dependent on MMEJ, repressed by c-NHEJ pathway, and stimulated by DNA damage. Nucleic Acids Research.2021:49(20): 11787-99. doi: 10.1093 / nar / gkab984. Chen YA, et al. Extrachromosomal telomere repeat DNA is linked to ALT development via cGAS-STING DNA sensing pathway. Nature Structural & Molecular Biology. 2017;24(12):1124-31. Epub 20171106. doi: 10. 1038 / nsmb.3498. Altmann T, Gennery AR. DNA ligase IV syndrome: a review. Orphanet J Rare Dis. 2016;l 1(1): 137. Epub 2016 / 10 / 07. doi: 10. 1186 / sl3023-016-0520-l.62. Menchon G, et al. Structure-Based Virtual Ligand Screening on the XRCC4 / DNA Ligase IV Interface. Sci Rep. 2016;6:22878. Epub 2016 / 03 / 11. doi: 10.1038 / srep22878.63. Tomkinson AE, et al. DNA ligases as therapeutic targets. Transl Cancer Res.2013;2(3). PubMed PMID: 24224145.64. Srivastava M. et al. An inhibitor of nonhomologous end-joining abrogates doublestrand break repair and impedes cancer progression. Cell. 2012:151(7): 1474-87. doi: 10. 1016 / j. cell.2012.1 1.054.65. Zhong S, et al, et al. Identification and validation of human DNA ligase inhibitors using computer-aided drug design. J Med Chem. 2008;51(15):4553-62. Epub 2008 / 07 / 17. doi: 10.1021 / jm8001668.66. Grawunder, U., et al. DNA ligase IV binds to XRCC4 via a motif located between rather than within its BRCT domains. Curr Biol. 1998; 8(15):873-876; doi: 10.1016 / s0960-9822(07)00349- 167. Her, J., et al. Factors forming the BRCA1-A complex orchestrate BRCA1 recruitment to the sites of DNA damage. Acta Biochimica et Biophysica Sinica, 2016; 48(7): 658— 664; doi: 10.1093 / abbs / gmw047
[0279] Disclosed are materials, compositions, and components that can be used for, can be used in conjunction with, can be used in preparation for, or are products of the disclosed embodiments. These and other materials are disclosed herein, and it is understood that when combinations, subsets, interactions, groups, etc. of these materials are disclosed that while specific reference of each various individual and collective combinations and permutations of these compositions may not be explicitly disclosed, each is specifically contemplated and described herein. For example, if a method is disclosed and discussed and a number of modifications that can be made to a number of molecules included in the method are discussed, each and every combination and permutation of the method, and the modifications that are possible are specifically contemplated unless specifically indicated to the contrary. Likewise, any subset or combination of these is also specifically contemplated and disclosed. This concept applies to all aspects of this disclosure including, but not limited to, steps in methods using the disclosed compositions. Thus, if there are a variety of additional steps that can be performed, it is understood that each of these additional steps can be performed with any specific method steps or combination of method steps of the disclosed methods, and that eachsuch combination or subset of combinations is specifically contemplated and should be considered disclosed.
[0280] One skilled in the art will readily appreciate that the present disclosure is well adapted to carry out the objects and obtain the ends and advantages mentioned, as well as those inherent therein. The present disclosure described herein are presently representative of preferred embodiments, are exemplary, and are not intended as limitations on the scope of the present disclosure. Changes therein and other uses will occur to those skilled in the art which are encompassed within the spirit of the present disclosure as defined by the scope of the claims.
[0281] No admission is made that any reference, including any non-patent or patent document cited in this specification, constitutes prior art. In particular, it will be understood that, unless otherwise stated, reference to any document herein does not constitute an admission that any of these documents forms part of the common general knowledge in the art in the United States or in any other country. Any discussion of the references states what their authors assert, and the applicant reserves the right to challenge the accuracy and pertinence of any of the documents cited herein. All references cited herein are fully incorporated by reference, unless explicitly indicated otherwise. The present disclosure shall control in the event there are any disparities between any definitions and / or description found in the cited references.TABLES OF SEQUENCESTable 1: Sequences of ecDNA Reporter Construct Polynucleotide Constructs (n = a or c or g or t / u' unless otherwise specified)Table 2: Sequences of Oligonucleotides
Claims
WHAT IS CLAIMED IS:
1. A method of detecting circular DNA formation in a cell, the method comprising:(i) contacting a plurality of cells with a reporter polynucleotide,(ii) culturing the plurality of cells comprising the polynucleotide, and(iii) detecting a first signal from the first reporter polypeptide; wherein detection of the first signal indicates circular DNA formation in the plurality of cells.
2. The method of claim 1, wherein the reporter polynucleotide comprises a first nucleotide sequence encoding a first reporter polypeptide followed by a first promoter, wherein circularization of the linear polynucleotide leads to the first promoter being operably linked to the first nucleotide sequence.
3. The method of claim 2, wherein the first reporter polypeptide comprises a fluorescent protein.
4. The method of any one of claims 1-3. wherein the reporter polynucleotide comprises a nucleotide sequence of at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 1, wherein the nucleotide sequence comprises the first nucleotide sequencing encoding the first reporter polypeptide followed by the first promoter element.
5. The method of any one of claims 1-4, wherein the reporter polynucleotide further comprises a second promoter operably linked to a second nucleotide sequence encoding a second reporter polypeptide, wherein the second promoter and the second nucleotide sequence are positioned in the linear polynucleotide between the first nucleotide sequence and the first promoter.
6. The method of claim 5, wherein the first reporter polypeptide and the second reporter polypeptide are different proteins.
7. The method of claims 5 or 6, wherein the first promoter and the second promoter each comprise a eukary otic promoter nucleotide sequence.
8. The method of any one of claims 5-7, wherein the first promoter element and the second promoter element are different.
9. The method of any one of claims 5-8, further comprising detecting a second signal from the second reporter polypeptide, wherein detection of the second signal from the second reporter polypeptide indicates reporter polynucleotide presence in the plurality of cells.
10. The method of any one of claims 5-9, wherein the reporter polynucleotide further comprises a third nucleotide sequence encoding a selection marker, wherein the third nucleotide sequence encoding the selection marker is operably linked to the second promoter element.
11. The method of any one of claims 5-10, wherein the reporter polynucleotide comprises a nucleotide sequence comprising at least 80%. at least 85%. at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 2, wherein the nucleotide sequence, from its 5'-end to its 3'-end, comprises(i) the first nucleotide sequence encoding the first reporter polypeptide,(ii) the second promoter element,(iii) the second nucleotide sequence encoding the second reporter polypeptide,(iv) the third nucleotide sequence encoding the selection marker, and,(v) the first promoter element; wherein the second nucleotide sequence encoding the second reporter polypeptide and the third nucleotide sequence encoding the selection marker are operably linked to the second promoter element.
12. The method of claim 10 or 11, further comprising culturing the plurality7of cells with a compound that corresponds to a selectable marker.
13. The method of any one of claims 1-12, further comprising(i) culturing the plurality of cells w ith a compound of interest, and(ii) detecting a third signal from the first reporter polypeptide; whereinthe compound of interest decreases circular DNA presence if the third signal from the first reporter polypeptide is less than the first signal from the first reporter polypeptide.
14. The method of claims , wherein the plurality of cells is cultured with the compound of interest before the plurality of cells is contacted with the reporter polynucleotide.
15. The method of any one of claims 1-14, wherein the first signal from the first reporter polypeptide, second signal from the second reporter polypeptide, and / or third signal from the first reporter polypeptide are fluorescence signals.
16. A linear reporter polynucleotide comprising a first nucleotide sequence encoding a first reporter polypeptide followed by a first promoter, wherein circularization of the linear reporter polynucleotide leads to the first promoter being operably linked to the first nucleotide sequence.
17. The linear reporter polynucleotide of claim 16, wherein the first reporter polypeptide comprises a fluorescent protein.
18. The linear reporter polynucleotide of claim 17, comprising a nucleotide sequence of at least 80%. at least 85%. at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 1, wherein the nucleotide sequence comprises the first nucleotide sequencing encoding the first reporter polypeptide followed by the first promoter element.
19. The linear reporter polynucleotide of any one of claims 16-18, further comprising a second promoter operably linked to a second nucleotide sequence encoding a second reporter polypeptide, wherein the second promoter and the second nucleotide sequence are positioned in the linear polynucleotide between the first nucleotide sequence and the first promoter.
20. The linear reporter polynucleotide of claim 19, wherein the first reporter polypeptide and the second reporter polypeptide are different proteins.
21. The linear reporter polynucleotide of any one of claims 16-20, wherein the first promoter and the second promoter each comprise a eukaryotic promoter nucleotide sequence.
22. The linear reporter polynucleotide of any one of claims 16-21, wherein the first promoter element and the second promoter element are different.
23. The linear reporter polynucleotide of any one of claims 16-22, further comprising a third nucleotide sequence encoding a selection marker, wherein(i) the third nucleotide sequence encoding the selection marker lies between the second nucleotide sequence encoding the second reporter polypeptide and the first promoter element, and(ii) the third nucleotide sequence encoding the selection marker is operably linked to the second promoter element.
24. The linear reporter polynucleotide of any one of claims 16-23, comprising a nucleotide sequence comprising at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO:
2. wherein the nucleotide sequence, from its 5’-end to its 3’-end, comprises(i) the first nucleotide sequence encoding the first reporter polypeptide,(ii) the second promoter element,(iii) the second nucleotide sequence encoding the second reporter polypeptide.(iv) the third nucleotide sequence encoding the selection marker, and,(v) the first promoter element; wherein the second nucleotide sequence encoding the second reporter polypeptide and the third nucleotide sequence encoding the selection marker are operably linked to the second promoter element.
25. The linear reporter polynucleotide of any one of claims 16-24, further comprising a fourth nucleotide sequence encoding a cleavage site, wherein the fourth nucleotide sequence encoding the cleavage site lies between the second nucleotide sequence encoding the second reporter polypeptide and the third nucleotide sequence encoding the selection marker.
26. A method of determining whether a compound impairs circular DNA formation in a cell, the method comprising:(i) contacting a plurality of cells with the reporter polynucleotide of anyone of claims 16-24,(ii) culturing the plurality of cells comprising the polynucleotide,(iii) detecting a first signal from a first reporter polypeptide expressed from the polynucleotide,(iv) contacting the plurality of cells with a compound of interest,(v) detecting a second signal from the first reporter polypeptide, and(vi) comparing the second signal from the first reporter polypeptide to the first signal from the first reporter polypeptide; wherein the compound impairs circular DNA formation if the second signal is less than the first signal.
27. The method of claim 26, wherein the polynucleotide comprises the sequence as set forth in SEQ ID NO: 1 or SEQ ID NO: 2.
28. The method of claim 26 or 27, further comprising culturing the plurality of cells with a compound that corresponds to a selectable marker.
29. The method of any one of claims 26-28, further comprising detecting a third signal from a second reporter polypeptide expressed from the polynucleotide, wherein detection of the second signal from the second reporter polypeptide indicates polynucleotide presence in the plurality of cells.
30. The method of any one of claims 1-15 or 26-29, wherein the plurality of cells is a plurality of cancer cells.
31. The method of any one of claims 1-15 or 26-29, wherein the plurality of cells is a plurality of plant cells.
32. A kit for detecting circular DNA presence in a cell, the kit comprising a linear reporter polynucleotide of any one of claims 16-25 and instructions for use.
33. A method of inhibiting ecDNA biogenesis, comprising contacting a plurality of cells with an ecDNA inhibitor.
34. The method of claim 33, wherein the ecDNA inhibitor suppresses the activity of Lig4, XCCR4, or a protein of the BRCA1-A complex.
35. The method of claim 34, wherein the protein of the BRCA1-A complex is one of BABAM2, ABRAXAS 1, or Rap80.
36. The method of any one of claims 33-35, wherein the plurality of cells is in a subject having or suspected of having a cancer or a plant.
37. A method of screening for ecDNA formation inhibition in cells, comprising: delivering a reporter polynucleotide of any one of claims 16-25 to a plurality of cells, contacting a plurality of cells with an ecDNA inhibitor, thereby producing a population of treated cells; detecting an amount of ecDNA formation in the population of treated cells; and comparing the amount of ecDNA formation in the population of treated cells to a control population of untreated cells.
38. The method of claim 37, wherein the ecDNA inhibitor suppresses the activity of Lig4, XCCR4, or a protein of the BRCA1-A complex.
39. The method of claim 38, wherein the protein of the BRCA1-A complex is one of BABAM2, ABRAXAS 1, or Rap80.
40. The method of any one of claims 37-39, further comprising measuring the level of a reporter encoded by the reporter polynucleotide.
41. A composition comprising an inhibitor of ecDNA biogenesis and, optionally, a pharmaceutically acceptable excipient.
42. The composition of claim 41, wherein the inhibitor of ecDNA biogenesis suppresses the activity of Lig4, XCCR4, or a protein of the BRCA1-A complex.
43. The composition of claim 41 or 42, wherein the inhibitor of ecDNA biogenesis targets Lig4, XCCR4, or a protein of the BRCA1-A complex.
44. The composition of any one of claims 41-43. wherein the inhibitor of ecDNA is present in an effective amount to reduce the formation of ecDNA in a plurality of cells or a subject compared to a baseline level of ecDNA formation in the subject without the inhibitor.
45. The composition of any one of claims 41-44, further comprising a pharmaceutically acceptable carrier.
46. The composition of any one of claims 41-45, wherein the inhibitor is present in a unit dose formulation.
Citation Information
Patent Citations
Transposition-mediated identification of specific binding or functional proteins
US20150152406A1
Mutant of adeno-associated virus (AAV) capsid protein
US20200002384A1
AP50 polymerases and uses thereof
US20240084373A1
In vitro peptide or protein expression library
US7416847B1