R-loop sensor

The RHINO sensor, with multiple hybrid-binding domains and a reporter domain, addresses the limitations of existing R-loop detection methods by providing sensitive and specific imaging of R-loops in living cells, aiding in the study of diseases related to R-loops.

JP2026509984APending Publication Date: 2026-03-26FUNDAÇÃO GIMM- GULBENKIAN INSTITUTE FOR MOLECULAR MEDICINE +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-01
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Current methods for detecting R-loops in living cells, such as using anti-DNA:RNA hybrid antibody S9.6 and catalytically inactivated RNase H1, suffer from poor antigen specificity and low image quality, limiting the ability to study the dynamics of R-loop formation and dissolution in various diseases.

Method used

A genetically encoded R-loop sensor (RHINO) comprising multiple hybrid-binding domains (HBDs) and a reporter domain is expressed in cells to specifically and sensitively detect and image R-loops, allowing for determination of their abundance, distribution, and dynamic behavior.

Benefits of technology

Enables accurate and sensitive detection and imaging of R-loops in living cells, facilitating the study of diseases associated with R-loops, including cancer and degenerative disorders.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026509984000001_ABST
    Figure 2026509984000001_ABST
Patent Text Reader

Abstract

The present invention relates to a method for detecting R-loops in cells, comprising culturing isolated cells expressing a fusion protein that binds to a DNA:RNA hybrid, comprising (a) two or more hybrid-binding domains (HBDs) and (b) a detectable reporter domain, such that the fusion protein binds to the R-loop in the cell. The distribution of the fusion protein within the cell is then determined. The accumulation of the fusion protein indicates R-loops in the cell. Methods for screening fusion proteins that bind to DNA:RNA hybrids, coding nucleic acids, and compounds that modulate R-loops are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an agent that selectively binds to R-loops in cells, and a method of using the same.

Background Art

[0002] R-loops are abundant non-standard nucleic acid structures composed of double-stranded DNA:RNA hybrids and displaced DNA strands. R-loops are normally formed during transcription when nascent RNA molecules hybridize with template DNA strands, preventing the non-template strand from annealing to the template strand.

[0003] The most well-described physiological role of R-loops relates to their effect on the transcription of protein-coding genes by RNA polymerase II. For example, R-loops can activate or suppress gene expression by controlling the recruitment of transcription factors and by promoting epigenetic and chromatin remodeling (Non-Patent Document 1). R-loops are also involved as modulators of genomic instability and DNA damage (Non-Patent Document 2). As a result, R-loops are thought to contribute to the pathogenesis of various cancers as well as degenerative diseases such as amyotrophic lateral sclerosis (ALS), frontotemporal dementia (FTD), polyglutamine-related disorders, spinocerebellar ataxia, Huntington's disease, Friedreich's ataxia, and Ehlers-Danlos syndrome (Non-Patent Document 3).

[0004] The anti-DNA:RNA hybrid antibody S9.6 is standardly used to identify R-loops. However, S9.6 is known to bind to double-stranded RNA (dsRNA) molecules and is not suitable for use in living cells (Non-Patent Documents 4 and 5).

[0005] Other commonly used methods for mapping and quantifying R-loops utilize the RNase H1 enzyme. RNase H1 degrades the RNA portion of the R-loop, thereby restoring the double-stranded structure of the two DNA strands (Non-Patent Literature 6). Its specificity for DNA:RNA hybrids is conferred by the hybrid-binding domain (HBD). Overexpression of RNase H1 is useful for studying R-loops, which play many roles in processes such as gene expression regulation, class switch recombination, and genomic damage and repair.

[0006] Catalytically inactivated RNase H1 has been used as an R-loop detection tool for imaging of fixed cells and immunoprecipitation-based assays (Non-Patent Documents 7 and 8). Similarly, single fluorescently labeled RNase H1 HBDs have also been used to detect R-loops in living cells (Non-Patent Document 9). However, as with S9.6, the use of catalytically inactivated RNase H1 or isolated RNase HBDs in imaging of living cells has proven problematic due to poor antigen specificity and low image quality (e.g., low signal-to-noise ratio).

[0007] Therefore, access to tools that can specifically and sensitively detect R-loops and measure the dynamics of R-loop formation and dissolution is advantageous in the study of various diseases, including cancer and degenerative diseases. [Prior art documents] [Non-patent literature]

[0008] [Non-Patent Document 1] Niehrs C. Nat Rev Mol Cell Biol. 2020 Mar;21(3):167-178 [Non-Patent Document 2] Marnef A. Nat Cell Biol. 2021 Apr;23(4):305-313 [Non-Patent Document 3] Khan ES. Genes. 2022; 13(12):2181 [Non-Patent Document 4] Bou-Nader C. Nature Communications. 2022 Mar;13(1):1641 [Non-Patent Document 5] Smolka JA. J Cell Biol. 2021 Jun 7;220(6):e202004079 [Non-Patent Document 6] Cerritelli SM. Methods Mol Biol. 2022;2528:91-114 [Non-Patent Document 7] Wang K. Sci Adv. 2021 Feb 17;7(8):eabe3516 [Non-Patent Document 8] Crossley MP. J Cell Biol. 2021 Sep 6;220(9):e202101092 [Non-Patent Document 9] Silva, S. Methods Mol. Biol. 2022 June 16 2528 [Overview of the project]

[0009] The inventors unexpectedly discovered that the expression of a fusion protein that binds to DNA:RNA hybrids, comprising two or more hybrid-binding domains (HBDs) and a reporter domain, makes it possible to detect and image R-loops in living cells. The expression of this fusion protein may enable specific and highly sensitive determination of, for example, the abundance, distribution, and / or dynamic behavior of R-loops in cells, and may be useful in R-loop mapping and other applications. [Means for solving the problem]

[0010] A first aspect of the present invention is a method for detecting R-loops in cells, (i) The fusion protein binds to the R loop in the cell, (a) Two or more hybrid binding domains (HBDs), (b) culturing an isolated cell expressing a fusion protein that binds to a DNA:RNA hybrid comprising a detectable reporter domain and and (ii) determining the distribution of the fusion protein within the cell; and comprising providing a method wherein accumulation of the fusion protein indicates R-loops in the cell.

[0011] A second aspect of the invention is (i) two or more hybrid binding domains (HBDs); and (ii) a reporter domain; and providing a fusion protein that binds to a DNA:RNA hybrid.

[0012] Preferred hybrid binding domains (HBDs) for the first and second aspects comprise HBDs derived from RNase enzymes such as human RNase H1 (i.e., RNase HBDs) and can include, for example, the amino acid sequence of SEQ ID NO: 1 or variants thereof.

[0013] A third aspect of the invention provides a nucleic acid encoding a fusion protein that binds to a DNA:RNA hybrid of the second aspect.

[0014] A fourth aspect of the invention provides a vector comprising the nucleic acid of the third aspect.

[0015] A fifth aspect of the invention provides an isolated cell expressing a fusion protein that binds to a DNA:RNA hybrid of the second aspect. The isolated cell can comprise the nucleic acid of the third aspect or the vector of the fourth aspect.

[0016] A sixth aspect of the invention provides a method of producing a cell expressing a fusion protein that binds to a DNA:RNA hybrid, e.g., an isolated cell of the fifth aspect, comprising introducing the nucleic acid of the third aspect or the vector of the fourth aspect into an isolated cell.

[0017] A seventh aspect of the present invention is a method for screening a compound that regulates R-loops, comprising: (i) culturing the isolated cell according to the fifth aspect in the presence of a test compound such that the fusion protein binds to the R-loop in the cell; (ii) determining the distribution of the fusion protein in the cell; and providing a method wherein a change in the distribution of the fusion protein in the cell in the presence compared to the absence of the test compound indicates that the compound regulates the R-loop.

[0018] Aspects and embodiments of the present invention will be described in more detail below. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] [Figure 1]This figure shows images of the R-loop in living cells. (A) Illustration of the RHINO structure (left) and DNA:RNA hybrid labeling in the R-loop on the back of the RNA Pol II complex during transcription (right). The positions of the substituted amino acids in the RHINO mutant are indicated by red (W43A) and blue (KK59AA) asterisks. (B) Confocal microscopy image of RHINO expressed in U2OS cells with mCherry-H2B for marking the cell nucleus and the corresponding maximum intensity projection. Shows the RHINO spot throughout the nucleoplasm (highlighted in the inset) and the strongly labeled foci in the nucleolus. (C) Comparison of RHINO and S9.6 antibody labeling of DNA:RNA hybrids. The nucleolar signals of S9.6 and RHINO are mutually exclusive (inset). (D) Expression of two RHINO mutant constructs and mCherry-histone-H2B. Nucleoli are defined by dashed lines. (E) Maximum intensity projection and quantification of RHINO foci under each condition for live U2OS cells expressing RHINO alone (left) or live U2OS cells co-expressing mScarlet-i tagged human RNase H1 (right). (F) Quantification of RHINO foci when using rapidly inducible RNase H1. Monoconfocal plane of HeLa cells co-expressing RHINO and inducible RNase H1 before induction (top). Bottom, the same cells 1 hour after RNase H1 induction (+TA). Quantification of RHINO foci before and 1 hour after TA addition. (G) Maximum intensity projection and quantification of the number of RHINO foci of random live U2OS cells expressing RHINO before incubation with tryptride and after 1 hour of incubation. (H) Maximum intensity projection and quantification of the number of RHINO foci of the same live U2OS cells expressing RHINO before incubation with tryptride and after 1 hour of incubation. [Figure 2]This figure shows the dynamic behavior of the R-loop in nuclear substructures. (A) Labeling of the R-loop of nucleoli in live U2OS cells. Nucleoli are labeled with a red fluorescent nucleolar marker. Bright RHINO foci in several nucleoli are indicated by arrows. (B) Elucidation of R-loop localization in nucleolar substructures. RHINO-expressing U2OS cells were fixed and stained with an anti-UBF antibody that marks the fibrous center / high-density fibrous component interface. Presentation of two sets of super-resolution microscopy images acquired with structured illumination (top) and Airyscan 2 technology (bottom). (C) Measurement of the dynamic behavior of the R-loop of nucleoli by FRAP at the RHINO foci inside the nucleolus. Fluorescence within the red circle was decolorized (false color image), and recovery was monitored over time. The curves show the mean intensity (±SEM) of scattered GFP at the RHINO foci and as a control. (D) Labeling of the Tel-R-loop in live U2OS cells co-expressing RHINO and iRFP-TRF2. Some telomeres show labeling directly adjacent to or overlapping with the R-loop (enlarged details 1-5). Region 1 in the orthogonal 3D view shows colocalization of the R-loop and telomere labeling. (E) Measurement of Tel-R-loop dynamic behavior by FRAP of wild-type RHINO compared with both mutants and scattered GFP in labeled telomeres. The pseudocolor image shows wild-type RHINO-labeled cells with decolorized regions (red circles). Curves show the mean intensity (±SD) of the RHINO focus, RHINO mutant, and scattered GFP as control. [Figure 3]This figure shows images of R-loops in actively transcribed genes in living cells. (A) Illustration of a β-globin reporter gene for live cell transcription imaging using the MS2 system. (B) Colocalization of mCherry-rtTA and MS2-GFP signals upon doxycycline addition in living U2OS reporter gene cells. The merged image and intensity line scan show the combined fluorescence signals. (C) Labeling of R-loops with RHINO in living U2OS cells in an actively transcribed β-globin reporter gene array. Fluorescent colocalization is characterized in the merged image and intensity line scan in the inset. (D) Monoconfocal plane of RHINO in a 3D mCherry-rtTA-labeled β-globin reporter gene array over time. The graph on the right shows the RHINO fluorescence intensity over time at the reporter gene locus. (E) Illustration of an IgM reporter gene containing an R-loop forming sequence (RFS) integrated as a single copy in the U2OS host cell line genome. Labeling of a single reporter gene locus with mCherry-rtTA parallel to GFP-PP7, directly visualizing transcriptional activity in merged images and enlarged insets using intensity line scans. (F) Colocalization of RHINO and mCherry-rtTA at the site of the transcriptional activity reporter gene locus. [Modes for carrying out the invention]

[0020] This invention relates to the development of a genetically encoded R-loop sensor (hereinafter referred to herein as "RHINO" or a sensor that binds to RNA:DNA hybrids). The R-loop sensor is a fusion protein that selectively binds to RNA:DNA hybrids and comprises multiple hybrid-binding domains (HBDs) and a reporter domain. Expression of the sensor in living cells enables imaging of R-loops within living cells and the specific and highly sensitive determination of R-loop abundance, distribution, and dynamic behavior.

[0021] DNA:RNA hybrids are double-stranded nuclear molecules containing a DNA strand and a complementary RNA strand. DNA-RNA hybrids arise as transient intermediates in numerous physiological processes. In particular, DNA:RNA hybrids reside in the R-loop within the cell nucleus.

[0022] The R-loop is a non-standard triple-stranded nucleic acid structure containing a DNA:RNA hybrid and a displaced DNA strand. R-loops are typically transiently formed in cells during transcription when a nascent RNA molecule hybridizes with a template DNA strand. However, stable R-loops, associated with gene damage and instability, can reside in the cell nucleus, e.g., the nucleolus or nucleoplasm. R-loops can also occur in mitochondria within cells.

[0023] R-loop sensors expressed by cells can bind to the R-loop in the cell nucleus, for example, in the nucleolus or nucleoplasm. Preferably, the R-loop sensor selectively binds to the R-loop. For example, the R-loop sensor may selectively bind to the DNA:RNA hybrid structure of the R-loop, but may not be able to bind to other DNA:RNA hybrid structures (e.g., primers for Okazaki fragments), or may not be able to bind substantially to them.

[0024] The binding of an R-loop sensor to an R-loop leads to the accumulation or enrichment of the sensor at the R-loop site within the cell. This accumulation or enrichment may be detected, for example, as a separate focus above a nonspecific scattering background. For example, detecting the distribution of expressed sensors within a cell by imaging allows for the detection of accumulation or enrichment and the identification of the R-loop site. In imaging assays, the accumulation or enrichment of an R-loop sensor may be identified as a signal focus (e.g., a fluorescence focus). A signal focus may indicate, for example, an R-loop, a population of R-loops (e.g., formed at multiple locations within the same gene), or the presence of individual R-loops large enough to facilitate the binding of two or more R-loop sensors. In some embodiments, a signal focus is identified as a region having a signal intensity (e.g., mean fluorescence intensity, MFI) at least 12%–15% higher than the background, as determined by confocal microscopy.

[0025] The R-loop sensor is a fusion protein expressed within a cell. A fusion protein is a recombinant protein that is neither found in nature nor expressed in cells; that is, the fusion protein and its encoding nucleic acid are heterologous to the cell. The term "heterologous" refers to a polypeptide or nucleic acid that is foreign to a particular biological system, such as a host cell, and does not exist naturally in that system. Heterologous polypeptides or nucleic acids can be introduced into a biological system by artificial means, for example, using recombinant technology. For example, a heterologous nucleic acid encoding a polypeptide may be inserted into a suitable expression construct and then used to transform a host cell to produce the polypeptide. Heterologous polypeptides or nucleic acids may be synthetic or artificial, or they may exist in different biological systems, such as different species or cell types.

[0026] The fusion protein may contain two or more, three or more, four or more, five or more, or six or more, seven or more, eight or more, or nine or more hybrid binding domains (HBDs). Preferably, the fusion protein contains three to nine HBDs, for example, three, four, five, six, seven, eight or nine HBDs.

[0027] HBDs are protein domains that bind to DNA:RNA hybrids. DNA:RNA hybrid binding can be conflated by protein structures such as alpha-beta plaits, P-loop triphosphate hydrolase domains, nucleic acid-binding OB folds, DNA:RNA helicase domains, DEAD / DEAH box-type domains, or K-homologous domains (Wang, I. Genome Res. 2018 August, 14 28:1405-1414). HBDs suitable for use in R-loop sensors may be derived from proteins that bind to any DNA:RNA hybrid containing HBDs. Suitable proteins that bind to DNA:RNA hybrids are known in the art (see, for example, Wang et al 2018 Genome Res. Genome Res. 2018 August, 14 28:1405-1414) and may include DDX5, FUS, HNRNPM, MATR3, NCL, NONO, PSPC1, RECQL, SFPQ, SRSF1, XRN1, XRN2, or RNase H1. Preferably, the HBD is an RNase hybrid binding domain (HBD), i.e., an HBD derived from an RNase enzyme such as human RNase H1. Suitable RNase HBDs are well known in the art and include HBDs of H1 RNase enzymes such as human RNase H1 (residues 1-50 of the mature sequence of human RNase H1 (Gene ID 246243, NP_001273763.1), Nowotny M. EMBO J. 2008 Apr 9;27(7):1172-81).

[0028] HBDs such as RNase HBDs may consist of 25 to 100 amino acids, 30 to 75 amino acids, 40 to 60 amino acids, or 45 to 55 amino acids, preferably about 50 amino acids. In some preferred embodiments, the HBD may be an RNase HBD containing or consisting of one of the amino acid sequences of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 4, or a variant thereof. A preferred HBD may be encoded by the nucleotide sequence of SEQ ID NO: 3 or SEQ ID NO: 5, or a variant thereof.

[0029] Variants of reference amino acid sequences described herein, such as reference HBD sequences, reference linker sequences, reference reporter sequences, or reference fusion protein sequences, may include amino acid sequences having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 98% sequence identity with respect to the reference sequence. A particular amino acid sequence variant may differ from the reference sequence described herein by one, two, three, four, five, six, seven, eight, nine, or more amino acid insertions, additions, substitutions, or deletions.

[0030] Preferably, the RNase HBD may contain an amino acid sequence having a Tyr residue at the position corresponding to Trp16 in SEQ ID NO: 1, a Lys residue at the position corresponding to Lys32 in SEQ ID NO: 1, and a Lys residue at the position corresponding to Lys33 in SEQ ID NO: 1. RNase HBD may further include an amino acid sequence having a Phe residue at the position corresponding to Phe1 in SEQ ID NO: 1, a Tyr residue at the position corresponding to Tyr2 in SEQ ID NO: 1, a Val residue at the position corresponding to Val4 in SEQ ID NO: 1, an Arg residue at the position corresponding to Arg5 in SEQ ID NO: 1, an Arg residue at the position corresponding to Arg8 in SEQ ID NO: 1, a Phe residue at the position corresponding to Phe13 in SEQ ID NO: 1, a Cys residue at the position corresponding to Cys19 in SEQ ID NO: 1, a Val residue at the position corresponding to Val23 in SEQ ID NO: 1, a Phe residue at the position corresponding to Phe31 in SEQ ID NO: 1, a Phe residue at the position corresponding to Phe34 in SEQ ID NO: 1, and / or a Phe residue at the position corresponding to Phe43 in SEQ ID NO: 1 (Nowotny et al. The EMBO Journal (2008) 27, 1172-1181).

[0031] Sequence similarity and identity are generally defined by referring to the GAP algorithm (Wisconsin Package, Accelerys, San Diego, USA). GAP uses the Needleman-Unsch algorithm to align two complete sequences, maximizing the number of matches and minimizing the number of gaps. Typically, default parameters are used, with a gap generation penalty of 12 and a gap extension penalty of 4. While the use of GAP may be preferred, other algorithms that generally use default parameters may also be used, such as BLAST (using the method of Altschul SF. J Mol Biol. 1990 Oct 5;215(3):403-10), FASTA (using the method of Pearson WR and Lipman DJ. Proc Natl Acad Sci US A. 1988 Apr;85(8):2444-8), or the Smith-Waterman algorithm (Smith TF, Waterman MS. J Mol Biol. 1981 Mar 25;147(1):195-7), or the TBLASTN program of Altschul et al. (1990) mentioned above. In particular, the psi-Blast algorithm (Altschul SF. Nucleic Acids Res. 1997 Sep 1;25(17):3389-402) may be used.

[0032] Sequence comparisons may be performed over the entire length of the relevant sequences described herein.

[0033] In some preferred embodiments, the fusion protein may comprise three RNase HBDs. For example, the fusion protein may comprise a first RNase HBD, a second RNase HBD, and a third RNase HBD. The first, second, and third HBDs each comprise or consist of the amino acid sequence of SEQ ID NO: 1 or a variant thereof. In some embodiments, the second and third HBDs may be modified to remove an internal translation initiation site that can lead to the production of a cleaved fusion protein. For example, the first HBD comprises or consists of the amino acid sequence of SEQ ID NO: 3 or a variant thereof and may be encoded by the nucleotide sequence of SEQ ID NO: 4 or a variant thereof. The second and third HBDs comprise or consist of the amino acid sequence of SEQ ID NO: 5 or a variant thereof and may be encoded by the nucleotide sequence of SEQ ID NO: 6 or a variant thereof. Preferably, the first, second, and third RNase HBDs are arranged consecutively in the fusion protein in an N-terminal to C-terminal orientation.

[0034] A detectable reporter domain is a protein domain that produces a detectable signal. Preferably, the detectable signal is light. The signal produced by the reporter allows for the determination of the distribution of fusion proteins within a cell and enables the identification of enrichment or accumulation exhibiting an R-loop.

[0035] A light-generating reporter allows the distribution of fusion proteins within cells to be determined by standard live-cell imaging techniques such as fluorescence microscopy. Suitable reporter domains include fluorescent proteins and bioluminescent proteins.

[0036] Suitable fluorescent proteins are well known in the art and include green fluorescent protein (GFP) and its derivatives, e.g., mGreenLantern, mNeonGreen, DsRed, and its derivatives, e.g., mCherry and tdTomato; flavin mononucleotide-binding fluorescent protein (FbFP) and its derivatives; small ultrared fluorescent protein (smURFP) and its derivatives; mRuby and its derivatives; TagRFP and its derivatives; and synthetic fluorescent proteins, e.g., mScarlet and its derivatives. In a preferred embodiment, the fluorescent protein may be mNeonGreen. A suitable fluorescent protein may contain or consist of the amino acid sequence of SEQ ID NO: 7 or a variant thereof, and may be encoded by the nucleotide sequence of SEQ ID NO: 8 or a variant thereof.

[0037] Suitable bioluminescent proteins are well known in the art and include firefly luciferases such as Photinus pyralis luciferase, sea urchin luciferases such as Renilla reniformis luciferase, and bioluminescent proteins such as aequorin. Other suitable bioluminescent proteins include NanoLuc and its derivatives (Hall, M. ACS Chem. Biol. 2012, 7, 11, 1848-1857, Suzuki, K. Nat Commun. 2016 7, 13718).

[0038] The RNase HBD domain and reporter domain may be linked directly or independently via a linker in the fusion proteins described herein. In some embodiments, the linker may introduce charged amino acids and polar amino acids, thereby increasing the solubility of the fusion protein by preventing the formation of protein aggregates and increasing the bioavailability of the fusion protein in living cells. In addition, the linker may improve the flexibility of the fusion protein, thereby increasing its binding to the R-loop (by promoting the optimal alignment of the DNA:RNA hybrid of the HBD and R-loop).

[0039] A suitable linker may consist of 1 to 60 amino acids, for example, 1 to 5 amino acids, 5 to 10 amino acids, 10 to 15 amino acids, 15 to 20 amino acids, 20 to 25 amino acids, 25 to 30 amino acids, 30 to 35 amino acids, 35 to 40 amino acids, 45 to 50 amino acids, 50 to 55 amino acids, or 55 to 60 amino acids. Preferably, the linker consists of 1 to 30 amino acids, for example, 5 to 25 amino acids or 5 to 15 amino acids.

[0040] Two or more RNase HBDs in the fusion protein can be linked facing each other, or they can be linked by a linker located between the HBDs. A suitable linker may include or consist of the amino acid sequence of SEQ ID NO: 9 or a variant thereof, and may be encoded by the nucleotide sequence of SEQ ID NO: 10 or a variant thereof.

[0041] The RNase HBD and reporter domain can be linked facing each other, or linked by a linker located between the HBD and the reporter domain. In some embodiments, the linker between the HBD and the reporter domain is longer than the linker between the HBDs, for example, 5 to 15 amino acids longer, to minimize interference between the reporter domain (e.g., a fluorescent protein) and the HBD, and thus promote improved hybrid binding and signal identification. A suitable linker may include or consist of the amino acid sequence of SEQ ID NO: 11 or a variant thereof, and may be encoded by the nucleotide sequence of SEQ ID NO: 12 or a variant thereof.

[0042] In some preferred embodiments, the fusion protein that binds to the DNA:RNA hybrid suitable for use as an R-loop sensor described herein may include or consist of the amino acid sequence of SEQ ID NO: 13 or a variant thereof.

[0043] A nucleic acid is provided that encodes a fusion protein that binds to a DNA:RNA hybrid suitable for use as an R-loop sensor. The nucleic acid may encode a fusion protein that binds to the above-mentioned DNA:RNA hybrid. For example, the nucleic acid encoding a fusion protein that binds to a DNA:RNA hybrid may contain or consist of the nucleotide sequence of SEQ ID NO: 14 or a variant thereof. In some embodiments, the nucleic acid encoding the fusion protein may include one or more specific sequence modifications to the N-terminus of the HBD. For example, the nucleic acid may encode an additional Val residue or an additional Met residue prior to the HBD sequence (e.g., between the linker and the HBD sequence). This may, for example, exclude or substantially reduce translation from the internal start site that leads to the production of the cleaved fusion protein.

[0044] Variants of reference nucleotide sequences described herein, such as reference fusion protein coding sequences, may include nucleotide sequences having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 98% sequence identity with respect to the reference sequence. A particular nucleotide sequence variant may differ from the reference sequence described herein by one, two, three, four, five, six, seven, eight, nine, or more nucleotide insertions, additions, substitutions, or deletions.

[0045] Nucleic acid molecules may include DNA and / or RNA and may be partially or completely synthesized. Unless otherwise required by context, references to nucleotide sequences described herein encompass DNA molecules having the specified sequence and RNA molecules having the specified sequence in which T is substituted with U.

[0046] The nucleic acids may be codons optimized for expression in mammalian cells.

[0047] The nucleic acid encoding the fusion protein that binds to the DNA:RNA hybrid may be operably ligated to a promoter. The promoter may be a constitutive promoter, or more preferably an inductive promoter. Preferred constitutive promoters include the human ubiquitin C promoter and the mouse ribosomal protein L30 promoter. Preferred inductive promoters include chemically inductive promoters such as tetracycline-regulated transactivators, lac operons, or lac repressors, temperature-inductive promoters such as heat shock protein-derived promoters, and optogenetically activated promoters. In preferred embodiments, the inductive promoter may be an inductive and regulated promoter system such as the Teton3G (Registered) system or a Cumate switch system.

[0048] Suitable methods for introducing and expressing heterologous nucleic acids into cells are well known in the art and will be described in more detail below.

[0049] In some embodiments, nucleic acids encoding the fusion proteins described herein can be directly introduced into cells using gene editing techniques.

[0050] In other embodiments, the nucleic acids encoding the fusion proteins described herein may be incorporated into an expression vector. Suitable vectors are well known in the art and are described in more detail herein. Suitable vectors containing appropriate regulatory sequences, including promoter sequences, terminator fragments, polyadenylation sequences, enhancer sequences, marker genes, and other sequences, may be selected or constructed as needed. Suitable vectors may, for example, include, in addition to the coding sequence, a Kozak sequence, a 3'UTR, a cleavage site, a polyadenylation signal, and / or a 5'UTR. Inducible vectors may further include co-expressed transcriptional activators that bind to the promoter in response to inducing signals, such as the addition of specific promoter elements and compounds.

[0051] Preferably, the vector contains appropriate regulatory sequences for driving nucleic acid expression in mammalian cells. The vector may also contain sequences such as origins of replication, promoter regions, and selectable markers that enable its selection, expression, and replication in a bacterial host such as E. coli.

[0052] The vector may be a plasmid, such as a phagemid, or a virus, such as a phage, as needed. The vector may also be a DNA plasmid. Suitable viral vectors may include retroviral vectors, lentiviral vectors, adenovirus vectors, adeno-associated virus vectors, herpes simplex virus vectors, and chimeric virus vectors. Preferably, the nucleic acid construct is contained in the viral vector, most preferably in a lentiviral vector such as a gamma retroviral vector or a VSVg pseudotyped lentiviral vector.

[0053] Cells may be transduced by contact with viral particles containing nucleic acids. Viral particles for transduction may be produced by known methods. For example, HEK293T cells may be transduced using a lentiviral vector containing plasmids encoding viral packaging and envelope elements, as well as encoding nucleic acids. VSVg pseudotyped viral vectors can be produced in combination with viral envelope glycoprotein G of vesicular stomatitis virus (VSVg) to produce pseudotyped viral particles. For further details, see, for example, Sambrook & Russell (2001) Molecular Cloning: a Laboratory Manual: 3rd edition, Cold Spring Harbor Laboratory Press.

[0054] For example, many known techniques and protocols for the preparation of nucleic acid constructs, mutagenesis, sequencing, DNA introduction into cells, and manipulation of nucleic acids in gene expression, as well as protein analysis, are described in Ausubel et al. (1999) 4 th This is described in detail in the eds., *Short Protocols in Molecular Biology: A Compendium of Methods from Current Protocols in Molecular Biology*, John Wiley & Sons.

[0055] Isolated cells expressing fusion proteins that bind to the DNA:RNA hybrids described herein are also provided. Preferred cells may include nucleic acids or vectors encoding the fusion proteins that bind to the DNA:RNA hybrids described above.

[0056] Isolated cells can transiently or stably express fusion proteins that bind to DNA:RNA hybrids.

[0057] The isolated cells may be mammalian cells such as human cells or rodent cells, or non-mammalian cells such as bird cells, insect cells such as fruit fly cells, nematode cells such as C. elegans cells, fish cells such as zebrafish (Danio rerio) cells, reptile cells, or amphibian cells.

[0058] Suitable mammalian cells may be derived from any tissue at any stage of differentiation and may include stem cells such as hematopoietic stem cells, mesenchymal stem cells, and epithelial stem cells; bone cells such as osteoclasts, osteoblasts, and osteocytes; blood cells such as lymphocytes, neutrophils, basophils, and megakaryocytes; muscle cells such as myocytes, cardiomyocytes, and smooth muscle cells; adipocytes and other adipocytes; and brain or nerve cells such as neurons and glial cells.

[0059] Suitable isolated cells may be derived from cultured cell lines or obtained directly from the subject.

[0060] In some embodiments, the isolated cells described herein may exhibit a disease phenotype. For example, the isolated cells may exhibit a cancer phenotype (i.e., the cells may be cancer cells) or a neurodegenerative phenotype. Suitable disease cells include those derived from cultured cell lines or cells obtained directly from subjects with the disease. In some embodiments, cancer may be HCC, non-Hodgkin lymphoma, Burkitt lymphoma, multiple myeloma, chronic lymphocytic leukemia, hairy cell leukemia, prolymphocytic leukemia, anal cancer, appendiceal cancer, cholangiocarcinoma (i.e., cholangiocarcinoma), bladder cancer, brain tumor, breast cancer, cervical cancer, colon cancer, cancer of unknown primary origin (CUP), esophageal cancer, eye cancer, fallopian tube cancer, gastrointestinal cancer, kidney cancer, liver cancer, lung cancer, medulloblastoma, melanoma, oral cancer, ovarian cancer, pancreatic cancer, parathyroid disease, penile cancer, pituitary tumor, prostate cancer, rectal cancer, skin cancer, gastric cancer, testicular cancer, pharyngeal cancer, thyroid cancer, uterine cancer, vaginal cancer, or vulvar cancer.

[0061] Isolated cells expressing fusion proteins that bind to DNA:RNA hybrids as described herein may be produced by a method comprising introducing nucleic acids or vectors encoding fusion proteins that bind to DNA:RNA hybrids as described herein into isolated cells.

[0062] Nucleic acids or vectors may be introduced into isolated cells by any suitable method. Suitable methods for introducing or incorporating nucleic acids into isolated cells are well known to those skilled in the art. The nucleic acid to be inserted may be assembled in a construct or vector containing effective regulatory elements that drive transcription in the isolated cells. Suitable techniques for transporting constructs or vectors into isolated cells are well known in the art and include calcium phosphate transfection, DEAE-dextran, electroporation, liposome-mediated transfection, and transduction using retroviruses or other viruses, e.g., vaccinia or lentiviruses. For example, solid-phase transduction may be performed without selection by culturing on retroviral vector-pretreated tissue culture plates coated with retronectin.

[0063] In some embodiments, the nucleic acid may be operably linked to an inductive promoter, and the method may further include inducing a cell to express a fusion protein that binds to a DNA:RNA hybrid such that the fusion protein binds to an R loop in the cell.

[0064] Methods for detecting R-loops and double-stranded DNA:RNA hybrids in cells are also provided. In some embodiments, methods for detecting R-loops in cells are: (i) Culture isolated cells expressing a fusion protein that binds to the DNA:RNA hybrid described herein so that the fusion protein binds to the R loop in the cell, (ii) To determine the distribution of the fusion protein expressed in the cell, It may include, The distribution of the expressed fusion protein reflects the distribution of R-loops within the cell.

[0065] Preferably, the isolated cells are living cells.

[0066] In some embodiments, the isolated cells may be located within an in vitro organoid.

[0067] In another embodiment, a method for detecting R-loops in cells is: (i) Expressing a fusion protein that binds to the DNA:RNA hybrid described herein in one or more cells of a non-human organism such that the fusion protein binds to the R loop in one or more cells, (ii) Determining the distribution of the fusion protein expressed in one or more cells, It may include, The distribution of the expressed fusion protein reflects the distribution of R-loops in one or more cells of non-human organisms.

[0068] Suitable non-human organisms include small or naturally occurring transparent / translucent organisms such as the nematode *Synorhabdis elegans* and the zebrafish (*Danio relio*).

[0069] The distribution of expressed fusion proteins within the nucleus of one or more cells may be determined, for example, within the nucleolus or nucleoplasm. The distribution of fusion proteins may also be determined by determining the amount of fusion protein present at different sites or locations within one or more cells. Because expressed fusion proteins bind to the R-loop, they may accumulate or concentrate at R-loop sites within one or more cells, forming separate foci; that is, larger amounts of fusion protein may be present at R-loop sites than at other sites. The presence of accumulation or concentration of expressed fusion proteins at a site in a cell indicates the presence of an R-loop at that site. Sites of accumulation or concentration of fusion proteins may indicate populations near the R-loop (e.g., formed at multiple locations within the same gene) or the presence of individual R-loops large enough to facilitate the binding of two or more fusion proteins.

[0070] The distribution of the fusion protein may be determined by detecting the reporter domain in one or more cells. For example, the reporter domain may produce a detectable signal, preferably light. Signals produced by the reporter domain at different sites within the cell may be detected. The intensity of the signal at a site within the cell may indicate the amount of fusion protein at that site. The intensity of the signal throughout the cell may indicate the distribution of the fusion protein within the cell. An increased signal intensity at a site compared to other sites, e.g., a site known not to contain an R-loop, indicates the accumulation or enrichment of the fusion protein at the R-loop site. The amount of fusion protein at one or more sites may be increased compared to the background amount of fusion protein in the cell. For example, a site of fusion protein accumulation or enrichment may be identified as a region having a signal intensity (e.g., mean fluorescence intensity, MFI) at least 12% to 15% higher than the background signal intensity.

[0071] In some embodiments, the detectable signal may be light. The distribution of fusion proteins in cells can be determined, for example, by imaging the cells, by detecting the light produced by the reporter domain. Suitable imaging techniques such as fluorescence microscopy, chemiluminescence / bioluminescence, and immunoelectron microscopy are well established in the art (Suzuki, K. Nat Commun. 2016 7, 13718, Ariotti et al., Dev Cell. 2015 Nov 23;35(4):513-25). The distribution of fusion proteins may be determined by standard image analysis techniques, for example, using automated tools or software. For example, signal foci indicating enrichment or accumulation of fusion proteins in cells can be detected and counted, and their intensity, size, and / or shape can be preferably measured in three dimensions.

[0072] The presence, number, location, duration, and dynamic behavior of R-loops within a cell can be determined from the distribution of fusion proteins within the cell. For example, the presence, number, and / or location of R-loops within the cell or nucleus may be determined.

[0073] The total number of R-loop sensors concentrated or accumulated, indicated by the foci of detectable signals, may indicate the proliferation of R-loops in a given cell type or tissue compared to a control, or in a treated cell type or tissue. An increase in the number of signal foci may indicate an increase in the number of R-loops. The intensity of signal foci compared to the background may indicate the number or size of R-loops at a given location within the cell compared to the background.

[0074] The distribution of fusion proteins within a cell may be determined at a single time point, at multiple time points over a period of time, or continuously. In some preferred embodiments, the distribution of fusion proteins within a cell may be determined in real time in a living cell. Changes in the distribution of fusion proteins over a period of time may indicate changes in R-loops within the cell. For example, the presence, number, location, duration, and / or dynamic behavior of R-loops within a cell may change over a period of time. In some embodiments, the rates of R-loop formation and dissolution may be determined within the cell.

[0075] In some embodiments, the method for detecting R-loops in cells described above may further include preparing isolated cells expressing a fusion protein that binds to a DNA:RNA hybrid. Suitable isolated cells can be prepared by introducing the nucleic acids or vectors described herein into the isolated cells described herein.

[0076] The fusion proteins of the present invention are also suitable for use in nucleic acid blotting (so-called "dot blot" techniques) and immunoprecipitation assays, each of which is well established in the art.

[0077] The methods described herein may be useful, for example, for screening compounds that modulate the presence, number, location, duration, and / or dynamic behavior of DNA:RNA hybrids or R-loops in cells. A method for screening compounds that modulate R-loops in cells is: (i) Culturing isolated cells expressing a fusion protein that binds to the DNA:RNA hybrid described herein in the presence of the test compound, (ii) To determine the distribution of the fusion protein expressed in the cell, It may include, The change in the distribution of the fusion protein in the presence of the test compound compared to the absence of the test compound indicates that the compound regulates the R-loop in cells.

[0078] In another embodiment, a method for screening compounds that regulate the R-loop in non-human organisms is: (i) Expressing a fusion protein that binds to the DNA:RNA hybrid described herein in one or more cells of a non-human organism treated with the test compound so that the fusion protein binds to the R loop in one or more cells, (ii) Determining the distribution of the fusion protein expressed in one or more cells, It may include, The changes in the distribution of fusion proteins in one or more cells in treated non-human organisms compared to untreated non-human organisms indicate that the compound modulates the R-loop in non-human organisms.

[0079] Compounds that modulate R-loops can alter the presence, number, location, duration, and / or dynamic behavior of R-loops within one or more cells. For example, a compound may increase or decrease the number of R-loops in a cell, alter the location of R-loops in a cell, and / or increase or decrease the duration of R-loops in a cell. Such compounds may satisfy clinical or experimental utility as novel therapeutic agents, for example, as active agents for use in the treatment of cancer or degenerative diseases. Therefore, the screening methods for test compounds disclosed herein may be useful and advantageous in identifying novel therapeutic agents.

[0080] An increase in the amount of fusion protein concentrated or accumulated at one or more sites in cells, or an increase in the number of sites where the fusion protein is concentrated or accumulated in one or more cells, may indicate that the test compound promotes or stabilizes the R loop. A decrease in the amount of fusion protein concentrated or accumulated at one or more sites in cells, or a decrease in the number of sites where the fusion protein is concentrated or accumulated, may indicate that the test compound reduces or destabilizes the R loop.

[0081] The exact form of either the screening or assay method of the present invention may be modified by those skilled in the art using routine skills and knowledge. Those skilled in the art will fully understand the need to utilize appropriate control experiments.

[0082] The test compound may be an isolated molecule or may be contained in a sample, mixture, or extract, such as a biological sample. Compounds that can be screened using the methods described herein may be natural or synthetic compounds used in drug screening programs. Extracts of plants, microorganisms, or other organisms containing several characterized or uncharacterized components may also be used. In some embodiments, the test compound may be a pharmaceutical agent, such as a chemotherapeutic agent.

[0083] Suitable test compounds also include analogues, derivatives, variants, and mimics of any of the compounds listed above, for example, compounds produced using relevant drug designs to yield a candidate test compound having specific molecular shape, size, and charge characteristics suitable for modulating the R-loop in cells.

[0084] Combinatorial library technology provides an efficient method for testing the ability of a potentially vast number of different compounds to modulate the R-loop in cells. Such libraries and their uses are known in the art, among other things, for all kinds of natural products, small molecules, and peptides. The use of peptide libraries may be preferable under certain circumstances.

[0085] The amount of test compound that can be added to the assay of the present invention is usually determined by trial and error depending on the type of compound used. Typically, putative inhibitor compounds may be used at concentrations of about 0.001 nM to 1 mM or higher, for example, 0.01 nM to 100 μM, for example 0.1 μM to 50 μM, for example about 10 μM. Even compounds with weak effects may be useful lead compounds for further investigation and development.

[0086] Test compounds identified as regulating the R-loop in cells may be further investigated.

[0087] The test compound identified as an R-loop modulator may be isolated and / or purified, or it may be synthesized using conventional recombinant expression techniques or chemical synthesis techniques. Furthermore, the test compound may be manufactured and / or used in the preparation of compositions such as pharmaceuticals, pharmaceutical compositions or drugs, i.e., in manufacturing or formulation. Accordingly, the methods described herein may include formulating the test compound in a pharmaceutical composition with pharmaceutically acceptable excipients, vehicles or carriers for therapeutic use.

[0088] Other aspects and embodiments of the present invention provide the aspects and embodiments described above with the terms "comprising" replaced by the term "consisting of", and the aspects and embodiments described above with the terms "comprising" replaced by the term "consisting essentially of".

[0089] It is understood that this application discloses all aspects and all combinations of the embodiments described above, unless the context should be interpreted otherwise. Similarly, this application discloses all combinations of preferred and / or any features, either individually or with any other aspects, unless the context should be interpreted otherwise.

[0090] Modifications of the above embodiments, further embodiments, and variations thereof will become apparent to those skilled in the art by reading this disclosure, and they themselves fall within the scope of the present invention.

[0091] All references and sequence database entries mentioned herein are incorporated herein by reference in their entirety for all purposes.

[0092] As used herein, “and / or” shall be understood as a specific disclosure that includes or excludes each of two expressed features or components. For example, “A and / or B” shall be understood as if each of (i) A, (ii) B, and (iii) A and B were described separately herein.

[0093] Herein, certain aspects and embodiments of the present invention will be described as examples and with reference to the drawings shown above. [Examples]

[0094] Experiment 1: Imaging of R-loops in living cells RHINO is a genetically encoded RNA:DNA hybrid-binding sensor that detects the R-loop in vivo, allowing for novel insights into the dynamic behavior of the R-loop at different genomic locations. RHINO is a fusion protein containing three HBDs and a reporter domain (e.g., a fluorescent protein) that enables real-time imaging of R-loop foci in living cells (Figure 1A). Imaging of living human osteosarcoma (U2OS) cells transiently expressing RHINO and the fluorescent histone H2B protein (mCherry-H2B) revealed nuclear-only staining, with clearly defined nucleoli and scattered foci of heterogeneous size across a low scattering background (inset) (Figure 1B). Staining of U2OS cells expressing RHINO using antibody S9.6 revealed very strong cytoplasmic staining, which has been described as primarily resulting from dsRNA binding (Non-Patent Literature 8). The difference in the RHINO and S9.6 staining profiles highlights the superior specificity of RHINO, which colocalizes with the S9.6 focus only in the nucleus (Figure 1C).

[0095] Next, to investigate the HBD structure necessary for DNA:RNA binding, we generated two RHINO mutants with either the W43A or KK59AA amino acid substitution on the HBD (Nowotny M. EMBO J. 2008 Apr 9;27(7):1172-81). By imaging U2OS cells expressing each of the two mutant sensors, we revealed nonspecific scattering staining covering the entire nucleus, which in the case of the W43A mutant included the nucleolus (Figure 1D). The absence of separate foci formed by either of the mutants suggests that wild-type (wt) RHINO specifically accumulates at sites containing RNA:DNA hybrids.

[0096] To further evaluate the specificity of RHINO, RNase H1 was overexpressed in U2OS cells. This supported the specific binding of RHINO to RNA:DNA hybrids, and a 50% reduction in the number of R-loop foci was observed in cells overexpressing RNase H1 (Figure 1E). Similar results were obtained with rapid digestion of the R-loop by glucocorticoid receptor (GR)-fused RNase H1 (RNase H1-GR) (Figure 1F). Addition of the GR ligand triamcinolone acetonide (TA) to cells expressing RNase H1-GR drove the translocation of RNase H1-GR to the nucleus, significantly reducing the number of R-loop foci within 1 hour of treatment (Figure 1F).

[0097] Since almost all R-loops formed in cells arise from co-transcription, the specificity of RHINO was tested using tryptride, a comprehensive transcription inhibitor. One hour of treatment of U2OS cells with tryptride was sufficient to significantly reduce the number of R-loop foci detected by RHINO in a large cell population (Figure 1G). To track the effect of tryptride in individual cells, cells expressing the same RHINO were imaged one hour before and one hour after tryptride treatment. This experiment revealed a 46% reduction in the number of R-loop foci within one hour after transcriptional inhibition (Figure 1H).

[0098] Experiment 2: Detection of R-loop dynamic behavior in nuclear substructures Due to the lack of robust tools for imaging and collecting quantitative parameters of the R-loop in living cells, a clear understanding of the dynamic behavior of the R-loop at different genomic loci and their specific dynamics is limited. The nucleolus is a nuclear structure rich in R-loops (Santos-Pereira JM. Nat Rev Genet. 2015 Oct;16(10):583-97). Consistent with findings from Experiment 1, strong nucleolar staining, including separate foci, was observed in RHINO-expressing cells (Figure 2A). In particular, the nucleolar foci colocalized with upstream binding factor (UBF) foci, an RNA polymerase I transcription factor (Figure 2B). UBF labels sites of active ribosomal RNA transcription in the periphery of the fibrillary center and in the high-density fibrillary components of the nucleolus, where the R-loop is rich (Potapova, TA. Chromosome Res 27, 109-127 (2019)). To investigate the dynamic behavior of the R-loop in nucleoli, fluorescence recovery was performed after photobleaching (FRAP) experiments at the RHINO focus of individual nucleoli (Figure 2C). We observed a rapid turnover of the fluorescent RHINO molecule, with 90% of the initial fluorescence intensity recovered within 200 seconds (Figure 2C). The RHINO dynamic profile was comparable to that of the RNA polymerase I elongation phase on ribosomal genes, taking up to 182 seconds to complete the transcription cycle (Dundr M. Science. 2002 Nov 22;298(5598):1623-6). Overall, these data suggest that the R-loop is rapidly formed and dissolved at rDNA genes, explaining RHINO's ability to derive dynamic parameters for these structures in specific nuclear regions in living cells.

[0099] Human cancer cells, specifically telomerase-negative cells that rely on telomere alternative extension (ALT) mechanisms for telomere elongation (Sobinoff AP Trends Genet. 2017 Dec;33(12):921-932), are characterized by high levels of telomere R-loops (Arora R. Nat Commun. 2014 Oct 21;5:5220). U2OS is one type of ALT cell rich in telomere R-loops (Bertrand E. Cell. 1998 Oct;2(4):437-45). Several foci observed in RHINO-expressing U2OS cells co-localized with TRF2 foci, components of the sheltarin complex that assembles at telomeres (Figure 2D). Interestingly, FRAP experiments revealed two distinct populations of telomere R-loops: a major fraction (70%) in which fluorescence recovered to 88% of its initial intensity within 10–15 seconds after photobleaching, and another fraction in which little to no fluorescence recovery was observed (Figure 2E). FRAP experiments in cells expressing W43A and KK59AA RHINO mutants or GFP alone revealed much faster fluorescence recovery determined solely by the protein diffusion rate (Figure 2E), suggesting that RHINO FRAP is determined by its specific binding to the RNA:DNA hybrid portion of the telomere R-loop.

[0100] These data reveal the existence of two types of telomere R-loops with different dynamic behaviors: a more frequent and unstable type of structure that may correspond to co-transcribed R-loops formed during repetitive telomere transcription cycles and rapidly dissolving, and a type of persistent R-loop that can form trans telomeres in a RAD51-dependent pathway and is not rapidly replaced by telomere transcription (Feretzaki M. Nature. 2020 Nov;587(7833):303-308).

[0101] Experiment 3: Imaging of R-loops in actively transcribed genes in living cells The most documented physiological role of the R-loop concerns its influence on the transcription of protein-coding genes by RNA polymerase II (Non-Patent Literature 1). Furthermore, the lack of tools to accurately detect and monitor these structures in living cells makes a deeper understanding of the mechanisms and factors governing such functions impossible. Using RHINO, we collected novel dynamic parameters of the R-loop formed in protein-coding genes in human cells. We established a cell line containing tandem genome integration of human β-globin reporter gene arrays carrying the MS2 system for in vivo transcription imaging (Figure 3A) (Martins SB.. Nat Struct Mol Biol. 2011 Sep 4;18(10):1115-23). To avoid interference with R-loop formation by the MS2-binding protein (MS2CP) (Bonnet A. Mol Cell. 2017 Aug 17;67(4):608-621.e6), the inventors tagged reverse tetracycline trans-activator (rtTA) with red fluorescent protein (mCherry) to simultaneously label the reporter gene locus and induced its transcription upon doxycycline (Dox) addition (Figures 3A and 3B). The R-loop formed by the reporter gene was imaged in living cells expressing RHINO and mCherry-rtTA upon doxycycline addition, but not by MS2CP (Figure 3C). Imaging of living cells for 40 minutes provided a unique real-time record of R-loop level fluctuations during continuous transcription cycles (Figure 3D).

[0102] To image R-loops formed during single-gene transcription, an R-loop tendency sequence found in the β-actin gene (Skourti-Stathaki K.. Mol Cell. 2011 Jun 24;42(6):794-805) was inserted into the intron of a mouse IgM reporter gene, which also contained the PP7 stem-loop in exon II for direct visualization of its transcription (Figure 3E) (Vitor AC. Sci Adv. 2019 Jan 9;5(1):eaau1249). The reporter gene locus was labeled with mCherry-rtTA. Intensity line scans showed overlap between mCherry-rtTA and PP7-GFP (Figure 3E) and RHINO (Figure 3F), verifying RHINO's ability to detect R-loops formed on individual genes with high sensitivity and spatial resolution.

[0103] Overall, these data, utilizing RHINO imaging, can provide a dynamic framework for the dynamic behavior of R-loops at different genomic loci: rDNA, telomeres, and protein-encoding genes. RHINO's ability to collect precise quantitative parameters from living cells offers unique opportunities for future research seeking to explain inter-R-loop interactions and processes such as transcription, telomere maintenance, rRNA biosynthesis, and other functions exhibiting these nucleic acid structures.

[0104] References Altschul SF. J Mol Biol. 1990 Oct 5;215(3):403-10. Altschul SF. Nucleic Acids Res. 1997 Sep 1;25(17):3389-402. Arora R. Nat Commun. 2014 Oct 21;5:5220 Ausubel et al. (1999) 4th eds., Short Protocols in Molecular Biology: A Compendium of Methods from Current Protocols in Molecular Biology, John Wiley & Sons. Bertrand E. Mol Cell. 1998 Oct;2(4):437-45. Bou-Nader C.. Nature Communications. 2022 Mar;13(1):1641. Bonnet A. Mol Cell. 2017 Aug 17;67(4):608-621.e6. Cerritelli SM. Methods Mol Biol. 2022;2528:91-114. Crossley MP. J Cell Biol. 2021 Sep 6;220(9):e202101092. Dundr M. Science. 2002 Nov 22;298(5598):1623-6. Feretzaki M. Nature. 2020 Nov;587(7833):303-308 Marnef A. Nat Cell Biol. 2021 Apr;23(4):305-313. Martins SB. Nat Struct Mol Biol. 2011 Sep 4;18(10):1115-23. Niehrs C.. Nat Rev Mol Cell Biol. 2020 Mar;21(3):167-178. Nowotny M. EMBO J. 2008 Apr 9;27(7):1172-81. Pearson WR and Lipman DJ. Proc Natl Acad Sci U S A. 1988 Apr;85(8):2444-8. Potapova, TA. Chromosome Res 27, 109-127 (2019) Sambrook & Russell (2001) Molecular Cloning: a Laboratory Manual: 3rd edition, Cold Spring Harbor Laboratory Press Santos-Pereira JM. Nat. Rev Genet. 2015 Oct;16(10):583-97 Smith TF, Waterman MS.. J Mol Biol. 1981 Mar 25;147(1):195-7 Smolka JA.. J Cell Biol. 2021 Jun 7;220(6):e202004079. Sobinoff AP.. Trends Genet. 2017 Dec;33(12):921-932. Skourti-Stathaki K.. Mol Cell. 2011 Jun 24;42(6):794-805 Vitor AC. Sci Adv. 2019 Jan 9;5(1):eaau1249 Wang K. Sci Adv. 2021 Feb 17;7(8):eabe3516.

[0105] Sequence FYAVRRGRKTGVFLTWNECRAQVDRFPAARFKKFATEDEAWAFVRKSAS Sequence No. 1 RNase H1 HBD

[0106] TTCTATGCCGTGAGGAGGGGCCGCAAGACCGGGGTCTTTCTGACCTGGAATGAGTGCAGAGCACAGGTGGACAGATTTCCTGCTGCCAGATTTAAGAAGTTTGCCACAGAGGATGAGGCCTGGGCCTTTGTCAGGAAATCAGCAAGC Coding sequence of Sequence No. 2 RNase H1 HBD

[0107] MFYAVRRGRKTGVFLTWNECRAQVDRFPAARFKKFATEDEAWAFVRKSAS Sequence ID 3 RNase H1 HBD#1

[0108] ATGTTCTATGCCGTGAGGAGGGGCCGCAAGACCGGGGTCTTTCTGACCTGGAATGAGTGCAGAGCACAGGTGGACAGATTTCCTGCTGCCAGATTTAAGAAGTTTGCCACAGAGGATGAGGCCTGGGCCTTTGTCAGGAAATCAGCAAGC Code sequence of RNase H1 HBD#1 (sequence number 4)

[0109] VFYAVRRGRKTGVFLTWNECRAQVDRFPAARFKKFATEDEAWAFVRKSAS Sequence ID 5 RNase H1 HBD#2

[0110] GTGTTCTATGCCGTGAGGAGGGGCCGCAAGACCGGGGTCTTTCTGACCTGGAATGAGTGCAGAGCACAGGTGGACAGATTTCCTGCTGCCAGATTTAAGAAGTTTGCCACAGAGGATGAGGCCTGGGCCTTTGTCAGGAAATCAGCAAGC Code sequence of RNase H1 HBD#2, sequence number 6

[0111] MVSKGEEDNMASLPATHELHIFGSINGVDFDMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGYGFHQYLPYPDGMSPFQAAMVDGSGYQVHRTMQFEDGASLTVNYRYTYEG SHIKGEAQVKGTGFPADGPVMTNSLTAADWCRSKKTYPNDKTIISTFKWSYTTGNGKRYRSTARTTYTFAKPMAANYLKNQPMYVFRKTELKHSKTELNFKEWQKAFTDVMGMDELYK Sequence ID 7: Fluorescent protein (mNeonGreen)

[0112] ATGGTGAGCAAGGGCGAGGAGGATAACATGGCCTCTCTCCCAGCGACACATGAGTTACACATCTTTGGCTCCATCAACGGTGTGGACTTTGACATGGTGGGTCAGGGCACCGGCAATCCAAATGATGGTTATGAGGAGTTAAACCTGAAGTCCACCAAGGGTGACCTCCAGTTCTCC CCCTGGATTCTGGTCCCTCATATCGGGTATGGCTTCCATCAGTACCTGCCCTACCCTGACGGGATGTCGCCTTTCCAGGCCGCCATGGTAGATGGCTCGGCTACCAAGTCCATCGCACAATGCAGTTTGAAGATGGTGCCTCCCTTACTGTTAACTACCGCTACACCTACGAGGGAA GCCACATCAAAGGAGAGGCCCAGGTGAAGGGGACTGGTTTCCCTGCTGACGGTCCTGTGATGACCAACTCGCTGACCGCTGCGGACTGGTGCAGGTCGAAGAAGACTTACCCCAACGACAAAACCATCATCAGTACCTTTAAGTGGAGTTACACCACTGGAAATGGCAAGCGCTACCG GAGCACTGCGCGGACCACCTACACCTTTGCCAAGCCAATGGCGGCTAACTATCTGAAGAACCAGCCGATGTACGTGTTCCGTAAGACGGAGCTCAAGCACTCCAAGACCGAGCTCAACTTCAAGGAGTGGCAAAAGGCCTTTACCGATGTGATGGGCATGGACGAGCTGTACAAGTAA Sequence ID 8: Coding sequence for the fluorescent protein (mNeonGreen)

[0113] VDGTAGPGSA Sequence ID 9 Linker #1

[0114] GTGGACGGAACAGCGGGCCCGGGATCTGCC Sequence ID 10: Code array of linker #1

[0115] VDGTAGPGSTGVTGKLDPPVAT Sequence ID 11 Linker #2

[0116] GTGGACGGCACCGCTGGCCCGGGATCGACCGGAGTTACAGGTAAGCTGGATCCACCGGTCGCCACC Sequence ID 12: Code sequence of linker #2

[0117] TIFF2026509984000002.tif34170 Sequence ID No. 13: Fusion protein that binds to DNA:RNA hybrids (HBD is underlined, spacer is italicized, reporter is dashed)

[0118] TIFF2026509984000003.tif96170 Sequence ID No. 14: Code sequence of a fusion protein that binds to a DNA:RNA hybrid (HBD is underlined, spacer is italicized, reporter is dashed)

Claims

1. (i) Two or more hybrid binding domains (HBDs), (ii) Reporter domain and A fusion protein that includes and binds to DNA:RNA hybrids.

2. A fusion protein according to claim 1, which binds to a double-stranded DNA:RNA hybrid.

3. A fusion protein according to claim 1 or claim 2 that binds to the R-loop.

4. The fusion protein according to any one of claims 1 to 3, wherein the HBD is RNase HBD, and optionally human RNase H1 HBD, or a fragment thereof.

5. A fusion protein according to any one of claims 1 to 4, comprising three HBDs.

6. The fusion protein according to any one of claims 1 to 5, wherein each HBD comprises a peptide containing a Tyr residue at a position corresponding to Trp16 of SEQ ID NO: 1, a Lys residue at a position corresponding to Lys32 of SEQ ID NO: 1, and a Lys residue at a position corresponding to Lys33 of SEQ ID NO:

1.

7. The fusion protein according to any one of claims 1 to 6, wherein each HBD contains or consists of the amino acid sequence described in SEQ ID NO: 1, or a variant thereof.

8. A fusion protein according to any one of claims 1 to 7, comprising a first HBD, a second HBD, and a third HBD.

9. The fusion protein according to claim 8, wherein the first HBD comprises or consists of the amino acid sequence described in SEQ ID NO: 3, or a variant thereof.

10. The fusion protein according to claim 8, wherein the second HBD and / or the third HBD comprises or consists of the amino acid sequence described in SEQ ID NO: 5, or a variant thereof.

11. The fusion protein according to any one of claims 1 to 10, wherein the reporter domain comprises a fluorescent protein.

12. The fusion protein according to claim 11, wherein the fluorescent protein comprises green fluorescent protein (GFP) or a derivative thereof.

13. The fusion protein according to claim 12, wherein the GFP derivative is mNeonGreen.

14. A fusion protein according to any one of claims 1 to 13, further comprising one or more linker peptides.

15. The fusion protein according to claim 14, comprising a linker peptide located between the HBDs.

16. The fusion protein according to claim 14 or claim 15, comprising a linker peptide located between the HBD and the reporter domain.

17. The fusion protein according to any one of claims 14 to 16, wherein each linker peptide comprises or consists of the amino acid sequence described in SEQ ID NO: 9, or a variant thereof.

18. A fusion protein according to any one of claims 14 to 17, comprising a first linker peptide, a second linker peptide, and a third linker peptide.

19. The fusion protein according to claim 18, wherein the first linker peptide and / or the second linker peptide comprises or consists of the amino acid sequence described in SEQ ID NO: 9, or a variant thereof.

20. The fusion protein according to claim 18, wherein the third linker peptide comprises or consists of the amino acid sequence described in SEQ ID NO: 11, or a variant thereof.

21. A fusion protein according to any one of claims 1 to 20, comprising or consisting of the amino acid sequence described in SEQ ID NO: 13, or a variant thereof.

22. A nucleic acid encoding a fusion protein according to any one of claims 1 to 21.

23. The nucleic acid according to claim 22, comprising the nucleic acid sequence described in Sequence ID No. 2, which codes for RNase H1 HBD.

24. The nucleic acid according to claim 22 or claim 23, comprising the nucleic acid sequence described in Sequence ID No. 8, which encodes a fluorescent protein.

25. The nucleic acid according to any one of claims 22 to 24, comprising the nucleic acid sequence described in SEQ ID NO: 10, which encodes a linker sequence.

26. The nucleic acid according to any one of claims 22 to 25, comprising or consisting of the nucleic acid sequence described in Sequence ID No.

14.

27. The nucleic acid according to any one of claims 22 to 26, which is operably connected to a promoter.

28. The nucleic acid according to claim 27, wherein the promoter is an inducible promoter.

29. The nucleic acid according to claim 28, wherein the inducible promoter comprises a tetracycline-regulating transactivator.

30. An isolated cell comprising a fusion protein according to any one of claims 1 to 21 or a nucleic acid according to any one of claims 22 to 29.

31. A method for detecting R-loops in cells, (i) Culturing isolated cells expressing a fusion protein that binds to a DNA:RNA hybrid as described in any one of claims 1 to 21, (ii) Identifying one or more sites in the cell where the fusion protein accumulates, A method comprising the accumulation of the fusion protein, wherein the accumulation of the fusion protein indicates the presence of an R-loop at the site.

32. A method for screening compounds that regulate the R-loop, (i) Culturing isolated cells expressing a fusion protein that binds to the DNA:RNA hybrid according to any one of claims 1 to 21 in the presence of the test compound such that the fusion protein binds to the R loop in the cell, (ii) Determining the distribution of the fusion protein within the cell, A method comprising a change in the distribution of the fusion protein in the cells in the presence of the test compound compared to the absence of the test compound, thereby demonstrating that the compound modulates the R-loop.