Means and methods for preparing engineered target proteins by genetic code expansion in target protein-selective manner

An orthogonal translation system using RNA-TP and O-RS segments, along with AFPs, enables site-specific introduction of amino acid residues onto POIs in mammalian cells, addressing the challenges of background suppression and cellular disruption in existing methods.

JP2025081525AActive Publication Date: 2025-05-27EURO LAB FUER MOLEKULARBIOLOGIE EMBL
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025025264
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-02-14
Filing Date
2025-02-19
Publication Date
2025-05-27
Estimated Expiration
2040-02-14

AI Technical Summary

Technical Problem

Current methods for orthogonal translation in living cells, particularly in eukaryotic cells, face challenges due to the complexity of the genome and the high abundance of amber codons, leading to background suppression of main proteins and undesired incorporation of non-standard amino acids.

Method used

The development of an orthogonal translation system that utilizes RNA-targeting polypeptide (RNA-TP) segments and orthogonal aminoacyl-tRNA synthetase (O-RS) segments, combined with assembler fusion proteins (AFPs), to achieve site-specific introduction of canonical amino acid residues onto a polypeptide of interest (POI) by spatially concentrating these components.

Benefits of technology

This approach allows for the selective and efficient introduction of non-canonical amino acid residues onto specific POIs in mammalian cells, minimizing background incorporation and maintaining cellular physiology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025081525000016
    Figure 2025081525000016
  • Figure 2025081525000017
    Figure 2025081525000017
  • Figure 2025081525000018
    Figure 2025081525000018
Patent Text Reader

Abstract

To provide orthogonal translation systems which allow for the site-specific introduction of non-canonical amino acid residues into a target protein (POI).SOLUTION: The present invention relates to assembler fusion proteins, which bring an RNA-targeting polypeptide (RNA-TP) segment and an orthogonal aminoacyl tRNA synthetase (O-RS) segment into spatial proximity of one another, either by direct linkage in RNA-TP / O-RS fusion proteins, or though the action of "assemblers" fused to each of these segments in assembler fusion proteins (AFPs). The invention also relates to AFP combinations and nucleic acid molecules comprising a POI-encoding sequence together with a targeting nucleotide sequence that is able to interact with an RNA-TP. The invention further relates to nucleic acid molecules, expression cassettes and expression vectors and the like encoding the RNA-TP / O-RS fusion proteins or AFPs.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a method for selectively inserting a non-coding region onto a polypeptide of interest (POI) in a POI mRNA. The present invention relates to an orthogonal translation system that allows the site-specific introduction of canonical amino acid (ncAA) residues. Specifically, the present invention provides an RNA-targeting polypeptide (RNA-TP) segment and an orthogonal amino acid sequence. Fusion of acyl-tRNA synthetase (O-RS) segments into close spatial proximity to each other This is a protein that combines an RNA-TP segment and an O-RS segment into one Combined as the same fusion protein (RNA-TP / O-RS fusion protein) or acts as an "assembler" (AP) to assemble one or more APs into an RNA-TP Assembler fusion proteins (AFPs) containing the .alpha.-amino acid sequence and the O-RS segment. Facilitating local condensation, thereby condensing the RNA-TP segment and the O-RS segment This is achieved by the action of one or more polypeptide segments that bring the The present invention relates to a targeting nucleotide sequence (TN) capable of interacting with RNA-TP. The present invention also relates to a nucleic acid molecule comprising a POI coding sequence together with an AFP combination. The present invention further provides a nucleic acid molecule encoding the RNA-TP / O-RS fusion protein or AFP. , expression cassettes and expression vectors, cells containing the same, and translationally preparing POIs The present invention relates to a method and a kit for the same. [Background technology]

[0002] The ability to engineer orthogonal (i.e., non-cross-reactive) translation systems into living cells in a site-specific manner allows the introduction of new functional groups into proteins. However, this is a challenging task. Because translation is a complex multi-step process. At least 20 different aminoacylated tRNAs and their corresponding aminoacyl tRNAs were identified. The ribosome, tRNA synthetase (RS), and various other factors work together to The ideal orthogonal system would be able to synthesize polypeptide chains from RNA transcripts, and interact with factors of the host machinery. showed no differential reactivity and their impact on cellular housekeeping translation activity and normal physiology Minimize the project.

[0003] To this end, genetic code expansion (GCE) is a method for reprogramming specific codons. GCE is a method that allows for the use of orthogonal (suppressed) RS (O-RS) The corresponding inhibitor tRNA can be aminoacylated with a non-standard amino acid (ncAA). Typically, these ncAAs are custom designed and have chemical functionality, which have demonstrated, for example, that protein function can be light-regulated or that proteins that encode post-translational modifications or use click chemistry to introduce fluorescent labels for microscopy studies To introduce ncAAs site-specifically onto a polypeptide of interest (POI), Therefore, the anticodon loop of the tRNA decodes and therefore suppresses one of the stop codons. (e.g., Liu et al., Annu Rev Biochem 2010,79:413-444;Lemke,ChemBioChem 2014, 15:1691-1694;Chin,Nature 2017,550;53-60 To minimize the impact on the host cell machinery, endogenous proteins should be terminated. Due to its particularly low abundance in E. coli for binding (<10%), Amber stop codon (corresponding tRNA CUA ) is often used. Regardless of the nature of the amber codon, in principle, any amber codon in the genome can be suppressed, and the non-target host can be suppressed. This may lead to undesired background suppression of the main protein. When proteins are produced recombinantly for in vitro applications, this background Incorporation of the endonucleases may be tolerated as long as the yield of purified full-length protein is acceptable. However, the host is only capable of sacrificing bioreactors for its proteins. The challenge is different when the in situ host cell PO is considered to be a large-scale in vitro culture. To study the function of I, the physiological state of the host cell is an important factor. In this study, minimization of background incorporation of ncAAs was achieved through well-controlled experiments. It is specifically required to ensure

[0004] To enable orthogonal translation in E. coli, i.e., PO At least three excellent approaches have been proposed to decode specific codons only for RNA of type I. i) Orthogonal ribosomes that recognize unique Shine-Dalgarno sequences have been developed. These have been developed to decode quadruplet codons instead of stop codons. Instead, it is used to site-specifically encode the ncAA on the POI (e.g., Heuma nn et al.,Nature 2010,464:441:444;Orelle et al.,Nature 2015,524:119-124;Fried et al. (see, e.g., Angew Chem 2015, 54:12791-12794). i) Recently, genome engineering has allowed E. coli strains to be depleted of selected native codons. The gene for selectively reading specific codons only on the POI has been developed. A genetically clean (eg, amber codon-free) host background was provided. (See, e.g., Isaacs et al., Science 2011, 333:348- 353;Lajoie et al.,Science 2013,342:357-3 60;Ostrov et al.,Science 2016,353:819-82 2; see Wang et al., Nature 2016, 539:59-64). i ii) Unique non-standard codons are encoded using artificial base pairs only in the coding sequence of the POI. This reduces the risk of non-specific readings in other parts of the genome. (See Zhang et al., Nature 2017, 551:644-647) However, due to the complexity of the genome, these orthogonal translation approaches have not been applied to eukaryotic cells. Transfer to living organisms is not straightforward (e.g., Thompson et al., CS Chem Biol 2018,13:313-325). In addition, Amber codons are highly abundant in mammalian cells (20%).

[0005] Therefore, E.coli, which is versatile and relatively easy to handle and operate, is It works not only in well-characterized prokaryotes such as li, but is also applicable to eukaryotic cells. There is a high demand for a strategy for POI-selective orthogonal translation. It was the object of the present invention to address this issue. Summary of the Invention

[0006] The present inventors have demonstrated that ncAA residues can be translationally introduced onto the growing polypeptide chain of a POI. By providing spatial proximity between the POI mRNA and the O-RS, An orthogonal translation system (OT system) that can selectively translate POI mRNA is created The present inventors have found that various POIs, including membrane proteins, These OT systems share the same stop codon (which is the sequence that encodes the ncAA residue of the POI). The POI is compared to other mRNAs in the cytoplasm that contain The site-specific addition of ncAA residues to the POI in mammalian cells is a key step in the selection of mRNA. It has been demonstrated that it is possible to introduce the

[0007] In the orthogonal translation system of the present invention, spatial proximity is achieved by RNA-targeting polypeptide (RNA-TP) The mRNA of the POI contains a targeting sequence (TN) capable of selectively interacting with the targeting sequence. This is accomplished by linking an O-RS to such an RNA-TP. The fusion protein contains both O-RS and RNA-TP (RNA-TP / O-RS fusion protein). The protein may be a recombinant protein.

[0008] In another approach, this is achieved by integrating at least two assembler fusion proteins (AFPs) One or more polymorphs that act as "assemblers" (APs) to facilitate localized concentration of This can be achieved by the action of peptide segments, at least one of which is Contains one or more AP and RNA-TP segments, and at least one other AFP contains one or more and therefore, the RNA-TP and O-RS segments ment (RNA-TP and O-RS are also called "effector" or "EP") The local enrichment of AFPs can be used to generate artificial orthogonal translation organelles. An assembly capable of acting as a network integrator (OT assembly, herein referred to as an "OT integrator" This allows the formation of a 'ganera' (also called a 'ganera').

[0009] The inventors have demonstrated that different types of APs can be used. The first type is Subcellular structures (e.g., microtubules, or cell or nuclear membranes, ER, mitochondria, etc.) APs drive localized enrichment at the cytoplasmic side of the membrane (such as the amyloid plaque or Golgi organelles) The AP comprises the intracellular targeting polypeptide (IC-TP) segment. Type 2 is characterized by high local AFP concentrations due to self-association (particularly by phase separation) in the cytoplasm. These segments are referred to herein as phase-separated polypeptide (PSP) segments. The AP type is particularly characterized by the presence of coiled cores formed by synthetic SYNZIP polypeptide pairs. It can also be combined with other polypeptide elements that have the ability to form multimeric structures such as yl heterodimers. Similarly, the EP type may be a synthetic SYNZIP polypeptide pair. Ability to form multimeric structures, particularly coiled-coil heterodimers formed by Such multimerization can also be combined with other polypeptide elements having the following structure: This further improves local concentration of.

[0010] The present inventors further demonstrated that AFPs combining different AP types are particularly useful. Found it.

[0011] In yet another approach, the EP serine / peptide sequences are expressed on a single polypeptide, i.e., fused together. Both types of segments, i.e., RNA-TP and O-RS segments, and AP segments AFPs containing one or both types, i.e., IC-TP and / or PSP segments, Optionally, the polypeptide element (SYN) is capable of forming a multimeric structure. ZIP polypeptide) is required to create the OT system of the present invention. It offers the advantage that all elements are contained on one single AFP. Thus, in a first aspect, the present invention provides a method for producing a method for treating a cancer cell comprising: (a) (a1) A polypeptide segment derived from an intracellular targeting polypeptide (IC-TP segment (The intracellular targeting polypeptide is a target polypeptide for a cell within or directly adjacent to the cytoplasm.) (targets and is therefore locally concentrated in intracellular structural elements), and (a2) A polypeptide segment (PSP segment) derived from a phase-separated polypeptide ( The phase-separating polypeptides act to induce cell proliferation in a manner that creates sites of high local concentration in the cytoplasm. (having the ability to undergo self-association in the cytoplasm), At least one first polypeptide acting as an assembler (AP) selected from A segment, (b) b1) an RNA-targeting polypeptide (RNA-TP) segment, and b2) orthogonal aminoacyl-tRNA synthetase (O-RS) segment; At least one second polypeptide acting as an effector (EP) selected from A segment and The polypeptide segments are connected to each other in the assembler fusion protein (AFP) comprising: It is functionally attached on the AFP.

[0012] In a second aspect, the present invention provides a method for the preparation of a medicament comprising administering to a subject the method of the present invention at least two AFPs of the present invention as described herein. The present invention relates to an assembler fusion protein (AFP) combination comprising the assembler fusion protein (AFP) of the present invention. The alignment is performed by combining at least one AFP containing an RNA-TP segment and an O-RS segment. and at least one AFP comprising: including a first SYNZIP element and on at least another AFP of the combination. The inclusion of a second SYNZIP element represents another advantageous embodiment of the second aspect, The first and second SYNZIPs work together by forming a heterodimeric structure. Use.

[0013] In a third aspect, the present invention provides a method for producing a composition comprising: (i) at least one RNA-targeting polypeptide (RNA-TP) segment; (ii) at least one orthogonal aminoacyl-tRNA synthetase (O-RS) segment; To and; Regarding a fusion protein (RNA-TP / O-RS fusion protein) comprising The polypeptide segment comprises a functional group on the RNA-TP / O-RS fusion protein. are effectively combined.

[0014] In a further aspect, the present invention provides a nucleic acid molecule, or a combination of two or more nucleic acid molecules. : (i) at least one RNA-TP / O-RS fusion protein of the invention described herein; a nucleotide sequence encoding a protein; or (ii) a nucleic acid sequence complementary to (i); or (iii) Both (i) and (ii); Includes.

[0015] In a further aspect, the present invention provides a nucleic acid molecule, or a combination of two or more nucleic acid molecules. : (i) a nucleotide sequence encoding at least one AFP of the invention as described herein; Array, or (ii) a nucleic acid sequence complementary to (i); or (iii) Both (i) and (ii); Includes.

[0016] In a further aspect, the present invention provides a nucleic acid molecule, or a combination of two or more nucleic acid molecules. : (i) a nucleic acid encoding at least one AFP combination of the invention described herein; a nucleotide sequence, or (ii) a nucleic acid sequence complementary to (i); or (iii) Both (i) and (ii); Includes.

[0017] In a further aspect, the present invention provides an expression cassette comprising the nucleic acid of the invention as described herein. It includes the nucleotide sequence of a molecule or combination of nucleic acid molecules.

[0018] In a particular embodiment, the present invention provides an expression cassette comprising: (i) at least one RNA-TP / O-RS fusion protein of the invention described herein; a nucleotide sequence encoding a protein; or (ii) a nucleic acid sequence complementary to (i); or (iii) Both (i) and (ii); Includes.

[0019] In a further particular embodiment, the present invention provides an expression cassette comprising: (i) a nucleotide sequence encoding at least one AFP of the invention as described herein; Array, or (ii) a nucleic acid sequence complementary to (i); or (iii) Both (i) and (ii); Includes.

[0020] In a further particular embodiment, the present invention provides an expression cassette comprising: (i) a nucleic acid encoding at least one AFP combination of the invention described herein; a nucleotide sequence, or (ii) a nucleic acid sequence complementary to (i); or (iii) Both (i) and (ii); Includes.

[0021] In a further aspect, the present invention provides at least one expression cassette of the invention described herein. An expression vector comprising the vector is provided.

[0022] In a further aspect, the present invention provides at least one nucleic acid molecule of the invention as described herein. In certain embodiments, the cell comprises a nucleic acid molecule of the present invention. At least one expression cassette or at least one expression vector of the invention described in the specification This includes the culprit.

[0023] In a further aspect, the present invention relates to a method for the preparation of a medicament comprising the step of: The present invention relates to a method for preparing a polypeptide of interest (POI) comprising a POI-A residue. The method comprises expressing a POI by a cell of the invention in the presence of said one or more ncAAs. See, cells: (i) at least one AF comprising an RNA-TP segment, as described herein; at least one AFP comprising a P and an O-RS segment; (ii) a nucleotide sequence encoding the POI (CS POI ) (one or more of the POIs the ncAA residue is encoded by a selector codon(s), (iii)CS POI and functionally binding at least one of AFPs in the cell. a targeting nucleotide sequence (TN) capable of interacting with the RNA-TP segment; (iv)CS POI The selector codon(s) of One or more orthogonal tRNAs carrying ncAA (O-tRNA ncAA ) molecule (before O-tRNA ncAA The molecule is then coupled to one or more O-RS segments of the AFP in the cell. and one or more orthogonal O-RS / O-tRNA ncAA They form pairs, and these are PO allowing the introduction of said one or more ncAA residues onto the amino acid sequence of I); and optionally, the method further comprises recovering the expressed POI. (i) at least one AFP comprising an RNA-TP segment as described above; The at least one AFP containing an O-RS segment is an AFP of the same type, i.e. That is, it may be an AFP that contains both an RNA-TP segment and an O-RS segment. Alternatively, the at least one RNA-TP segment described in (i) The AFP and the at least one AFP comprising an O-RS segment are different AFPs. It's fine.

[0024] In a further aspect, the present invention relates to a method for the preparation of a medicament comprising the step of: The present invention relates to a method for preparing a polypeptide of interest (POI) comprising a POI-A residue. The method comprises expressing a POI by a cell of the invention in the presence of said one or more ncAAs. See, cells: (i) an RNA-TP / O-RS fusion protein of the invention as described herein; (ii) a nucleotide sequence encoding the POI (CS POI ) (one or more of the POIs the ncAA residue is encoded by a selector codon(s), (iii)CS POI The RNA-TP / O-RS fusion protein is functionally linked to the A target that can interact with at least one RNA-TP segment of the fusion protein Targeted nucleotide sequence (TN); (iv)CS POI The selector codon(s) of One or more orthogonal tRNAs carrying ncAA (O-tRNA ncAA ) molecule (before O-tRNA ncAA The molecule is one of the RNA-TP / O-RS fusion proteins in the cell. Together with the O-RS segments, one or more orthogonal O-RS / O-tRNA nc AA and forming pairs, which are the binding of said one or more ncAA residues onto the amino acid sequence of the POI. enabling adoption); Including, The method optionally further comprises recovering the expressed POI.

[0025] In a further aspect, the present invention relates to a method for the preparation of a medicament comprising the step of: The present invention relates to a method for preparing a polypeptide of interest (POI) comprising a POI-A residue. teeth: (a) a nucleic acid sequence comprising at least one RNA-TP segment, as described herein; and one or more AFPs comprising at least one O-RS segment, and expressing the (b) one or more orthogonal tRNAs ncAA (O-tRNA ncAA ) molecules by the cells And it is expressed, - the orthogonal tRNA ncAAOne or more of the O-RS segments of the molecule and the AFP Within the , one or more orthogonal aminoacyl-tRNA synthetases / tRNA ncAA (O -RS / O-tRNA ncAA ) form a pair, - the O-RS / O-tRNA ncAA The pair is a 1 on the amino acid sequence of the POI. It allows the introduction of one or more ncAA residues, Steps (a) and (b) may be simultaneous or sequential in any order. and; (c) then expressing the POI by the cell in the presence of the one or more ncAAs. death, - the nucleotide sequence encoding the POI (CS POI ) is a nucleotide sequence of one or more ncAA residues one or more selector codons encoding a group, - the selector codon is selected from the one or more O-tRNAs ncAA The anticodon of the molecule Match; - Said CS POI is functionally linked to the targeting nucleotide sequence (TN), and thus CS POI / TN fusion sequence is formed, - Said CS POI The / TN fusion sequence is capable of directing at least one of the AFPs in a cell via the TN. a step capable of interacting with one RNA-TP segment; (d) optionally recovering the expressed POI; Includes.

[0026] In a further aspect, the present invention relates to a method for the preparation of a medicament comprising the step of: The present invention relates to a method for preparing a polypeptide of interest (POI) comprising a POI-A residue. teeth: (a) Inducing the inventive RNA-TP / O-RS fusion protein described herein by a cell and expressing the compound; (b) one or more orthogonal tRNAs ncAA (O-tRNA ncAA ) molecules by the cells And it is expressed, - one or more O-RS segments of the RNA-TP / O-RS fusion protein and said cross-tRNA ncAA The molecule is capable of catalyzing one or more orthogonal aminoacyl-tRNA synthases in a cell. Sesatase / tRNA ncAA (O-RS / O-tRNA ncAA ) form a pair, - the O-RS / O-tRNA ncAA The pair is a 1 on the amino acid sequence of the POI. It allows the introduction of one or more ncAA residues, Steps (a) and (b) may be simultaneous or sequential in any order. and; (c) then expressing the POI by the cell in the presence of the one or more ncAAs. death, - the nucleotide sequence encoding the POI (CS POI ) is a nucleotide sequence of one or more ncAA residues one or more selector codons encoding a group, - the selector codon is selected from the one or more O-tRNAs ncAA The anticodon of the molecule Match; - Said CS POI is functionally linked to the targeting nucleotide sequence (TN), and thus CS POI / TN fusion sequence is formed, - Said CS POI The / TN fusion sequence mediates intracellular RNA-TP / O- It is possible to interact with at least one RNA-TP segment of the RS fusion protein. Steps to take; (d) optionally recovering the expressed POI; Includes.

[0027] In a further aspect, the present invention provides a method for producing a method for treating a cancer cell comprising: (i) a nucleotide sequence encoding a polypeptide of interest (POI) (CS POI )and( The POI is linked to a CS by a selector codon POI One or more of the same or similar contains different non-canonical amino acid (ncAA) residues), (ii) a targeting nucleotide sequence (TN), The RNA molecule containing the TN is targeted to an RNA targeting polypeptide (RNA-T P).

[0028] In a further aspect, the present invention provides a method for the preparation of a medicament having at least one non-standard amino acid (ncAA) residue. The present invention relates to a kit for preparing a polypeptide of interest (POI), the kit comprising: - at least one ncAA corresponding to at least one ncAA residue of the POI or Salt and - at least one expression vector of the invention as described herein, Includes. The expression vector comprises: (i) at least one RNA-TP / O-RS fusion protein of the invention, at least one AFP or at least one AFP combination of the invention a nucleotide sequence, or (ii) a nucleic acid sequence complementary to (i); or (iii) Both (i) and (ii); The nucleic acid sequence of the present invention comprises at least one expression cassette comprising: [Brief description of the drawings]

[0029] [Figure 1]A schematic diagram of the spatial separation of components that enables orthogonal translation to decode a specific stop codon on a uniquely tagged mRNA is shown. (A) Conventional expression of the synthetase PylRS leads to aminoacylation of its corresponding stop codon-suppressing tRNAPyl with a custom-designed ncAA. This leads to site-specific ncAA incorporation whenever the respective stop codon is present on the mRNA of the POI. Given that many endogenous mRNAs terminate at the same stop codon, utilizing this approach in the cytoplasm would potentially lead to the mis-incorporation of the ncAA on unwanted proteins (left box). (B) To avoid this, the present invention allows that the mRNA encoding the POI and the orthogonal aminoacyl-tRNA synthetase (e.g. PylRS) can be brought into close proximity with each other by the use of an RNA-targeting polypeptide segment (e.g. MCP) and an assembler (AP). This allows spatial concentration of all components, creating an OT assembly ("OT organelle") that contains the mRNA encoding the POI, the orthogonal aminoacyl-tRNA synthetase, the tRNA, and the ribosome (right box). Here, the aminoacylated tRNAPyl is particularly available in the direct vicinity of the OT organelle, so that stop codon suppression (of the POI mRNA) can occur, particularly here. This leads to selective suppression (and therefore expression) of the stop codon of the POI mRNA compared to the corresponding stop codon on the mRNA that is not targeted to the OT assembly. In (A), GCE occurs in a stop codon-specific manner, whereas in (B) it should occur in a stop codon-specific and mRNA-specific manner.

[0030] [Figure 2A] Schematic diagram of the different assembler classes: B = bimolecular MCP::PylRS fusion, P1 = fusions with FUS and EWSR1, P2 = SPD5, K1 = truncation of kinesin KIF13A (KIF13A1-411, ΔP390), K2 = truncation of kinesin KIF16B (KIF16B1-400), and combinations thereof (K1::P1, K1::P2, K2::P1, K2::P2).

[0031] [Figure 2B] A schematic diagram of a dual-color reporter is shown. mRNAs encoding the fluorescent proteins GFP and mCherry, which contain stop codons at permissive sites, are expressed from one plasmid. Each has its own CMV promoter, ensuring a constant ratio of mRNAs in each experiment. The mCherry reporter mRNA, mRNA(mCherry)::ms2, is tagged with two MS2 RNA stem-loops ("ms2", also referred to herein as MS2 tags). In the presence of ncAA and tRNAPyl, in the case of cytoplasmic PylRS, both GFP39 stops and mCherry185 stops are produced, leading to a diagonal line in fluorescence flow cytometry (FFC) analysis (left box). However, under the same conditions, orthogonal translation in the OT organelle allows selective stop codon suppression of mRNA(mCherry):ms2, resulting in a population that is mCherry positive and GFP negative (schematically depicted as a vertical population in the right box). In both schemes, untransfected HEK293T cells are represented by the lower grey circle.

[0032] [Figure 2C]Selectivity and relative efficiency of various exemplary OT systems are shown. In all experiments, the indicated constructs were co-expressed with tRNAPyl (anticodon corresponding to the indicated codon) and a dual reporter (GFP39stop, mCherry185stop::ms2). GCE was performed in the presence of the indicated ncAA and cells were analyzed by FFC. Dark grey bars (normalized to cytoplasmic PylRS) represent the ~fold change in the ratio r of the mean fluorescence intensity of mCherry to GFP (derived from FFC, see Fig. 2D,E) for all systems tested. Light grey bars represent the relative efficiency, defined by the mean fluorescence intensity of mCherry (derived from FFC, see Fig. 2D,E) for each condition divided by the cytoplasmic PylRS control. Mean values ​​of at least three independent experiments are shown; error bars represent SEM. Box highlights the best performing OT organelle (OTK2::P1).

[0033] [Figure 2D] Figure 1 shows the results of FFC analysis of the four indicated lines of transfected HEK293T cells in the presence of the lysine derivative ncAA SCO with a cyclooctyne side chain and the dual-color reporter expressed by tRNAPyl. Highly selective and efficient orthogonal translation was observed for OT assemblies (black arrows indicate bright highly mCherry-positive populations). The sum of at least three independent experiments is shown in the dot plot. The axis indicates fluorescence intensity in arbitrary units.

[0034] [Figure 2E] Shown are FFC plots of OT assemblies selectively translating only the opal and ochre codons of recruited mRNA(mCherry185TGA):ms2 and mRNA(mCherry185TAA):ms2, respectively.

[0035] [Diagram 3]Schematic diagrams of the constructs that make up the following systems are shown: PylRS, MCP::PylRS, FUS::MCP::PylRS, and LcK::FUS::PylRS·LcK::EWS::MCP.

[0036] [Figure 4] Flow cytometric analysis of dual reporter expression by the four different systems depicted in Figure 3. HEK293T cells were transfected with constructs encoding the dual reporter, tRNA, LcK::FUS::PylRS, and LcK::EWS::MCP, or PylRS, MCP::PylRS, FUS::MCP::PylRS, and pcDNA3.1. A total of at least three independent experiments is shown. The axis indicates the fluorescence intensity in arbitrary units.

[0037] [Diagram 5] A bar graph of the ratio of mean fluorescence intensity of mCherry to GFP fluorescence for all lines tested is shown. Plots represent the average of at least three biological replicates, and error bars indicate the standard error of the mean.

[0038] [Figure 6]FIG. 2 provides an overview of the different approaches of the invention to generate OT organelles targeted to the surface of different intracellular structures. Different constructs are expressed and the results of the respective fluorescence flow cytometry (FFC) analysis are shown. At the top of the figure, the dual-color reporter construct GFP39T AG·mCherry185TAG::ms2 (see also FIG. 2B) applied in each of the experiments A to G, which are illustrated in a schematic manner, is illustrated, showing a schematic illustration of the different target cell compartments. A control experiment performed without the effector polypeptide MCP (-MCP) is also illustrated for each of the experiments A to G. A: OT organelles targeted to microtubules and obtained by expressing the system KIF16B1-400::FUS::PylRS·KIF16B1-400::EWSR1::MCP or the construct KIF16B1-400::FUS::PylRS (control). B: OT organelles targeted to the microtubule plus ends and obtained by expressing the construct EB1::FUS::MCP::PylRS or EB1::FUS::PylRS (control). C: OT organelles targeted to the plasma membrane and obtained by expressing the system LcK::FUS::PylRS·LcK::EWSR1::MCP or the construct LcK::FUS::PylRS (control). D: OT organelles targeted to the mitochondrial membrane and obtained by expressing the system TOM201-70::FUS::PylRS·TOM201-70::EWSR1::MCP or the construct TOM201-70::FUS::PylRS (control). E: OT organelles targeted to the nuclear membrane and obtained by expressing the system CG1::FUS::PylRS·CG1::EWSR1::MCP or the construct CG1::FUS::PylRS (control). F (left side): OT organelles targeted to the Golgi membrane and obtained by expressing the system EBAG91-29::FUS::PylRS·EBAG91-29::EWSR1::MCP or the construct EBAG91-29::FUS::PylRS (control). F (right side): OT organelles targeted to the Golgi membrane and obtained by expressing the system CMP Sia Tr::FUS::PylRS·CMP Sia Tr::MCP or the construct CMP Sia Tr::FUS::PylRS (control).G: OT organelles targeted to the ER membrane and obtained by expressing the system P450 2C11-27::FUS::PylRS·P450 2C11-27::EWSR1::MCP or the construct P450 2C11- 27::FUS::PylRS (control).

[0039] [Figure 7] FIG. 1 provides an overview of the different approaches of the present invention for recruiting RNA using the interactions of different RNA loops and the respective RNA-targeting proteins. The results of the respective fluorescence flow cytometry (FFC) analyses are shown and compared with the respective analyses obtained for untargeted PylRS alone. A: The ms-2-MCP system incorporates an ms2 loop on the UTR of an mRNA molecule and recruits the mRNA to an artificial organelle by the MCP protein. B: The boxB-lambda N22 system incorporates a boxB loop on the UTR of an mRNA molecule and recruits the mRNA to an artificial organelle by the lambda N22 protein. C: The pp7-PCP system incorporates a pp7 loop on the UTR of an mRNA molecule and recruits the mRNA to an artificial organelle by the PCP protein.

[0040] [Figure 8]Further approaches of the invention to generate OT organelles that will function on the surface of different cellular structures are illustrated. Here, targeting to the plasma membrane is exemplified. A particular approach features the incorporation of the so-called synthetic heterodimeric coiled-coil peptides SYNZIP1 and SYNZIP2 as a pair fused to the system LcK::FUS::SYNZIP1::PylRS·EWSR1::SYNZIP2::MCP. Upon expression, SYNZIP1 and 2 pair up and recruit MCP to the plasma membrane-based OT organelle, which in turn allows selective orthogonal translation of subsequently recruited mRNAs that contain ms2-targeted nucleotide loops. The selective translation is illustrated by the results of the respective FFC analysis (A). In a comparative approach with the LcK::FUS::PylRS·EWSR1::SYNZIP2::MCP system lacking SYNZIP1, no translation selectivity could be observed (B). DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0041] Unless otherwise defined herein, scientific and technical terms used in the context of the present invention The technical terms have the meanings commonly understood by those skilled in the art. However, in case of any possible ambiguity, the following is provided herein: The definitions set forth herein take precedence over any dictionary or extrinsic definitions. Unless otherwise specified, singular terms include plurals and plural terms include the singular.

[0042] Unless otherwise indicated, nucleotide sequences are referred to herein from 5' to 3' Unless otherwise indicated, amino acid sequences are depicted herein in the are depicted in the N-terminal to C-terminal direction.

[0043] Unless otherwise stated, the objects translationally expressed by the OT system according to the invention are The polypeptide of interest (POI) is a nucleotide sequence encoding the POI by a selector codon. Array (CS POI ) contains one or more ncAA residues encoded above.

[0044] 1. Fusion Proteins

[0045] 1.1 General

[0046] The fusion proteins of the present invention can be construed in different ways.

[0047] The first type is a small molecule that contains at least one RNA-TP and at least one O-RS. At least two types of effector polypeptides (EP) are contained in one and the same fusion protein (R The fusion protein contained in the NA-TP / O-RS fusion protein Includes.

[0048] The second type consists of at least one assembler polypeptide (AP) and an RNA-TP segment. and at least one type of EP selected from the O-RS segment. In particular, AFPs include proteins that are at least one of the APs. In addition to the above two types, both RNA-TP and O-RS segments, e.g., one or more and one or more O-RS segments in any sequential order. In particular, the AFP may be selected from the following fusion protein types: The fragments may be functionally linked in any order along the polypeptide chain. One or more segments of the type may be in any order on the polypeptide chain:

[0049] (RNA-TP / AP) (O-RS / AP) (RNA-TP / O-RS / AP)

[0050] The AP is selected from the IC-TP and the PSP, and may be assigned to one or more IC- It can be composed of one or more of TP and / or PSP. Therefore, AFP is more Specifically, the following fusion protein types are selected (the segments are connected to each other on the polypeptide chain): In any order, one or more segments of the same type may be functionally linked to the polymer. In any order on the peptide chain: (RNA-TP / IC-TP) (O-RS / IC-TP) (RNA-TP / O-RS / IC-TP)

[0051] (RNA-TP / PSP) (O-RS / PSP) (RNA-TP / O-RS / PSP)

[0052] (RNA-TP / PSP / IC-TP) (O-RS / PSP / IC-TP) (RNA-TP / O-RS / PSP / IC-TP)

[0053] The AP and / or EP may be involved in hetero-oligomerization (as part of a fusion protein), particularly Heterodimer-forming polypeptide segments such as synthetic coiled-coil SYNZIP peptides Each AFP may also contain only one member of such interacting SYNZIP pairs. Such interacting groups are distributed among the members of the AFP combination so as to include The AFP combination including the SYNZIP pair is a specific embodiment.

[0054] The term "segment" as used herein in the context of a fusion protein means Elements designated as such (e.g., RNA-TP, O-RS, IC-TP, PSP, SY NZIP) is part of the fusion protein, i.e., linked to the rest of the fusion protein. The segments of the fusion protein of the present invention are functionally linked to each other. That is, they are RNA-TP, O-RS, IC-TP, and PSP or are linked in such a way that they function as SYNZIPs. The linkage is preferably covalent. and in particular peptidic linkages.

[0055] For example, the RNA-TP segment contained in the fusion protein of the present invention may be A segment of a fusion protein that is derived from and functions in the context of a fusion protein. Therefore, it is possible for the fusion protein to interact (bind) with the target RNA. However, the interaction is expediently specific. Therefore, the RNA-TP segment The amino acid sequence (overall) or functional sequence of the RNA targeting polypeptide described herein It may contain fragments.

[0056] Similarly, the O-RS segment contained in the fusion protein of the invention is derived from an O-RS. and is a segment of a fusion protein that functions in the context of the fusion protein. Therefore, the fusion protein has O-RS enzymatic activity, i.e., the activating activity of O-tRNA by ncAA. The O-RS segment thus confers the ability to catalyze aminoacylation. The amino acid sequence may include the entire amino acid sequence or a functional fragment of the O-RS described herein.

[0057] The assembler fusion protein (AFP) described herein is a fusion protein that is a fusion protein that is a fusion protein of the assembler (AP) and As used herein, the term "protein" refers to a polypeptide that comprises at least one polypeptide segment that acts as a polypeptide. The term AP refers to the division of AFPs containing said segments at spatially distinct sites within living cells. For convenience, the term "spatially distinct" refers to any polypeptide segment that allows for enrichment. Such sites are located within or directly adjacent to the cytoplasm of a cell and are accessible to the translation machinery of the cell ( This includes standard aminoacylated tRNAs, translation factors, ribosomal subunits, etc. These are readily accessible to ncA-tRNA and O-tRNA. Allows the introduction of A residues.

[0058] There are different types of polypeptide segments that function as APs in the present invention. One type of P is derived from an intracellular targeting polypeptide (IC-TP) and is a fusion protein. These IC-TP segments are polypeptide segments that function in the context of proteins. The compound may comprise the amino acid sequence of IC-TP (entire sequence) or a functional fragment. C-TP targets intracellular structural elements within or directly adjacent to the cytoplasm, thus Examples of such structural elements are microtubules, cell membranes, nuclear membranes, mitochondria, and the like. This includes the cytoplasmic sides of membranes such as the chondriac membrane, Golgi membrane, and ER membrane.

[0059] Thus, in a particular embodiment, the fusion proteins of the invention are capable of binding to microtubules, particularly microtubule plaques. The present invention is directed to targeting the positive or negative ends of fusion proteins and facilitating localized concentration of the fusion proteins therein. For example, dynein and kinesin ( Proteins of the dynein or kinesin family and their functional fragments Fragments and variants thereof can be used as IC-TPs for such functions.

[0060] In a further specific embodiment, the fusion protein of the invention is derived from a membrane anchor and functional For example, the fusion protein of the present invention comprises at least one IC-TP segment that functions as a fusion protein. It targets and initiates fusion at the (inner) cell membrane (especially the cytoplasmic side of the cell membrane). At least one IC-TP segment that facilitates localized concentration of the protein. In an example, the fusion protein of the invention targets the (outer) nuclear membrane (particularly the cytoplasmic side of the nuclear membrane). and at least one IC-TP segment that facilitates localized enrichment of the fusion protein. In a further particular embodiment, the fusion protein of the invention comprises an outer mitochondrial membrane Targeting (especially the cytoplasmic side of the mitochondrial membrane) and localized concentration of the fusion protein In a further particular embodiment, the IC-TP segment includes at least one IC-TP segment that facilitates The fusion protein of the invention targets the outer ER membrane (particularly the cytoplasmic side of the ER membrane) and Contains at least one IC-TP segment that facilitates localized concentration of the fusion protein In a further particular embodiment, the fusion protein of the invention is expressed in the outer Golgi membrane (particularly the Golgi membrane). At least one targeting agent that targets the fusion protein to the cytoplasmic side of the fusion protein and facilitates local concentration of the fusion protein. It contains two IC-TP segments. For example, membrane proteins and their functional fragments The transmembrane domains of the mutants and mutants of the mutants can be used as IC-TPs for such functions. can.

[0061] Polypeptides that target and therefore are locally concentrated in the intracellular structural elements described above. Lipeptides are known in the art and are useful as IC-TPs in the present invention. Specific examples of IC-TP include: - In living cells, they constitutively move toward and are locally concentrated at the plus ends of microtubules. Optionally, a truncated kinesin polypeptide, such as an optionally truncated kinesin family member. - member 16B (KIF16B), e.g., an optionally truncated form of Homo sapiens KIF16B (Uniprot: Q96L93), in particular the amino acid sequence of SEQ ID NO: 20 A fragment covering KIF16B amino acid residues 1-400 (KIF16B 1 -400 or optionally a truncated form of kinesin family member 13A (KIF13A ), e.g., the arbitrarily truncated Homo sapiens KIF13A (Uniprot :Q9H1H9), in particular, an adenosine triphosphate (APT) deletion vector having the amino acid sequence of SEQ ID NO:22. A KIF13A fragment covering amino acid residues 1-411 (KIF13A 1-411 ,Δ390 ); the microtubule tip-binding protein polypeptide EB1, which binds growing microtubules Binds to the plus end (Nehlig A, Molina A, Rodrigues-Fer reira S, Honore S, Nahmias C. Regulation o f end-binding protein EB1 in the control of microtubule dynamics.Cell Mol Life S ci.2017;74(13):2381-2393.doi:10.1007 / s00 018-017-2476-2) (Uniprot:Q15691), hence organelle to the microtubule plus end, and comprising the amino acid sequence of SEQ ID NO: 302; - For example, optionally a truncated form of translocase of the outer mitochondrial membrane 20 (TOMM20) A polypeptide that targets the outer mitochondrial membrane derived from any transmembrane protein, e.g. , an arbitrarily truncated version of Homo sapiens TOMM20 (Uniprot:Q153 88). In particular, amino acid residues 1-70 of TOMM20, which comprises the amino acid sequence of SEQ ID NO:24 A fragment covering TOMM20 1-70 ); - For example, lymphocyte-specific protein tyrosine kinase (LcK; e.g., Mus mus culus LcK, Uniprot:P06240), CD4 (e.g. Mus mus culus CD4, Uniprot:P06332), FRB(Homo sapie) ns mTOR-similar; Uniprot:P42345), CD28 (e.g., Mus musculus CD28, Uniprot:P31041), and combinations thereof In particular, a cell membrane targeting polypeptide derived from a transmembrane protein such as SEQ ID NO: 26, a polypeptide comprising the amino acid sequence of SEQ ID NO:28 or SEQ ID NO:30; - Nucleoporin polypeptide CG1 (Fern andez-Martinez J, Kim SJ Shi Y, et al. Str. ucture and Function of the Nuclear Pore Complex Cytoplasmic mRNA Export Platform .Cell.2016;167(5):1215-1228.e25.doi:10.1 016 / j.cell.2016.10.028)(also called Nup42)( Uniprot: O15504). Targets the cytoplasmic side of the nuclear envelope. Contains the amino acid sequence; - EBAG9 (E ngelsberg A, Hermosilla R, Karsten U, Schul ein R, Dorken B, Rehm A. The Golgi protein RCAS1 controls cell surface expression f tumor-associated O-linked glycan antigen ens.J Biol Chem.2003;278(25):22998-23007 .doi:10.1074 / jbc.M301361200(Uniprot:O005 59) Targets the cytoplasmic side of the Golgi membrane. Contains the amino acid sequence of SEQ ID NO: 292 (full length). or comprising the first 29 N-terminal amino acid residues of SEQ ID NO:294; or comprising 10 Polypeptide C of the CMP sialic acid transporter, a Golgi protein with 1 transmembrane helix MP Sia Tr(Eckhardt M,Gotza B,Gerardy-Sch ahn R.Membrane topology of the mammalian CMP-sialic acid transporter.J Biol Chem .1999;274(13):8779-8787. doi:10.1074 / jbc .274.13.8779) (Uniprot:P78382). Targeting. Comprising the amino acid sequence of SEQ ID NO: 296; - Polypeptide fragment of endoplasmic reticulum resident protein P450 2C1 (Faz al FM, Han S, Parker KR, et al. Atlas of Sub cellular RNA Localization Revealed by AP EX-Seq.Cell. 2019;178(2):473-490.e26.doi :10.1016 / j.cell.2019.05.027)(Uniprot:P78 382). It targets the cytoplasmic side of the ER membrane. In particular, the first 27 amino acids at the N-terminus (SEQ ID NO: 298 ) or a fragment comprising the first 29 (SEQ ID NO:300) amino acid residues; - the transmembrane protein stomatin-like protein 3 (SLP-3) (A of SEQ ID NO: 310) Membrane containing amino acid sequence; aa 1-59 (Homo sapiens, Uniprot:Q 8TAV4) It is localized in the plasma membrane and vesicle membrane (Lapatsina L, Jira J) A,Smith ES,et al.Regulation of ASIC chan nels by a stomatin / STOML3 complex locate d in a mobile vesicle pool in sensory ne urons.Open Biol.2012;2(6):120096.doi:10. 1098 / rsob.120096); and functional fragments and variants of these polypeptides, but are not limited to these. Such functional fragments and variants are not limited to the amino acid sequence of the polypeptide from which they are derived. amino acids and at least 60%, at least 70%, at least 80%, at least 90%, At least 95%, at least 96%, at least 97%, at least 98%, or less The amino acid sequence may have at least 99% amino acid sequence identity.

[0062] A further type of AP is derived from phase-separated polypeptides (PSPs) and is a fusion protein. PSPs are polypeptide segments that function in a context. A polymer having the ability to self-assemble in the cytoplasm of a cell to create a desired concentration of sites. Specifically, PSPs can drive phase separation (especially liquid-liquid phase separation). This leads to the formation of membrane-free compartments in the cytoplasm. These compartments can be droplets, aggregates, condensates, or In particular, PSPs are important classes of proteins that drive phase separation. These include intrinsically disordered proteins (IDPs), which are l., Bioessays 2016,38:959-968 and references therein. See, e.g., Patel et al., Cell 2015, 162:1066 -1077;Han et al.,Cell 2012,149:768-779;K (See Ato et al., Cell 2012, 149:753-767). There are three different classes. Each protein or its functional fragment or One well-known IDP is a PSP, and its variants can be used as PSPs in the present invention. The class contains so-called prion-like domains. These are uncharged and interspersed. It contains polar amino acid residues (Q, N, S, G) with aromatic residues (F, Y). For example, Malinovska et al., Biochim Biophys Acta 2013,1834:918-931,Alberti et al.,2009,C ell 137:146-158, Malinovska et al., Prion 2015, 9:339-346. Another class of IDPs is also characterized by low sequence complexity. They often contain both acidic and basic amino acid side chains. It is an RGG repeat-containing IDP. Nott et al., Cell 2015, 57 Specific examples of suitable IC-TPs are: - Spindle defect protein 5 (SPD5) (e.g., Caenorhabditis e legans SPD5; Uniprot: P91349). In particular, the amine a polypeptide comprising a nucleotide sequence; - Fusion sarcoma (FUS) (e.g., Homo sapiens FUS; Uniprot :P35637). In particular, a polypeptide comprising the amino acid sequence of SEQ ID NO: 34; - Ewing sarcoma breakpoint region 1 (EWSR1) (e.g., Homo sap iens EWSR1; Uniprot: Q01844). In particular, the amino acid sequence of SEQ ID NO: 36 a polypeptide comprising an acid sequence; - ATP-dependent RNA helicase laf-1 (RGG domain, 1-168, LAF- 1 membrane. Caenorhabditis elegans (Caenorhabditis elegans) gans, Uniprot:D0PV95)(Schuster BS, Reed EH ,Parthasarathy R,et al.Controllable prot. ein phase separation and modular recruitment ment to form responsive membraneless org anelles.Nat Commun.2018;9(1):2985.Publis hed 2018 Jul 30.doi:10.1038 / s41467-018-0 5403-1); and functional fragments and variants of these polypeptides, but are not limited to these. Not determined. Such functional fragments and variants have at least one amino acid sequence similar to that of the polypeptide from which they are derived. at least 60%, at least 70%, at least 80%, at least 90%, at least 9 5%, at least 96%, at least 97%, at least 98%, or at least 99% The amino acid sequence identity may include

[0063] The number of APs contained in the fusion protein of the present invention is not particularly limited. Proteins may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more of the same or different At least one AP and P selected from the IC-TP segment may be included. and at least one AP selected from the SP segment. It is particularly preferred. Similarly, the number of RNA-TP segments is not particularly limited and may be independently 1, 2, 3, 4, 5, or more, such as 6, 7, 8, 9, or 10, different Alternatively, the O-RS segment may be selected from the same RNA-TP segment. The number of components is not particularly limited and may be independently 1, 2, 3, 4, 5, or, for example, 6, 7, 8, 9. , or more, such as 10, different or the same O-RS segments to select from. This applies to both AFP and RNA-TP / O-RS fusion proteins. Of course, the number of segments on the fusion protein of the present invention may vary depending on the size of the fusion protein. This affects the size of the protein, which is typically, but not limited to, 3500 amino acid residues. For example, less than 3000 amino acid residues.

[0064] The order of the segments in the fusion protein of the present invention is also not particularly limited. The NA-TP, O-RS, and / or AP segments may be functionally linked in any order. The RNA-TP / O-RS fusion protein structure (both EP segments) Examples of types of [RNA-TP] x -[O-RS] y [O-RS] y -[RNA-TP] x wherein x and y are each independently 1, 2, 3, is an integer selected from: 4, and 5; "-" indicates a peptidic linkage. [RNA-TP] for x ≥ 2 x The RNA-TP sequences may contain the same or different RNA-TP segments. [O-RS] for y≧2 y may contain the same or different O-RS segments. It is possible.

[0065] An example of an RNA-TP / O-RS fusion protein structure is: [IC-TP] m -[EP] o [EP] o -[IC-TP] m [PSP] n -[EP] o [EP] o - [PSP] n [IC-TP] m -[EP] o - [PSP] n [PSP] n -[EP] o -[IC-TP] m [IC-TP] m - [PSP] n -[EP] o [EP] o - [PSP] n -[IC-TP] m [PSP]n -[IC-TP] m -[EP] o [EP] o -[IC-TP] m - [PSP] n In the formula, m, n, and o are each independently 1, 2, , 3, 4, or 5, or an integer selected from 1, 2, 3, 4, 5, 6 and "-" indicates a peptidic linkage.

[0066] In a preferred embodiment, "m" is the integer 1.

[0067] In another preferred embodiment, "n" is an integer selected from 1 and 2.

[0068] In yet another preferred embodiment, "o" is when EP is selected from RNA-TP. In some cases, it is an integer selected from 1, 2, 3, 4, 5, or 6.

[0069] In yet another preferred embodiment, "o" is selected from when EP is selected from O-RS. is an integer selected from 1 or 2.

[0070] In yet another preferred embodiment of the RNA-TP / O-RS fusion protein structure, at least Preferably, at least one ICT-TP occupies the C- or N-terminal position in the polypeptide chain. stomach.

[0071] In yet another preferred embodiment of the RNA-TP / O-RS fusion protein structure, at least It is preferred that one EP occupies the C- or N-terminal position on the polypeptide chain.

[0072] In yet another preferred embodiment of the RNA-TP / O-RS fusion protein structure, At least one ICT-TP occupies the C- or N-terminal position on the polypeptide chain. Preferably, at least one EP takes the N- or C-terminal position on the polypeptide chain. If any PSP is present in such a structure, it is located within the polypeptide chain. do.

[0073] [IC-TP] with m≧2 m may contain the same or different IC-TP segments Preferably, the same functional groups (e.g., the same membrane type or organelle type) are used. IC-TP is applied to target the same type of cellular structure. [PSP] with n ≥ 2 n teeth , may contain the same or different PSP segments. o ≧2 [EP] o teeth, It may contain the same or different EPs. [EP] o where different EPs are included In the example, at least one EP is an RNA-TP segment and at least one may be an O-RS segment.

[0074] The fusion proteins of the present invention provide an orthogonal translation (OT) system, and can be used to insert one or more ncDNAs onto a POI. One or more O-RS (segments) required for the introduction of AA residues should be at least 1 The RNA-targeting polypeptide (RNA-TP) segments are placed in close spatial proximity to each other. The POI mRNA contains at least one RNA-TP segment of the OT-based fusion protein. At least one targeting nucleotide sequence (TN) capable of interacting with The interaction is expediently specific. The P segment is preferably an mRNA targeting polypeptide segment. The RNA-TP segment of the protein and the TN of the POI mRNA are, for convenience, For this purpose, the RNA-TP segment is selected to interact with (bind to) A suitable pair of MENT and TN is a coat protein of an RNA virus and a coat protein of the and a nucleic acid motif bound by a protein. Proteins and RNA motifs bound by proteins are known in the art. .

[0075] Specific examples of suitable RNA-TPs are: - MCP (coat protein of the Enterobacteriaceae phage MS2), in particular the amino acid sequence of SEQ ID NO: 14 a polypeptide comprising a nucleotide sequence; - λ N22 (22 amino acid RNA of lambda phage antiterminator protein N) In particular, a polypeptide comprising the amino acid sequence of SEQ ID NO: 16; - PCP (coat protein of bacteriophage PP7) Wu B, Chao JA ,Singer RH. Fluorescence fluctuation spe ctroscopy enables quantitative imaging f single mRNAs in living cells. Biophys J.2012;102(12):2936-2944.doi:10.1016 / jb pj.2012.05.017). In particular, the polypeptide comprising the amino acid sequence of SEQ ID NO: 306 Chid; and functional fragments and variants of these polypeptides, but are not limited to these. Such functional fragments and variants are not limited to the amino acid sequence of the polypeptide from which they are derived. amino acids and at least 60%, at least 70%, at least 80%, at least 90%, At least 95%, at least 96%, at least 97%, at least 98%, or less The amino acid sequence may have at least 99% amino acid sequence identity.

[0076] Specific examples of suitable TNs are: - the enterobacteriaceae phage MS2 RNA stem loop, in particular the nucleotide sequence of SEQ ID NO: 17 A polynucleotide having an RNA sequence that corresponds to (is encoded by) a DNA sequence. Ochid; - BoxB (lambda phage RNA stem loop. N22 Specific binding site of In particular, it corresponds to the nucleotide (DNA) sequence of SEQ ID NO: 18 (thereby coding a polynucleotide having an RNA sequence; - Bacteriophage pp7 RNA stem loop (Wu B, Chao JA, Si nger RH. Fluorescence fluctuation spectr oscopy enables quantitative imaging of s ingle mRNAs in living cells.Biophys J.20 12;102(12):2936-2944.doi:10.1016 / j.bpj.2 In particular, the nucleotide sequence of SEQ ID NO: 289 or SEQ ID NO: 290 (D A polynucleotide having an RNA sequence corresponding to (encoded by) an RNA sequence D; and functional fragments and variants thereof. Such functional fragments and variants have at least the same structure as the polynucleotide sequence from which they are derived. At least 60%, at least 70%, at least 80%, at least 90%, at least 95% %, at least 96%, at least 97%, at least 98%, or at least 99% It may include nucleotide sequence identity.

[0077] Such TN may be present as a single copy segment or as more than one, e.g., two, copies of TN. Multi-copy segments consisting of 3, 4, 5, 6 or more repeat units. It can be used as a

[0078] MCP specifically interacts with the MS2 RNA stem loop. The RNA-TP segment(s) of the protein(s) is derived from the MCP. Where the mRNA of the POI comprises a functional segment, the mRNA of the POI may be conveniently On the one or more MS2 RNA stem loops, for example 2, 3, 4, 5, or 6 MS2 Contains an RNA stem loop. N22 specifically interacts with BoxB. Therefore, The RNA-TP segment(s) of the fusion protein(s) is / are: N2 2 POI is derived from and contains (or consists of) a functional segment Conveniently, the mRNA of It contains six or more BoxB motifs. PCP specifically interacts with the pp7 RNA stem loop. Therefore, the RNA-TP segment (single) of the fusion protein (single or multiple) The fragment or fragments of the polypeptide (or fragments) are derived from and comprise functional segments of PCP. In the present specification, the mRNA of the POI is conveniently prepared by integrating one or more pp7 RNA stem loops, e.g., 2 , 3, 4, 5, or 6 or more pp7 RNA stem loops.

[0079] Several RSs have been used for genetic code expansion, including Methanococc us jannaschii tyrosyl-tRNA synthetase, E. coli tyrosyl-tR NA synthetase, E. coli leucyl-tRNA synthetase, certain Methan osarcina (e.g. M.mazei, M.barkeri, M.acetivor ans, M. thermophila), Methanococcoides (M. bu rtonii), or Desulfitobacterium (D. hafniense ) and the corresponding orthogonal RS / tRNA pair have been used to genetically code for various functionalities on polypeptides (Ch in,Annu Rev Biochem 2014,83:379-408;Chin et al.,J Am Chem Soc 2001,124:9026;Chin et al.,Science 2003,301:964;Nguyen et a l.,J Am Chem Soc 2009,131:8720;Yanagisaw a et al., Chem Biol 2008,15:1187). Depending on the cells used, these RSs can be used as O-RSs in the present invention. can.

[0080] Pyrrolysyl-tRNA synthetase that can be used in the methods and fusion proteins of the present invention The enzyme (PylRS) may be a wild-type or a genetically engineered PylRS. Examples of PylRS are from archaea and eubacteria, e.g., Methanosarcina mai ze, Methanosarcina barkeri, Methanococcoid es burtonii, Methanosarcina acetivorans, M ethanosarcina thermophila and Desulfitobac PylRS from terium hafniense, but is not limited to The engineered PylRS has been described, for example, by Neumann et al. Chem Biol 2008,4:232) by Yanagisawa et al. al. (Chem Biol 2008,15:1187) and EP2192 The efficiency of genetic code expansion using PylRS has been reported in The amino acid sequence of PylRS was modified so that it was not directed against the For this purpose, the nuclear localization signal (NLS) was removed from PylRS. or can be overlaid by introducing a suitable nuclear export signal (NES). The pylRS used in the fusion proteins and methods of the present invention can be, for example, For example, those lacking an NLS and / or containing an NES as described in WO2018 / 069481. It may also be pylRS.

[0081] Thus, the O-RS segments (single or multiple) that can be used in the fusion proteins of the invention Examples of "plural" are: Methanococcus jannaschii tyrosyl-tRNA synthetase Escherichia coli tyrosyl-tRNA synthetase Escherichia coli leucyl-tRNA synthetase Methanosarcina mazeii pyrrolysyl-tRNA synthetase Methanosarcina barkeri pyrrolysyl-tRNA synthetase Methanosarcina acetivorans pyrrolysyl-tRNA synthesizer Tase Methanosarcina thermophila pyrrolysyl-tRNA synthesizer Tase Methanococcoides burtonii pyrrolysyl-tRNA syntheta -ze Desulfitobacterium hafniense pyrrolysyl-tRNA synthase cetanase and functional (i.e., enzymatically active) fragments of these polypeptides. The functional fragments and variants include, but are not limited to, those At least 60%, at least 70%, and %, at least 80%, at least 90%, at least 95%, at least 96%, at least at least 97%, at least 98%, or at least 99% amino acid sequence identity. Good too.

[0082] The O-transferase useful in the present invention is derived from M. mazei pyrrolysyl-tRNA synthetase. Specific examples of RS segments are: - PylRS AF (Methanosarcina mazei pyrrolysyl-tRNA synthase The double mutant cisplatinase: Y306A, Y384F; Uniprot: Q8PWY1) For example, an O-RS segment comprising the amino acid sequence of SEQ ID NO:8. to; - PylRS AA (Methanosarcina mazei pyrrolysyl-tRNA synthase The double mutant cisplatinase (N346A, C348A; Uniprot: Q8PWY1) For example, an O-RS segment comprising the amino acid sequence of SEQ ID NO: 10. nt; - PylRS AAAF (Methanosarcina mazei pyrrolidyl tRN A synthetase quadruple mutant: Y306A, N346A, C348A, Y384F; Uni prot:Q8PWY1). For example, the amino acid sequence of SEQ ID NO:12. an O-RS segment containing an amino acid sequence; - Methanosarcina mazei pyrrolysyl tRNA mutant (L305M , Y306L, L309S, N346S, C348M) O-RS cell derived from IFRS1 For example, an O-RS segment comprising the amino acid sequence of SEQ ID NO: 224; - Methanosarcina mazei pyrrolysyl tRNA mutant (Y306M , L309G, C348T) O-RS segment derived from CbzRS. For example, SEQ ID NO: an O-RS segment containing the amino acid sequence of No. 226; - Methanosarcina mazei pyrrolysyl tRNA mutant (A302S ) an O-RS segment derived from CpkRS. For example, the amino acid sequence of SEQ ID NO: 228 Includes O-RS segments; - Methanosarcina mazei pyrrolysyl tRNA mutant (A302T , Y384F, N346V, C348W, V401L) O-RS cells derived from OMeRS For example, an O-RS segment comprising the amino acid sequence of SEQ ID NO: 236; As well as functional (i.e., enzymatically active) flags of these polypeptide segments. The functional fragments and variants of the present invention include, but are not limited to, Variants should be at least 60% identical to the aminoacyl-tRNA synthetase from which they are derived, and at least at least 70%, at least 80%, at least 90%, at least 95%, at least 9 6%, at least 97%, at least 98%, or at least 99% amino acid sequence identity It may include gender.

[0083] According to certain embodiments, the wild-type and mutant M. mazei P ylRS is described in WO2012 / 104422 or WO2015 / 107064 It is used for aminoacylation of tRNA with ncAA. For this purpose, examples An exemplary ncAA is 2-amino-6-(cyclooct-2-yn-1-yloxycarbo 2-amino-6-(cyclooct-2-yn-1-ylamino)hexanoic acid (SCO), 2-amino-6[(4E-cyclohexyloxyethoxycarbonylamino)hexanoic acid 2-(oct-4-en-1-yl)oxycarbonylamino]hexanoic acid (TCO), Amino-6[(2E-cyclooct-2-en-1-yl)oxycarbonylamino]he TCO*, 2-amino-6-(prop-2-yloxycarbonylamino) Hexanoic acid (PrK), and 2-amino-6-(9-biocyclo[6.1.0]nona-4 -ynylmethoxycarbonylamino)hexanoic acid (BCN), Not limited.

[0084] In another embodiment of the present invention, the above-mentioned AP (IC-TP and PSP) segmentation The segments and / or EPs (RNA-TP and O-RS) referred to above are Independently, furthermore, natural or more specific molecules that induce and control macromolecular interactions are In particular, such additional proteins may be combined with synthetic protein segments. The segments are operably fused into the polypeptide chain of the AFP of the present invention. One or more, such as 5, 6, 7, 8, 9, or 10, but preferably one or more The protein segments can be functionally fused to the AFP vector of the present invention. Fusion onto the FP polypeptide chain allows the activities of the other polypeptide segments AP and EP to be realized. Qualitatively unaffected, in particular uninhibited (i.e., AP and EP remain operational) The ability of additional polypeptide segments to direct and control macromolecular interactions The force is maintained. The literature describes so-called SYNZIP peptides that form multimeric structures. It has been reported that the ribosomal protein has the ability to form specific heterodimeric coiled-coil protein structures. Of particular interest in the context of the present invention are SYNZIPs that A pair of synthetic peptides that can interact with each other and induce and control macromolecular interactions. A non-limiting example is SYNZIP 1:2;SYNZIP 3:4; and SYNZIP 5:6 pairs. .,Grant,RA,KeatingA.E.(2010)J Am Chem. Heterospecific coiled coyote described by Soc 132 6025-6031 SYNZIP2:SYNZIP1 (SYNZIP1: sequence number 312, SYNZI P2: SEQ ID NO: 314, SYNZIP3: SEQ ID NO: 316, SYNZIP4: SEQ ID NO: 3 18, and functional fragments and variants of those SYNZIP polypeptides are particularly Preferably, the functional fragments and variants are amino acids of the polypeptide from which they are derived. At least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or (The amino acid sequence may contain 99% amino acid sequence identity with the target molecule.) Since the use of SYNZIPs as a pair is required, these SYNZIPs are preferably On different AFP fusion proteins The interaction of such a SYNZIP pair integrated in The formation of OT organelles can be further supported.

[0085] In yet another embodiment of the invention, the fusion protein of the invention further comprises an expressed At least one so-called "epitope tag" useful for detecting / quantifying the fusion product. ", i.e., by the introduction (fusion) of a short oligopeptide sequence that serves as an antibody binding site. Non-limiting examples of such tags are: VSV-G: vesicular stomatitis virus glycoprotein epitope tag (SEQ ID NO: 680) HA: human influenza hemagglutinin epitope tag (sequence Number 682) Myc: human c-Myc proto-oncogene epitope tag (SEQ ID NO: 684)

[0086] 1.2 Specific Examples of AFP Constructs of the Invention

[0087] Each individual exemplified construct can be interpreted in the N->C or C->N orientation. The scheme shown is given in the N->C direction.

[0088] Segment block where m, n, y, or x is an integer > 1 [IC-TP] m , [PSP ] n , [O-RS] y , and [RNA-TP] x In this case, the repetition within the block The return segments may be the same or different, preferably they are the same.

[0089] The applicable segments [IC-TP], [PSP], [O-RS], and [R NA-TP] x , and [SYNZIP], described above in Section 1.1 It can be prepared from each of the examples of the ment.

[0090] 1.2.1. Monofunctional AFP targeting intracellular structures

[0091] 1.2.1.1 Subcellular structure-targeted monofunctional AFPs (i.e., containing one type of EP)

[0092] Preferred examples of these are:

[0093] [IC-TP] m -[O-RS] y where m=1 or 2, preferably 1; y=1 or is 2, preferably 1;

[0094] [IC-TP] m -[RNA-TP] x where m=1 or 2, preferably 1; x= 1, 2, 3, 4, 5, or 6, preferably 2, 3, or 4;

[0095] [IC-TP] m - [PSP] n -[O-RS] y where m=1 or 2, preferably is 1; n=1, 2, or 3, preferably 1 or 2; y=1 or 2, preferably 1;

[0096] [IC-TP] m - [PSP] n -[RNA-TP] x Wherein m=1 or 2, preferably or 1; n=1, 2, or 3, preferably 1 or 2; x=1, 2, 3, 4, 5, or 6 , preferably 2, 3, or 4;

[0097] [IC-TP] m -[O-RS 1 ] y - [PSP] n -[O-RS 2 ] y Where m = 1 or 2, preferably 1; n = 1, 2, or 3, preferably 1 or 2; independently of each other y=1 or 2, preferably 1; O-RS 1 and O-RS 2 are the same or different, preferably are identical;

[0098] [IC-TP] m - [PSP 1 ] n -[O-RS] y - [PSP 2 ] n Where m= n is independently 1, 2, or 3, preferably 1 or 2; Independently, y=1 or 2, preferably 1; PSP 1 and PSP 2 are the same or different ;

[0099] [IC-TP] m -[RNA-TP 1 ] x - [PSP] n -[RNA-TP 2 ] x Yes where m=1 or 2, preferably 1; n=1, 2, or 3, preferably 1 or 2; Independently, x=1, 2, 3, 4, 5, or 6, preferably 2, 3, or 4; RNA-TP 1 and RNA-TP 2 are the same or different, preferably the same;

[0100] [IC-TP] m - [PSP 1 ] n -[O-RS 1 ] y - [PSP 2 ] n -[O-RS 2 ] y where m=1 or 2, preferably 1; n are each independently 1, 2 or 3, preferably Preferably 1 or 2; independently of each other y=1 or 2, preferably 1; O-RS 1 and OR S 2 are the same or different, preferably the same; PSP 1 and PSP 2 are the same or different;

[0101] [IC-TP] m - [PSP 1 ] n -[RNA-TP 1 ] x - [PSP 2 ] n -[RN A-TP 2 ] x where m=1 or 2, preferably 1; and independently of each other, n=1, 2, or is 3, preferably 1 or 2; independently of each other, x=1, 2, 3, 4, 5, or 6, preferably 2, 3, or 4; RNA-TP 1 and RNA-TP 2 are the same or different; PSP 1 Reach and PSP 2 are the same or different.

[0102] 1.2.1.2 Bifunctional AFPs targeted to intracellular structures (including two forms of EP)

[0103] Individually preferred examples are:

[0104] [IC-TP] m -[O-RS] y -[RNA-TP] x Where m=1 or 2, preferably x=1, 2, 3, 4, 5, or 6, preferably 2, 3, or 4; y=1 or 2, preferably 1;

[0105] [IC-TP] m -[RNA-TP] x -[O-RS] y Where m=1 or 2, preferably x=1, 2, 3, 4, 5, or 6, preferably 2, 3, or 4; y=1 or 2, preferably 1;

[0106] [IC-TP] m - [PSP] n -[O-RS] y -[RNA-TP] x Where m = 1 or 2, preferably 1; n = 1, 2, or 3, preferably 1 or 2; x = 1, 2, 3 , 4, 5, or 6, preferably 2, 3, or 4; y=1 or 2, preferably 1;

[0107] [IC-TP] m - [PSP] n -[RNA-TP] x -[O-RS] y Where m = 1 or 2, preferably 1; n = 1, 2, or 3, preferably 1 or 2; x = 1, 2, 3 , 4, 5, or 6, preferably 2, 3, or 4; y=1 or 2, preferably 1;

[0108] [IC-TP] m -[O-RS] y - [PSP] n -[RNA-TP] xWhere m = 1 or 2, preferably 1; n = 1, 2, or 3, preferably 1 or 2; x = 1, 2, 3 , 4, 5, or 6, preferably 2, 3, or 4; y=1 or 2, preferably 1;

[0109] [IC-TP] m -[RNA-TP] x - [PSP] n -[O-RS] y Where m = 1 or 2, preferably 1; n = 1, 2, or 3, preferably 1 or 2; x = 1, 2, 3 , 4, 5, or 6, preferably 2, 3, or 4; y=1 or 2, preferably 1;

[0110] [IC-TP] m - [PSP 1 ] n -[O-RS] y - [PSP 2 ] n -[RNA-T P] x where m=1 or 2, preferably 1; n is independently of each other, n=1, 2 , or 3, preferably 1 or 2; x=1, 2, 3, 4, 5, or 6, preferably 2, 3, or 4; y=1 or 2, preferably 1; PSP 1 and PSP 2 are the same or different;

[0111] [IC-TP] m - [PSP 1 ] n -[RNA-TP] x -[O-RS 1 ] y - [PS P 2 ] n -[O-RS 2 ] y where m=1 or 2, preferably 1; where n=1, 2, or 3, preferably 1 or 2; x=1, 2, 3, 4, 5, or 6, preferably or 2, 3, or 4; independently of each other, y=1 or 2, preferably 1; PSP 1 and P.S. P 2 are the same or different; O-RS 1 and O-RS 2 are the same or different, preferably the same one;

[0112] [IC-TP] m - [PSP 1 ] n -[O-RS 1 ] y - [PSP 2 ] n -[O-RS 2 ] y -[RNA-TP] x where m=1 or 2, preferably 1; where n=1, 2, or 3, preferably 1 or 2; x=1, 2, 3, 4, 5, or 6, preferably or 2, 3, or 4, independently of each other y=1 or 2, preferably 1; PSP 1 and P.S. P 2 are the same or different; O-RS 1 and O-RS 2 are the same or different, preferably the same It is one.

[0113] 1.2.2. No monofunctional AFP targeted to intracellular structures

[0114] These are sequences that lack the segment [IC-TP] but retain the segment [PSP]. The same A listed above in Section 1.2.1, with the only exception that It is FP.

[0115] SYNZIP variants

[0116] These are the segments [IC-TP], [PSP], [O-RS 2 ], or [RNA- TP] is supplemented with a SYNZIP element at the N- or C-terminus. Except for the same AFPs listed in sections 1.2.1 and 1.2.2 above. The AFP may be one, two, three, four or five, preferably one or two, identical or different Non-limiting examples of such molecules include: This is:

[0117] 1.2.3.1 Monofunctional SYNZIP AFP

[0118] Preferred examples of these are:

[0119] [PSP] n -[SYNZIP]-[O-RS] y where y=1 or 2, preferably is 1; n=1, 2, or 3, preferably 1 or 2;

[0120] [PSP] n -[SYNZIP]-[RNA-TP] x where x=1, 2, 3, 4 , 5, or 6, preferably 2, 3, or 4; n=1, 2, or 3, preferably 1 or 2;

[0121] [IC-TP] m -[SYNZIP]-[O-RS] y Wherein m=1 or 2, preferably or 1; y=1 or 2, preferably 1;

[0122] [IC-TP] m -[SYNZIP]-[RNA-TP] x Wherein m=1 or 2; preferably 1; x=1, 2, 3, 4, 5, or 6, preferably 2, 3, or 4;

[0123] [IC-TP] m - [PSP] n-[SYNZIP]-[O-RS] y Where m= n=1 or 2, preferably 1; n=1, 2, or 3, preferably 1 or 2; y=1 or 2, preferably Preferably 1;

[0124] [IC-TP] m - [PSP] n -[SYNZIP]-[RNA-TP] x And, m=1 or 2, preferably 1; n=1, 2, or 3, preferably 1 or 2; x=1, 2, 3, 4, 5, or 6, preferably 2, 3, or 4.

[0125] 1.2.3.2 Bifunctional SYNZIP AFP

[0126] Preferred examples of these are:

[0127] [IC-TP] m -[O-RS] y -[SYNZIP]-[RNA-TP] x And , m=1 or 2, preferably 1; x=1, 2, 3, 4, 5, or 6, preferably 2, 3, or 4; y=1 or 2, preferably 1;

[0128] [IC-TP] m -[RNA-TP] x -[SYNZIP]-[O-RS] y And , m=1 or 2, preferably 1; x=1, 2, 3, 4, 5, or 6, preferably 2, 3, or 4; y=1 or 2, preferably 1;

[0129] [IC-TP] m - [PSP] n -[SYNZIP]-[O-RS] y -[RNA-T P] xwhere m=1 or 2, preferably 1; n=1, 2, or 3, preferably 1 or 2; x=1, 2, 3, 4, 5, or 6, preferably 2, 3, or 4; y=1 or 2, preferably 1;

[0130] [IC-TP] m - [PSP] n -[SYNZIP]-[RNA-TP] x -[OR S] y where m=1 or 2, preferably 1; n=1, 2, or 3, preferably 1 or 2; x=1, 2, 3, 4, 5, or 6, preferably 2, 3, or 4; y=1 or 2, preferably or 1;

[0131] [IC-TP] m - [PSP] n -[SYNZIP a ]-[O-RS] y -[SYNZ IP b ]-[RNA-TP] x where m=1 or 2, preferably 1; n=1, 2, or is 3, preferably 1 or 2; x=1, 2, 3, 4, 5, or 6, preferably 2, 3, or 4; y=1 or 2, preferably 1; SYNZIP a and SYNZIP b are the same or different and preferably identical;

[0132] [IC-TP] m - [PSP] n -[SYNZIP a ]-[RNA-TP] x -[SY NZIP b ]-[O-RS] y where m=1 or 2, preferably 1; n=1, 2, or is 3, preferably 1 or 2; x=1, 2, 3, 4, 5, or 6, preferably 2, 3, or 4;y=1 or 2, preferably 1;SYNZIP a and SYNZIP b are the same or different and preferably identical;

[0133] [IC-TP] m - [PSP 1 ] n -[SYNZIP]-[RNA-TP] x -[O- RS 1 ] y - [PSP 2 ] n -[O-RS 2 ] y where m=1 or 2, preferably 1 n=1, 2, or 3, preferably 1 or 2; x=1, 2, 3, 4, 5, or 6, preferably y=1 or 2, preferably 1; PSP 1 and PSP 2 are the same or Different;O-RS 1 and O-RS 2 are the same or different, preferably the same.

[0134] 1.2.4. Monofunctional fusion proteins

[0135] Preferred examples of these are:

[0136] [SYNZIP]-[O-RS], where y=1 or 2, preferably 1 y ;

[0137] [SYNZIP]-, where x=1, 2, 3, 4, 5, or 6, preferably 2, 3, or 4 [RNA-TP] x ;is.

[0138] Here, IC-TP and PSP are missing, so these are preferably at least and in combination with an AFP molecule containing another C-TP and / or PSP segment. There can be.

[0139] 1.3 Examples of individual fusion proteins

[0140] Highly specific examples of fusion proteins of the invention and specific combinations thereof are shown in the table below. The contents of Tables 1, 2, and 3 are also included in the general description of this specification. forms part of the disclosure, the contents of which are expressly and verbatim incorporated herein in the general part. "Fusion protein (single or multiple) containing an O-RS and an RNA-TP segment" The disclosures in Tables 1 and 2 in each column entitled "(plural)" refer to specific reports and The contents of the other columns in Tables 1 and 2 that refer to host cell lines are to be considered as independently disclosed. will be done.

[0141] 2. Functional Fragments and Mutants

[0142] The specific RNA-TP, O-RS, IC-TP, PSP, TN, and SY Fragments and variants of NZIP have been described that are functional (i.e., These results suggest that the parent RNA-TP has RNA-binding activity and the parent IC-TP has intracellular targeting activity. activity, self-assembly activity of parent PSP, binding activity of parent TN to RNA-TP, parent O- Possess the enzymatic activity of RS or the heterodimeric coiled-coil formation ability of the parent SYNZIP Such fragments and variants are characterized by the minimum degree of sequence identity described herein. The amino acid or nucleotide sequence identity can be expressed as follows: Identification over the entire length of the amino acid or nucleotide sequence characterized by Percentage identity values ​​are calculated using BLAST matching, as known in the art. Alignment, based on the blastp algorithm (protein-protein BLAST) or using the Clustal method (Higgins et al. l.,Comput Appl.Biosci.1989,5(2):151-1).

[0143] Specific RNA-TP, O-RS, IC-TP, SYNZIP, or PS Useful in the Invention Fragments and variants of P retain the relevant functions of the parent polypeptide (binding, self-binding, , or enzymatic activity) and, for example, conservative amino acid substitutions, i.e., those known in the art. with different amino acid residues that have similar biochemical properties (e.g., charge, hydrophobicity, and size) A typical example is the replacement of Leu by Ile or or vice versa; replacement of Asp by Glu or vice versa; replacement of Asn by Gln or vice versa; and others.

[0144] 3. Orthogonal translation, tRNA, and POI coding sequences

[0145] The term "translation system" generally refers to a system that is naturally present on a growing polypeptide chain (protein). The components of the translation system are the set of components required to incorporate the amino acid that is to be translated. Components include, for example, ribosomes, tRNAs, aminoacyl-tRNA synthetases, mRNAs Aminoacyl-tRNA synthetases (RS) can include the following: It is an enzyme that can aminoacylate tRNA with amino acids or amino acid analogs. The RS used in the process of the present invention amplifies tRNA with its corresponding ncAA. Noacylation, i.e., tRNA ncAA The present invention can be aminoacylated. As used herein, the term "orthogonal" refers to a translation system (e.g., a cell) that is reduced by the intended translation system. The elements of the translation system used in the efficiency (e.g., orthogonal tRNA (O-tRNA) and / or orthogonal tRNA (O-tRNA) "Orthogonal" refers to an O-tRNA synthetase (O-RS). Alternatively, the O-RS may function with the endogenous RS or endogenous tRNA of the target translation system, respectively. Inability to operate or reduced efficiency, e.g., less than 20% efficiency, 10% It means less than 5% efficiency, or less than 1% efficiency, for example. When compared to the aminoacylation of endogenous tRNAs by the target RS, O-tRNAs In the target translation system, the endogenous R In another example, an endogenous tRNA is aminoacylated by an endogenous RS. Compared to noacylation, O-RSs can mediate translation of a target protein with reduced or even zero efficiency. In particular, the term "orthogonal translation" refers to the aminoacylation of any endogenous tRNA. "OT system" or "OT system" as used herein refers to a system that does not incorporate ncAA residues on a growing polypeptide chain. O-RS / O-tRNA that allows the introduction of ncAA Refers to a translation system that uses pairs It is used to

[0146] O-RS / O-tRNA used in the present invention ncAA The pair preferably has the following characteristics: O-tRNA ncAA is preferentially aminoacylated with ncAA by O-RS In addition, O-tRNA ncAA The NCAA residue in the growing polypeptide chain of the POI The orthogonal pair is used to incorporate the nucleotide sequence into a translation system (e.g., a cell) of interest. Incorporation occurs in a site-specific manner. Specifically, O-tRNA ncA A The selector codon (e.g., amber, ochre, or recognizes the opal stop codon). The term "preferentially aminoacylates" refers to an aminoacylation of endogenous tRNAs in a translation system (e.g., a cell) of interest. Or, compared to the amino acid, the O-RS catalyzes the O-tRNA with the unnatural amino acid. The efficiency of the filter can be, for example, about 50% efficiency, about 70% efficiency, about 75% efficiency, about 85% efficiency, about It means that the unnatural amine is 90% efficient, about 95% efficient, or about 99% or more efficient. The acid has high fidelity, e.g., greater than about 75% efficiency for a given selector codon. with an efficiency of greater than about 80% for a given selector codon With greater than about 90% efficiency, with greater than about 95% efficiency for a given selector codon, or Lectacodons are incorporated into the growing polypeptide chain with an efficiency of greater than about 99% .

[0147] At least one O-RS derived from M. mazei pyrrolysyl-tRNA synthetase The fusion protein of the invention may be used to ammoacylate the nucleotide sequence of ... The tRNAs that can be used include pyrrolysyl-tRNA from M. mazei and its functional mutants. Anticodons include, but are not limited to, anticodons for a selector codon; For example, the CUA anticodon for the amber stop codon TAG, the opal stop codon The anticodon UCA for the stop codon TGA, and the anticodon TAA for the ochre stop codon The codon is UUA. An example of such a pyrrolysyl tRNA is SEQ ID NO: 4 (tRNA Pyl, CUA ), SEQ ID NO:5 (tRNA Pyl,UCA ), or SEQ ID NO: 6 (tRNA Pyl, UUA ) Further non-limiting examples of suitable tRNAs include pyrrolysyl-tRNA from M. mazei. It is derived from:

[0148] tRNA pyl,CGA Pyrrolysyl tRNA (serine codon), SEQ ID NO: 229 tRNA pyl,CGG Pyrrolysyl tRNA (proline codon), SEQ ID NO: 230 tRNA pyl,UAA Pyrrolysyl tRNA (leucine codon), SEQ ID NO: 231 tRNA pyl,UAG Pyrrolysyl tRNA (leucine codon), SEQ ID NO: 232 tRNA pyl,CCG Pyrrolysyl tRNA (arginine codon), SEQ ID NO: 233 tRNA pyl,AUA Pyrrolysyl tRNA (isoleucine codon), SEQ ID NO: 234 As used herein, the term "selector codon" refers to a codon that is selected from an O-tRNA during the translation process. n cAA The term refers to a codon that is recognized (i.e., bound) by a messenger. Polynucleotides that are not mRNA (e.g., polypeptides of DNA plasmids) The new OT sequences described herein are also used for the corresponding codons in the peptide coding sequence. The system is selective for the mRNA of the POI relative to other mRNAs present in the cytoplasm of the cell. The selector codon allows for orthogonal translation of the POI in a manner that is alternative. , codons with low abundance in the cell chosen for expression, e.g., naturally occurring eukaryotic The new OT system is preferably a codon with low abundance in the cell. mRNA, O-RS, and tRNA ncAA close to each other, and therefore the POIs Potential binding of selector codons at amino acid positions encoded by lectacodons The ncAA was introduced by using a tRNA that supports the introduction of ncAAs (rather than the introduction of amino acids with different tRNAs that can be used). Therefore, the selector codon may be a sense codon. In a preferred embodiment, the selector codon is an endogenous codon of the cell used to prepare the POI. A codon that is not recognized by a specific tRNA.

[0149] O-tRNA ncAA The anticodon of the POI is the selector on the mRNA (the mRNA of the POI). The polypeptide (POI) that is bound to the codon and is therefore encoded by said mRNA The novel OT system described herein site-specifically incorporates ncAAs onto the growing chain of Examples of selector codons that are useful for: - Nonsense codons, e.g. stop codons, e.g. amber (UAG), ochre (UA A), and the opal (UGA) codon; - codons consisting of more than three bases (e.g. four-base codons); - codons derived from natural or unnatural base pairs; and - including, but not limited to, sense codons. Where a selector codon that is a sense codon (i.e., a natural three-base codon) is used In the present invention, the endogenous translation system of the cell used for expressing the POI according to the method of the present invention is the native It is preferable to not use (or only use) three-base codons. Cells that lack or have reduced abundance of tRNAs that recognize natural three-base codons or a cell in which the natural three-base codon is a rare codon. Use of one or more stop codons as stop codons, e.g., one or more of amber, ochre, and opal. Use of the above is particularly preferred.

[0150] Any number of selector codons, e.g., one or more, two or more, or more than three selector codons. Dong et al. have reported a method for the preparation of a polynucleotide encoding a desired polypeptide (target polypeptide POI). The POI can carry two or more ncAA residues. The ncAA residues are the same and can be encoded by the same type of selector codon. or can be different and encoded by different selector codons. An anticodon has the reverse complementary sequence of the corresponding codon.

[0151] Repressor tRNAs are tRNAs that are involved in the translation of messenger RNA (mRNA) by a given translation system (e.g., a cell). tRNA that modifies the reading of the tRNA (e.g., O-tRNA ncAA ) Inhibitory tRN A can read, for example, a stop codon, a four base codon, or a rare codon.

[0152] The O-tRNA is preferentially amino-activated by the O-RS (rather than by the endogenous synthetase). acylated and capable of decoding a selector codon as described herein. The RS recognizes and preferentially encodes O-tRNAs that have, for example, extended anticodon loops. The O-tRNA is aminoacylated with an ncAA. The O-tRNA and O-RS used in the methods and / or fusion proteins of the invention can be naturally occurring or naturally occurring tRNA and / or can be derived by mutation of RS. In various embodiments, tRNA and R In another embodiment, the tRNA is derived from a first organism. and the RS is derived from a naturally occurring or mutated naturally occurring tRNA of a second organism. The RS is derived from a naturally occurring or mutated naturally occurring RS from Suitable (orthogonal) tRNA / RS pairs can be selected based on, for example, library screening results. Alternatively, the desired tRNA and RS can be selected from a library of mutant tRNAs and RSs. A suitable tRNA / RS pair is a heterologous tRNA / synthetase that is imported into the translation system from the source species. Preferably, the cells used as the translation system are the same as the source species. Methods for evolving tRNA / RS pairs are described, for example, in WO02 / 08592. 3 and WO02 / 06075. Conventional site-directed mutagenesis was used to introduce a selector codon into the coding sequence of the POI. It can be used.

[0153] 4. Nucleic acid molecules

[0154] The present invention relates to a nucleotide sequence encoding at least one of the fusion proteins of the present invention and and / or a nucleic acid molecule containing a nucleotide sequence complementary thereto (single-stranded or double-stranded DNA and It also relates to RNA sequences, eg cDNA, mRNA) or combinations of such nucleic acid molecules.

[0155] Furthermore, the present invention relates to a method for the preparation of a polypeptide comprising: (i) a nucleotide sequence encoding at least one POI (CS POI ) and (the POI is linked to the CS by a selector codon POI One or more of the above codes (ii) a targeting nucleotide sequence (TN ) and (the RNA molecule containing the TN (RNA version) is an RNA target via the TN. A nucleic acid molecule (1) comprising a targeting polypeptide (RNA-TP) capable of interacting with the targeting polypeptide (RNA-TP). Single- or double-stranded DNA and RNA sequences, e.g., cDNA, mRNA) or such nucleic acid molecules It relates to a combination of.

[0156] The nucleic acid molecules of the present invention may additionally contain untranslated sequences at the 3' and / or 5' ends of the coding gene region. TN may preferably comprise a sequence of nucleic acid encoding the POI(s). At the 3' end of the molecule, e.g., a nucleic acid of the invention encoding a POI(s). The molecule may be cloned using conventional cloning techniques known in the art to remove the 3' untranslated region, particularly the 3' ) can be prepared by introducing at least one TN into

[0157] The nucleic acid molecules of the present invention may additionally contain untranslated sequences at the 3' and / or 5' ends of the coding gene region. It may contain columns.

[0158] The present invention further relates to a nucleic acid molecule or a combination of nucleic acid molecules of the invention as described herein. Particularly recombinant expression constructs, which contain a nucleic acid sequence under the genetic control of a regulatory nucleic acid sequence. Thus, the expression cassette of the present invention comprises at least one Nucleic acid sequence encoding one POI (plus TN) or at least one fusion protein The present invention relates to these expression constructs (expression vectors). The present invention also relates to a recombinant vector comprising at least one of the following:

[0159] The expression cassette typically contains the POI(s) or fusion protein to be expressed. Located 5' (upstream) of a nucleic acid sequence encoding a protein or proteins and functionally related thereto a promoter sequence operably linked to the 3' (downstream) terminator of said coding sequence; -sequence, and optionally further regulatory elements. Examples of such further regulatory elements are targeting sequences, Enhancers, polyadenylation signals, selection markers, amplification signals, replication origins, and Suitable regulatory sequences include, but are not limited to, those described in, for example, Goeddel, G. ene Expression Technology:Methods in Enz ymology 185, Academic Press, San Diego, CA( 1990).

[0160] In addition to these control sequences, the natural control of these sequences precedes the actual structural gene. Optionally, the natural regulation is switched off and expression of the gene is increased. However, the nucleic acid construct may be genetically modified in a manner similar to that described above. That is, no additional control signals are inserted in front of the coding sequence and its control is not affected. In addition, the native promoter has not been removed. Instead, the native regulatory sequence is The gene is mutated in such a way that regulation no longer occurs and gene expression is increased.

[0161] Elements of a nucleic acid molecule, such as promoters, polypeptide coding sequences, terminators, control sequences, An "operable" linkage of regulatory factors means that the coding sequence can be transcribed and that any regulatory elements are capable of binding to said transcription. This means that these elements are arranged in such a way that the control of This can be achieved by direct linkage of the elements on one and the same nucleic acid molecule. However, such a direct link is not necessarily required. Sequences, such as enhancer sequences, can be derived from more distant locations or even from other DNA molecules. They can exert their functions on the target sequence. The two sequences are covalently linked together. The nucleic acid sequence to be transcribed is inserted downstream of the promoter sequence ( That is, the promoter sequence and the target gene to be expressed are preferably located at the 3' end of the promoter sequence. The distance between the nucleic acid sequences is less than 200 base pairs, or less than 100 base pairs. , or less than 50 base pairs.

[0162] For expression by a cell, the expression cassette is advantageously inserted on an expression vector. The present vector is selected according to the cells to be used for expression. This depends on the cell. Vectors are well known to those skilled in the art and are capable of optimal expression of the coding nucleotide sequence. For example, “Cloning vectors” (Pouwels PHet al., Ed.,Elsevier,Amsterdam-New York-Oxford,1 Examples of expression vectors are plasmids, viral vectors (viral vectors), and the like. Viruses, such as SV40, CMV, baculovirus and adenovirus, transposons, These include, but are not limited to, sequences, IS elements, phasmids, cosmids, and linear or circular DNA. For example, see the book "Cloning Vectors" (Eds. Pouwe ls PHet al. Elsevier,Amsterdam-New Yor See K-Oxford, 1985, ISBN 0444904018. - can replicate autonomously within a (host) cell or can replicate chromosomally The expression vector comprising at least one expression cassette of the present invention can be used in combination with a further Represents an aspect.

[0163] For expression of a POI by a cell according to the invention, for example, a nucleic acid sequence encoding the POI can be Alternatively, a gene (e.g., an expression vector of the invention) can be introduced into the cell. The amino acid position where the existing gene of the cell is intended for the POI to carry an ncAA residue The polypeptide encoding the ( A method for introducing a recombinant nucleic acid molecule into a cell or for modifying an existing gene in a cell. The methods are known in the art.

[0164] The term "expression" in the context of the present invention refers to the expression of a gene encoded by a corresponding nucleic acid sequence in a cell. The term "expression" refers to the production of a polypeptide encoded by a nucleic acid sequence in a cell. They are also used to produce tRNA molecules that

[0165] The nucleic acid molecules of the present invention, including the expression cassettes and expression vectors of the present invention, are disclosed in The vector can be prepared by conventional recombination and cloning techniques. For example, T. Maniatis, EF Fritsch and J.Sambrook,Molecular Cloning:A Laborato ry Manual,Cold Spring Harbor Laboratory, Cold Spring Harbor, NY (1989), and TJ Silha vy,MLBerman and LWEnquist,Experiment s with Gene Fusions,Cold Spring Harbor L Aboratory, Cold Spring Harbor, NY (1984), and and Ausubel,FMet al.,Current Protocols in Molecular Biology,Greene Publishing Ass. As described in oc. and Wiley Interscience (1987) It is.

[0166] Nucleic acid molecules or combinations of nucleic acid molecules of the invention, including expression cassettes and expression vectors of the invention The combination can be isolated, for example, by methods known in the art.

[0167] An "isolated" nucleic acid molecule is one that is separated from other nucleic acid molecules that are present in the natural source of the nucleic acid. , and when it is produced by recombinant techniques, other cellular material or culture medium. or when it is chemically synthesized, is essentially free of chemical precursors or other chemicals may be absent.

[0168] Nucleic acid molecules according to the invention can be synthesized using standard techniques of molecular biology and the sequences provided according to the invention. For example, cDNA can be isolated from a suitable cDNA bank. One of the complete sequences specifically disclosed as hybridization probes or The segments can be isolated using standard hybridization techniques. (e.g. Sambrook, J., Fritsch, E. F. and Maniat is,T.Molecular Cloning:A Laboratory Manu al.2nd edition,Cold Spring Harbor Labora tory,Cold Spring Harbor Laboratory Press , Cold Spring Harbor, NY, 1989). In addition, a nucleic acid molecule containing one of the disclosed sequences or a segment thereof can be derived from this sequence. Isolation by polymerase chain reaction using oligonucleotide primers constructed according to the The nucleic acid thus amplified can be cloned into a suitable vector. and can be characterized by DNA sequence analysis. Oligonucleotides can be produced by standard synthetic methods, e.g., by an automated DNA synthesizer. It can be produced.

[0169] 5. ncAA and post-translational POI modifications

[0170] The abbreviation "ncAA" generally refers to the 22 naturally occurring proteinogenic amino acids. A rare, non-standard, or unnatural amino acid or amino acid residue. It is well known in the field (e.g., Liu et al., Annu Rev Biochem 2010,79:413-444;Lemke,ChemBioChem 2014,1 5:1691-1694. The term "ncAA" refers to amino acid derivatives, such as (α- Such derivatives may also be translationally incorporated. For example, Ohta et al., 2008, ChemB See BioChem 9:2773-2778. Thus, as used herein, "amino acid" refers to The terms "aminoacylation" and "aminoacylation" refer to the coupling of a tRNA with an α-amino acid. The ligation is not limited to RS-catalyzed ligation, but can also be used to catalyze ncAA derivatives such as tRNA and α-hydroxy acids. RS-catalyzed coupling with a conductor is also included.

[0171] Particularly preferred ncAAs for use in the present invention are those that are suitable for, e.g., click chemistry reactions. Such a click reaction can be used to further post-translationally modify the saccharide. Electron-demanded Diels-Alder cycloaddition (SPIEDAC; e.g. Devaraja et al.,Angew Chem Int Ed Engl 2009,48:70 13), as well as with azides, nitrile oxides, nitrones, and diazocarbonyl reagents. A strained cycloalkynyl group (or a triple bond substituted with an amino group) Cycloaddition between a cycloalkynyl analog group having one or more ring atoms that are not (e.g. Sanders et al., J Am Chem Soc 2010,13 3:949;Agard et al.,J Am Chem Soc 2004,12 6:15046), including, for example, strain-promoted alkyne-azide cycloaddition (SPAAC). Such a Click reaction involves coupling a target polypeptide with a suitable group on a coupling partner molecule. Enables ultrafast, mutually orthogonal, covalent site-specific coupling of multiple ncAA labeling groups The above-mentioned linkages and labeling groups that can react by the Click reaction are Pairs are known in the art. Examples of ncAA's suitable for use in the present invention that contain a linking group are described, for example, in WO2012 / 104422 and WO2015 / 107064. Examples of ncAA include, but are not limited to, ncAA ("unnatural amino acid", "UAA"). The optionally substituted strained alkynyl group is an optionally substituted trans-cyclooctenyl. Groups, for example, include, but are not limited to, those depicted. The strained alkenyl group may be an optionally substituted cyclooctynyl group, for example as described in WO2012 / 104422 and WO2015 / 107064, but The optionally substituted tetrazinyl group is not limited to the above. and WO2015 / 107064.

[0172] The ncAAs used in the context of the present invention can be used in the form of their salts. The salts of the ncAAs described herein are acid or base addition salts, particularly those that are physiologically tolerated. Addition salts with acids or bases are meant. Physiologically tolerable acid addition salts are salts of suitable organic or It can be formed by treatment of the base form of ncAA with an inorganic acid. The ncAAs contained therein can be converted to their non-toxic metal ions by treatment with appropriate organic and inorganic bases. or amine addition salt form. The salt of the carboxyl group of the ncAA can be They can be produced in a manner known in the art, and inorganic salts, such as sodium, calcium, ammonium, , iron, and zinc salts, and organic bases, such as amines, e.g., triethanolamine, alkane, These include salts with ginine, lysine, piperidine, etc. ncAAs are also salts of acid addition, e.g., mineral acids. in the form of salts with, for example, hydrochloric acid or sulfuric acid, and with organic acids, for example acetic acid and oxalic acid. The ncAA and their salts useful in the present invention can also be used. Also included are hydrates and solvent addition forms, such as hydrates, alcoholates, and the like.

[0173] Physiologically tolerated acids or bases may be used, particularly in the preparation of POIs bearing ncAA residues. and is tolerated by the translation system in which it is used, e.g., is substantially non-toxic to eukaryotic cells. do.

[0174] The ncAAs and their salts useful in the context of the present invention are well known in the art. For example, the method may be similar to that described in the various publications cited herein. It can be manufactured.

[0175] The nature of the coupling partner molecule depends on the intended use. The peptides can be coupled to molecules suitable for imaging methods; Alternatively, it can be functionalized by coupling to a biologically active molecule. In addition, the coupling partner molecule may be a dye (e.g., a fluorescent, luminescent, or phosphorescent Dyes, e.g. dansyl, coumarin, fluorescein, acridine, rhodamine, silicone rhodamine, BODIPY, or cyanine dyes), which emit fluorescence upon contact with a reagent Molecules that can bind to chromophores (e.g., phytochromes, phycobilins, bilirubin, etc.) ), radioactive labels (e.g., radioactive forms of hydrogen, fluorine, carbon, phosphorus, sulfur, or iodine, For example, tritium 18 F, 11 C. 14 C. 32 P, 33 P, 33 S, 35 S, 11 I n, 125 I, 123 I, 131 I, 212 B. 90 Y, or 186 Rh), MRI sensitivity spin labels, affinity tags (e.g. biotin, His tags, Flag tags, st rep tag, sugar, lipid, sterol, PEG linker, benzylguanine, benzylsitol synthon, or cofactor), polyethylene glycol group (e.g., branched PEG, linear PEG, different (e.g., PEG with molecular weights of 100 or more), photocrosslinkers (e.g., p-azidoiodoacetanilide), NMR Probes, X-ray probes, pH probes, IR probes, resins, solid supports, and biological activities The compounds may include, but are not limited to, groups selected from the group consisting of cyclic amines, cyclic amines, and cyclic amines. Suitable biologically active compounds include cytotoxic compounds (e.g., cancer chemotherapeutic compounds), anti-cancer drugs, and the like. viral compounds, biological response modifiers (e.g., hormones, chemokines, cytokines, insulins, interleukins, drugs that affect microtubules, hormone regulators, and steroid compounds Specific examples of useful coupling partner molecules include, but are not limited to, Member of receptor / ligand pair; member of antibody / antigen pair; member of lectin / carbohydrate pair Members of enzyme / substrate pairs; Biotin / Avidin; Biotin / Streptavidin These include, but are not limited to, digoxin, and digoxin / antidigoxin.

[0176] Covalently coupled in situ to the conjugation partner molecule (the binding group of the The ability of certain ncAA residues (labeling groups) to be labeled is particularly The ncAA residues in eukaryotic cells or tissues expressing the target polypeptide are identified by the cross-linking reaction. For detecting a target polypeptide having a group(s) and a target polypeptide In particular, it can be used to study the distribution and fate of (e.g., eukaryotic) cells. The method of the present invention for preparing a POI by expression in a cell or tissue of such cells may be It can be combined with super-resolution microscopy (SRM) to detect POIs within the tissue. Such SRM methods are known in the art and are useful for detecting target polypeptides expressed by the eukaryotic cells of the present invention. The method can be adapted to utilize click chemistry to detect tides. A specific example of such an SRM method is DNA-PAINT (nanoscale topography). DNA point accumulation for imaging in mice; for example, Jungmann et al. al., Nat Methods 11:313-318, 2014 (direct stochastic optical reconstruction microscopy), dSTORM (direct stochastic optical reconstruction microscopy), and STED (stimulated emission depletion) ) Including microscopy.

[0177] 6. Translational preparation of POI by cells

[0178] The OT system provided by the present invention allows for the translational preparation of a POI by a cell.

[0179] The cells used to prepare a POI according to the present invention may be prokaryotic cells. Alternatively, the cells used to prepare the POI according to the present invention are eukaryotic cells. The cells used to prepare the POI according to the present invention may be, for example, single-celled microorganisms. The cells may be distinct cells, such as cells from a living organism, or cell lines derived from cells of a multicellular organism. The cells used to prepare the POI according to the present invention may be derived from tissues, organs, body parts, or may be present in (and be part of) an entire multicellular organism. The method of the invention for the detection of inflammatory bowel diseases can be carried out by separate cells or cell cultures, or by tissue or tissue culture. , organs, body parts, or (whole) multicellular organisms.

[0180] Eukaryotic cells are often easier to handle than prokaryotes, such as E. coli. It is difficult to implement and operate, and thus, as described in the "Background Art" above, Known approaches for POI-selective orthogonal translation are either inaccessible or inaccessible. Thus, the OT system and method of the present invention are suitable for use in eukaryotic cells (e.g., single cells and and multicellular eukaryotic organisms, and eukaryotic cell lines) for expression of the POI. , is particularly advantageous.

[0181] In principle, any prokaryotic or eukaryotic cell can be used to prepare a POI according to the method of the invention. It can be used to treat microorganisms, such as bacteria, fungi, or yeasts, as well as bacteria, such as mammals. Eukaryotic cells, such as mammalian cells, insect cells, yeast cells, and plant cells, can be used. Nuclear cells, especially mammalian cells, are particularly preferred.

[0182] The cells used to prepare a POI according to the present invention contain a nucleic acid encoding the POI. Otide sequence (CS POI ), and the ncAA residue(s) of the POI is / are a selector It is encoded by a codon or codons. POI one or more targeting sequences (TN) is functionally linked to CS. POI and TN(singular or plural) The cells further contain one or more fusion proteins of the invention, The fusion protein(s) comprises at least one O-RS segment and at least The O-RS and the RNA-TP are distinct from each other in the present invention. Alternatively, the O-RS and RNA may be on a fusion protein (e.g., AFP). -TP is a single fusion protein of the invention (e.g., an RNA-TP / O-RS fusion protein). The TN(s) (at least one of the TN(s)) may be present on the Through this, the mRNA is expressed in the cell as an RNA-TP segment of the fusion protein of the present invention. The cell is further capable of interacting with (binding to) at least one of the CS POI 1 carrying an anticodon(s) for the selector codon(s) of Orthogonal tRNAs ncAA Molecule (O-tRNA ncAA The O-tRNA ncAA The molecule and one or more O-RS segments of the fusion protein are capable of acting in a cell as follows: One or more orthogonal O-RS / O-tRNA ncAA This forms a pair. It is possible to introduce ncAA residue(s) into the amino acid sequence of the POI (to be produced). To perform.

[0183] CS with RNA-TP segment(s) POI and TN(singular or plural) Interaction of mRNA with O-RS segment(s) by ncAA -tRNA ncAA and the introduction of ncAA residue(s). Translational preparation of the POI containing the ncAA occurs more specifically in the cytoplasm of cells in the presence of the ncAA. It is thought to occur in the OT assembly (OT organelle).

[0184] CS POI and mRNA containing TN(singular or plural) (mRNA POI ) is introduced into cells Alternatively, the gene can be produced from a recombinant construct (e.g., an expression vector) into which the gene has been introduced. One or more endogenous genes of the cell contain one or more selector codons and one or more TNs. Techniques for introducing recombinant constructs into cells; Methods for modifying endogenous genes of a cell are well known in the art.

[0185] The tRNA of the present invention ncAA The molecules and fusion proteins are recombinant constructs that are introduced into cells. The gene can be produced from an entity (eg, an expression vector).

[0186] The expression vector according to the invention can be used to prepare a POI using the method of the invention. Advantageously, a recombinant cell capable of carrying the invention as described above can be produced. The recombinant vectors described herein are introduced into suitable cells and expressed.

[0187] The cells used to prepare the POI as described herein are capable of expressing the fusion protein. (single or multiple), tRNA ncAA The molecule(s) and the nucleic acid encoding the POI The nucleic acid sequence can be prepared by introducing a nucleotide sequence into a cell. The sequences can be on separate nucleic acid molecules (vectors) or on the same nucleic acid molecule (e.g., vector). They can be arranged in any combination and can be introduced into cells in combination or sequentially. This can be done.

[0188] Preferably, the vector is cloned using conventional cloning and transfection techniques known to those skilled in the art, such as For example, coprecipitation, protoplast fusion, electroporation, and virus-mediated gene expression Gene delivery, lipofection, microinjection, or other methods of nucleic acid delivery are described. Suitable techniques are described, for example, in Curre et al. nt Protocols in Molecular Biology,F.Ausu bel et al.,Ed.,Wiley Interscience,New Yo rk 1997 or Sambrook et al. Molecular Cloning g:A Laboratory Manual.2 nd Cold S edition pring Harbor Laboratory,Cold Spring Harb or Laboratory Press,Cold Spring Harbor,N This is described in Y, 1989.

[0189] In the methods of the present invention, the cells used for POI expression are grown in a manner known to those skilled in the art. Depending on the type of cells, liquid media can be used for the culture. The culture may be batch, semi-batch, or continuous. Nutrients must be present at the start of the culture. or can be subsequently fed semi-continuously or continuously.

[0190] The expressed POI can be isolated by known techniques, such as molecular sieve chromatography. gel filtration, e.g. Q-Sepharose chromatography, ion exchange chromatography Chromatography, and hydrophobic chromatography, as well as other common protein purification techniques, e.g. Purification can be achieved by, for example, ultrafiltration, crystallization, salting out, dialysis, and native gel electrophoresis. Suitable methods are described, for example, in Cooper, TG, Biochemische Arbeitsmethoden[Biochemical processes], Verlag Walter de Gruyter,Berlin,New York Scopes, R., Protein Purification, Spring Inscribed in ger Verlag, New York, Heidelberg, Berlin It has been done.

[0191] To isolate the POI, the POI is linked to a tag that can effect easier purification. It may be advantageous to ligate the corresponding tag coding sequence to the CS POI Introduce to Suitable tags for protein purification are well known in the art. For example, a histidine tag (e.g., His 6 tag) and the antigen recognized by the antibody These include epitopes that can D.,1988,Antibodies:A Laboratory Manual.C (as described by the old Spring Harbor (NY) Press) These tags are used to attach proteins to solid supports, such as polymer matrices. This can act, for example, on the packing of chromatography columns. or on a microtiter plate or some other support. Can be used on the body.

[0192] The tag linked to the POI can also serve to detect the POI. Tags for detection of proteins are well known in the art and include, for example, fluorescent dyes, detectable after reaction with a substrate. These include enzyme markers that form detectable reaction products, and others.

[0193] To prepare a POI according to the methods of the invention, a suitable nucleic acid sequence to allow translation of the POI is required. Over a suitable period of time, one or more ncAA residues corresponding to the ncAA residue(s) of the POI are Culturing the cells in the presence of ncAA (which may conveniently be included in the culture medium). Expression can be achieved by: and / or tRNA ncAA Depending on the nucleic acid(s) encoding the molecule, e.g. Arabinose, isopropyl β-D-thiogalactoside (IPTG) to enable transcription Expression is induced by adding a transcription-inducing compound such as tetracycline. The government may request that the relevant authority provide guidance.

[0194] After translation, the POI can optionally be recovered from the translation system. POI may be partially or substantially extracted according to procedures used and known by those skilled in the art. The target polypeptide can be isolated in the culture medium and then purified to homogeneity in either the culture medium or in a culture medium containing the target polypeptide. Unless secreted, recovery usually requires cell disruption. Methods for cell disruption are well known in the art. Yes, e.g. sonication, liquid shear disruption (e.g. with a French press) , mechanical methods (e.g., using a blender or grinder), or freeze-thaw Physical disruption by ice and lipid-lipid, protein-protein and / or protein Chemical lysis using agents that block protein-lipid interactions (e.g., detergents), and This involves a combination of physical disruption techniques and chemical lysis. Standard procedures for purifying polypeptides are also well known in the art, e.g., by purifying Ammonium or ethanol precipitation, acid or base extraction, column chromatography, affinity chromatography Chromatography, anion or cation exchange chromatography, sulphocellulose chromatography, hydrophobic interaction chromatography, hydroxyl This includes apatite chromatography, lectin chromatography, gel electrophoresis, etc. If desired, a protein refolding step can be performed to ensure that the protein is correctly folded. It can be used to produce mature proteins. High performance liquid chromatography (HPLC), affinity chromatography, or other suitable methods may be used to obtain a high purity product. The polypeptides of the present invention may be used in final purification steps, including purification of the polypeptides of the present invention. The antibodies can be used as purification reagents, i.e., for affinity-based purification of polypeptides. Various purification / protein folding methods are well known in the art. For example, Scopes, Protein Purification, Spring Er, Berlin (1993); and Deutscher, Methods in E nzymology Vol.182:Guide to Protein Purif ication, Academic Press (1990); and references therein. This document includes those referenced in the accompanying text.

[0195] As stated, those skilled in the art will recognize that, following synthesis, expression, and / or purification, a polypeptide is It is possible for a polypeptide to possess a conformation different from the desired conformation. For example, polypeptides produced by prokaryotic systems are often In order to achieve proper folding, exposure to chaotropic agents must be performed to optimize the folding. For example, during purification from a cell lysate, the expressed polypeptide is optionally This can be done, for example, by treating a protein with a solution of guanidine HCI or similar. This is achieved by solubilizing the expressed polypeptide in a chaotropic agent. The polypeptide is then denatured and reduced, and then oriented in a preferred conformation. It may be desirable to refold the protein using, for example, guanidine, urea, DT T, DTE, and / or chaperonins can be added to the translation product of interest. Methods for reducing, denaturing, and refolding proteins are well known to those of skill in the art. For example, refolding was observed in a redox buffer containing oxidized glutathione and l-arginine. You can ding. Polypeptides produced by the methods of the invention are also described. , can be prepared by the methods of the invention utilizing the OT system described herein.

[0196] 7. Kit

[0197] The present invention provides a method for preparing a POI having at least one non-standard amino acid (ncAA) residue. The kit of the invention also includes a kit for the preparation of at least one fusion protein of the invention. The kit may include at least one expression vector for expression. The fusion protein(s) encoded by the nucleotide sequence(s) of The kit may include an O-RS segment and at least one RNA-TP segment. The target further comprises at least one ncAA corresponding to at least one ncAA residue of the POI. Conveniently, the O-RS segment may comprise at least one n The kit further comprises an orthogonal tRNA that can be aminoacylated with cAA. NA ncAA (O-tRNA ncAA ) molecule. Further components of the kit may include a multiple cloning site and a targeting nucleotide sequence. and at least one expression vector comprising a nucleotide sequence (TN), The RNA molecule containing the TN interacts with an RNA targeting polypeptide (RNA-TP) via the TN. Conveniently, the TN, when present on an RNA molecule, can be included in the kit. The fusion protein(s) encoded by the expression vector(s) contained therein A sequence capable of interacting with at least one RNA-TP segment of The kit further comprises a readily accessible nucleic acid sequence having at least one non-standard amino acid (ncAA) residue. At least one reporter polypeptide encoding a detectable (e.g., fluorescent) reporter polypeptide. The mRNA transcribed from the construct may comprise a target construct, such that the mRNA transcribed from the construct may be Including the TN listed.

[0198] The kit of the present invention comprises a method for preparing the ncAA residue-containing POI described herein. It can be used in the method of the invention.

[0199] Specific Embodiments

[0200] The present invention further provides the following non-limiting embodiments E1 to E50.

[0201] E1: an assembler fusion protein (AFP) comprising: (a) (a1) A polypeptide segment derived from an intracellular targeting polypeptide (IC-TP segment (The intracellular targeting polypeptide is a target polypeptide for a cell within or directly adjacent to the cytoplasm.) (i) targeting, and therefore locally enriching, intracellular structural elements; and (a2) A polypeptide segment (PSP segment) derived from a phase-separated polypeptide ( The phase-separating polypeptides act to induce cell proliferation in a manner that creates sites of high local concentration in the cytoplasm. (having the ability to undergo self-association in the cytoplasm of cells), At least one first polypeptide acting as an assembler (AP) selected from Segments and; (b) b1) an RNA-targeting polypeptide (RNA-TP) segment, and b2) an orthogonal aminoacyl-tRNA synthetase (O-RS) segment; At least one second polypeptide acting as an effector (EP) selected from dosegment and; Including, The polypeptide segments are functionally linked to each other on the AFP.

[0202] E2: AFP of E1, containing at least two APs, preferably at least one I It contains a C-TP segment and at least one PSP segment.

[0203] E3: An AFP of E1 or E2 having one of the following structures (N-terminus to C-terminus): : [IC-TP] m -[EP] o [EP] o -[IC-TP] m [PSP] n -[EP] o [EP] o - [PSP] n [IC-TP] m -[EP] o - [PSP] n [PSP]n -[EP] o -[IC-TP] m [IC-TP] m - [PSP] n -[EP] o [EP] o - [PSP] n -[IC-TP] m [PSP] n -[IC-TP] m -[EP] o [EP] o -[IC-TP] m - [PSP] n wherein m, n, and o are each independently an integer selected from 1, 2, 3, 4, or 5. and "-" indicates a peptidic linkage.

[0204] E4: Any one of AFPs E1 to E3, wherein at least one EP is RNA- It is selected from the TP segment.

[0205] E5: Any one of AFPs E1 to E3, and at least one EP is O-RS A segment is selected.

[0206] E6: AFP selected from any one of E1 to E3 and the RNA-TP segment At least one EP selected from the O-RS segment and Includes:

[0207] E7: AFP of any one of E1 to E6, including dynein, kinesin, and dynein. At least one IC-TP selected from kinesin and kinesin fragments and variants. segments that target and are concentrated at the plus or minus ends of microtubules. maintains the ability to

[0208] E8: AFP of any one of E1 to E6, which is a transmembrane domain of a membrane protein and and at least one I selected from functional fragments and variants of the transmembrane domain. These contain C-TP segments, which bind to membranes, particularly the cell membrane, nuclear membrane, and mitochondrial membrane. The cytoplasmic side of the membrane selected from the group consisting of β-lactams, β-lactams, and β-lactams retain the ability to target and be concentrated on the cytoplasmic side of the membrane selected from the group consisting of β-lactams, β-lactams, and β-lactams.

[0209] E9: Any one of AFPs E1 to E8: - KIF16B comprising the amino acid sequence of SEQ ID NO: 20 1-400 or SEQ ID NO: 20 At least 60%, at least 70%, at least 80%, at least 90% %, at least 95%, at least 96%, at least 97%, at least 98%, or or a functional fragment or variant thereof having at least 99% amino acid sequence identity. Heterogeneity; - KIF13A comprising the amino acid sequence of SEQ ID NO: 22 1-411 , Δ390, or sequence number At least 60%, at least 70%, at least 80%, or at least at least 90%, at least 95%, at least 96%, at least 97%, at least 9 8% or a functional fragment thereof having at least 99% amino acid sequence identity or variants; - TOMM20 comprising the amino acid sequence of SEQ ID NO: 24 1-70 or the amine of SEQ ID NO:24 At least 60%, at least 70%, at least 80%, at least 90% , at least 95%, at least 96%, at least 97%, at least 98%, or or a functional fragment or variant thereof having at least 99% amino acid sequence identity body; LcK comprising the amino acid sequence of SEQ ID NO: 26 or at least at least 60%, at least 70%, at least 80%, at least 90%, at least 9 5%, at least 96%, at least 97%, at least 98%, or at least 9 a functional fragment or variant thereof having 9% amino acid sequence identity; - FRB-CD28 comprising the amino acid sequence of SEQ ID NO: 28 or the amino acid sequence of SEQ ID NO: 28 At least 60%, at least 70%, at least 80%, at least 90%, At least 95%, at least 96%, at least 97%, at least 98%, or at least a functional fragment or variant thereof having at least 99% amino acid sequence identity; - FUS-CD28 comprising the amino acid sequence of SEQ ID NO: 30, or the amino acid sequence of SEQ ID NO: 30 At least 60%, at least 70%, at least 80%, at least 90%, At least 95%, at least 96%, at least 97%, at least 98%, or at least a functional fragment or variant thereof having at least 99% amino acid sequence identity; EB1 comprising the amino acid sequence of SEQ ID NO: 302 or the amino acid sequence of SEQ ID NO: 303 At least 60%, at least 70%, at least 80%, at least 90%, at least At least 95%, at least 96%, at least 97%, at least 98%, or at least a functional fragment or variant thereof having at least 99% amino acid sequence identity; CG1 comprising the amino acid sequence of SEQ ID NO: 304 or with the amino acid sequence of SEQ ID NO: 304 At least 60%, at least 70%, at least 80%, at least 90%, at least At least 95%, at least 96%, at least 97%, at least 98%, or at least a functional fragment or variant thereof having at least 99% amino acid sequence identity; EBAG9 comprising the amino acid sequence of SEQ ID NO: 292 (full length), or SEQ ID NO: 292 and At least 60%, at least 70%, at least 80%, at least 90%, at least At least 95%, at least 96%, at least 97%, at least 98%, or at least or the first N-terminal 29 amino acids of SEQ ID NO:294; or a functional fragment or variant thereof comprising a 5'-amino acid residue; or a sequence similar to SEQ ID NO:294 and at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, At least 96%, at least 97%, at least 98%, or at least 99% of the functional fragments or variants thereof having amino acid sequence identity; - CMP Sia Tr comprising the amino acid sequence of SEQ ID NO: 296, or Amino acid sequence and at least 60%, at least 70%, at least 80%, at least 9 0%, at least 95%, at least 96%, at least 97%, at least 98%, or a functional fragment thereof having at least 99% amino acid sequence identity thereto Mutants; and - P450 2C1, which targets the cytoplasmic side of the ER membrane, or at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the amino acid sequence A functional fragment or variant thereof having sequence identity, particularly the first 27 ( A fragment containing the first 29 (SEQ ID NO: 300) amino acid residues or at least 60%, at least 70%, or at least At least 80%, at least 90%, at least 95%, at least 96%, at least 97% %, at least 97%, at least 98%, or at least 99% amino acid sequence identity a functional fragment or variant thereof having the same identity; The IC-TP segment includes at least one selected from:

[0210] E10: Any one of AFPs E1 to E9, which is an intrinsically disordered protein (IDP). , in particular prion-like domains, as well as functional fragments of IDP or prion-like domains and and a variant thereof, which are located in the cytoplasm. self-associate in the cytoplasm of cells, creating sites of high local concentration of Retain the ability.

[0211] E11: Any one of AFPs E1 to E10: - SPD5 comprising the amino acid sequence of SEQ ID NO: 32 or a combination of the amino acid sequence of SEQ ID NO: 32 and at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least a functional fragment or variant thereof having 99% amino acid sequence identity; - FUS comprising the amino acid sequence of SEQ ID NO: 34 or at least at least 60%, at least 70%, at least 80%, at least 90%, at least 9 5%, at least 96%, at least 97%, at least 98%, or at least 9 a functional fragment or variant thereof having 9% amino acid sequence identity; and EWSR1 comprising the amino acid sequence of SEQ ID NO: 36 or At least 60%, at least 70%, at least 80%, at least 90%, at least At least 95%, at least 96%, at least 97%, at least 98%, or at least a functional fragment or variant thereof having at least 99% amino acid sequence identity; The PSP segment comprises at least one selected from:

[0212] E12: Any one of AFPs E1 to E11, including at least one RNA-T It contains the P segment, the RNA-binding segment of the viral coat protein, and the viral of viral coat proteins that retain the ability to specifically interact with RNA motifs in the genome The RNA-binding segment is selected from functional fragments and variants of the RNA-binding segment.

[0213] E13: Any one of AFPs E1 to E12: - an MCP comprising the amino acid sequence of SEQ ID NO: 14 or at least at least 60%, at least 70%, at least 80%, at least 90%, at least 9 5%, at least 96%, at least 97%, at least 98%, or at least 9 a functional fragment or variant thereof having 9% amino acid sequence identity; - λ comprising the amino acid sequence of SEQ ID NO: 16 N22 or a sequence similar to the amino acid sequence of SEQ ID NO:16 at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least a functional fragment or variant thereof having 99% amino acid sequence identity; and - a PCP comprising the amino acid sequence of SEQ ID NO: 306 or At least 60%, at least 70%, at least 80%, at least 90%, at least At least 95%, at least 96%, at least 97%, at least 98%, or at least a functional fragment or variant thereof having at least 99% amino acid sequence identity; The RNA-TP segment comprises at least one RNA-TP segment selected from:

[0214] E14: Any one of AFPs E1 to E13: - Methanococcus jannaschii tyrosyl-tRNA synthetase ; - Escherichia coli tyrosyl-tRNA synthetase; - Escherichia coli leucyl-tRNA synthetase; - Methanosarcina mazeii pyrrolysyl-tRNA synthetase; - Methanosarcina barkeri pyrrolysyl-tRNA synthetase; - Methanosarcina acetivorans pyrrolysyl-tRNA synthase Tase; - Methanosarcina thermophila pyrrolysyl-tRNA synthase Tase; - Methanococcoides burtonii pyrrolysyl-tRNA synthetase -ze; - Desulfitobacterium hafniense pyrrolysyl-tRNA synthase acetase; Their functional fragments that retain aminoacyl-tRNA synthetase enzyme activity and Mutants; The O-RS segment comprises at least one O-RS segment selected from:

[0215] E15: Any one of AFPs E1 to E14: - PylRS comprising the amino acid sequence of SEQ ID NO: 8 AF or the amino acid sequence of SEQ ID NO:8 At least 60%, at least 70%, at least 80%, at least 90%, at least At least 95%, at least 96%, at least 97%, at least 98%, or at least a functional fragment or variant thereof having at least 99% amino acid sequence identity; - PylRS comprising the amino acid sequence of SEQ ID NO: 10 AA or the amino acid sequence of SEQ ID NO: 10 Column and at least 60%, at least 70%, at least 80%, at least 90%, at least At least 95%, at least 96%, at least 97%, at least 98%, or a functional fragment or variant thereof having at least 99% amino acid sequence identity; - PylRS comprising the amino acid sequence of SEQ ID NO: 12 AAAF or the amino acid sequence of SEQ ID NO:12 At least 60%, at least 70%, at least 80%, at least 90% of the acid sequence, At least 95%, at least 96%, at least 97%, at least 98%, or A functional fragment or variant thereof having at least 99% amino acid sequence identity ; - IFRS1 comprising the amino acid sequence of SEQ ID NO: 224, or Column and at least 60%, at least 70%, at least 80%, at least 90%, at least At least 95%, at least 96%, at least 97%, at least 98%, or a functional fragment or variant thereof having at least 99% amino acid sequence identity; - a CbzRS comprising the amino acid sequence of SEQ ID NO: 226, or Column and at least 60%, at least 70%, at least 80%, at least 90%, at least At least 95%, at least 96%, at least 97%, at least 98%, or a functional fragment or variant thereof having at least 99% amino acid sequence identity; - a CpkRS comprising the amino acid sequence of SEQ ID NO: 228, or Column and at least 60%, at least 70%, at least 80%, at least 90%, at least At least 95%, at least 96%, at least 97%, at least 98%, or a functional fragment or variant thereof having at least 99% amino acid sequence identity; and Beauty - an OMeRS comprising the amino acid sequence of SEQ ID NO: 236, or Column and at least 60%, at least 70%, at least 80%, at least 90%, at least At least 95%, at least 96%, at least 97%, at least 98%, or a functional fragment or variant thereof having at least 99% amino acid sequence identity; The O-RS segment comprises at least one O-RS segment selected from:

[0216] E16: an assembler fusion protein (AFP) combination, which is any one of E1 to E15 Contains at least two AFPs from either

[0217] E17: AFP combination of E16, comprising at least one RNA-TP segment At least one first AFP containing a nucleotide sequence and at least one O-RS segment and at least one second AFP.

[0218] E18: A fusion protein (RNA-TP / O-RS fusion protein): (i) at least one RNA-targeting polypeptide (RNA-TP) segment; (ii) at least one orthogonal aminoacyl-tRNA synthetase (O-RS) segment; To and; Including, The polypeptide segments are connected to each other on the RNA-TP / O-RS fusion protein. The functional groups are bonded to the amines.

[0219] E19: an RNA-TP / O-RS fusion protein of E18 having the following structure (N-terminus to the C-terminus): [RNA-TP] x- [O-RS] y [O-RS] y- [RNA-TP] x wherein x and y are, independently of each other, integers selected from 1, 2, 3, 4, and 5. , "-" indicates a peptidic linkage.

[0220] E20: An RNA-TP / O-RS fusion protein of E18 or E19, comprising at least Both contain an RNA-TP segment and are related to the RNA-binding segment of the viral coat protein. ment, as well as viruses that retain the ability to specifically interact with viral RNA motifs. A functional fragment or variant of the RNA-binding segment of the scort protein is selected from the group consisting of a can be.

[0221] E21: Any one of the RNA-TP / O-RS fusion proteins E18 to E20. Is this: - an MCP comprising the amino acid sequence of SEQ ID NO: 14 or at least at least 60%, at least 70%, at least 80%, at least 90%, at least 9 5%, at least 96%, at least 97%, at least 98%, or at least 9 a functional fragment or variant thereof having 9% amino acid sequence identity; - λ comprising the amino acid sequence of SEQ ID NO: 16 N22 or a sequence similar to the amino acid sequence of SEQ ID NO:16 at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least a functional fragment or variant thereof having 99% amino acid sequence identity; and - a PCP comprising the amino acid sequence of SEQ ID NO: 306 or At least 60%, at least 70%, at least 80%, at least 90%, at least At least 95%, at least 96%, at least 97%, at least 98%, or at least a functional fragment or variant thereof having at least 99% amino acid sequence identity; The RNA-TP segment comprises at least one RNA-TP segment selected from:

[0222] E22: Any one of the RNA-TP / O-RS fusion proteins E18 to E21. Is this: - Methanococcus jannaschii tyrosyl-tRNA synthetase ; - Escherichia coli tyrosyl-tRNA synthetase; - Escherichia coli leucyl-tRNA synthetase; - Methanosarcina mazeii pyrrolysyl-tRNA synthetase; - Methanosarcina barkeri pyrrolysyl-tRNA synthetase; - Methanosarcina acetivorans pyrrolysyl-tRNA synthase Tase; - Methanosarcina thermophila pyrrolysyl-tRNA synthase Tase; - Methanococcoides burtonii pyrrolysyl-tRNA synthetase -ze; - Desulfitobacterium hafniense pyrrolysyl-tRNA synthase acetase; Their functional fragments that retain aminoacyl-tRNA synthetase enzyme activity and Mutants; The O-RS segment comprises at least one O-RS segment selected from:

[0223] E23: Any one of the RNA-TP / O-RS fusion proteins E18 to E22. Is this: - PylRS comprising the amino acid sequence of SEQ ID NO: 8 AF or the amino acid sequence of SEQ ID NO:8 a functional fragment or variant thereof having at least 60% sequence identity; - PylRS comprising the amino acid sequence of SEQ ID NO: 10 AA or the amino acid sequence of SEQ ID NO: 10 a functional fragment or variant thereof having at least 60% sequence identity with the sequence; - PylRS comprising the amino acid sequence of SEQ ID NO: 12 AAAF or the amino acid sequence of SEQ ID NO:12 or a functional fragment or variant thereof having at least 60% sequence identity with the nucleic acid sequence. ; - IFRS1 comprising the amino acid sequence of SEQ ID NO: 224, or Column and at least 60%, at least 70%, at least 80%, at least 90%, at least At least 95%, at least 96%, at least 97%, at least 98%, or a functional fragment or variant thereof having at least 99% amino acid sequence identity; - a CbzRS comprising the amino acid sequence of SEQ ID NO: 226, or Column and at least 60%, at least 70%, at least 80%, at least 90%, at least At least 95%, at least 96%, at least 97%, at least 98%, or a functional fragment or variant thereof having at least 99% amino acid sequence identity; - a CpkRS comprising the amino acid sequence of SEQ ID NO: 228, or Column and at least 60%, at least 70%, at least 80%, at least 90%, at least At least 95%, at least 96%, at least 97%, at least 98%, or a functional fragment or variant thereof having at least 99% amino acid sequence identity; and Beauty - an OMeRS comprising the amino acid sequence of SEQ ID NO: 236, or Column and at least 60%, at least 70%, at least 80%, at least 90%, at least At least 95%, at least 96%, at least 97%, at least 98%, or a functional fragment or variant thereof having at least 99% amino acid sequence identity; The O-RS segment comprises at least one O-RS segment selected from:

[0224] E24: A nucleic acid molecule or a combination of two or more nucleic acid molecules comprising: (i) At least one AFP of any one of E1 to E15, or E16 or E1 7, or a nucleotide sequence encoding at least one AFP combination of (ii) a nucleic acid sequence complementary to the nucleotide sequence of (i); (iii) Both (i) and (ii); Includes.

[0225] E25: A nucleic acid molecule or a combination of two or more nucleic acid molecules comprising: (i) At least one RNA-TP / O-RS fusion tag from any one of E18 to E23 a nucleotide sequence encoding a protein; (ii) a nucleic acid sequence complementary to (i); or (iii) Both (i) and (ii); Includes.

[0226] E26: The nucleotide sequence of a nucleic acid molecule or a combination of nucleic acid molecules of E24 or E25 An expression cassette comprising:

[0227] E27: An expression vector comprising at least one expression cassette of E26.

[0228] E28: A cell, comprising at least one nucleic acid molecule or a nucleic acid molecule of E24 or E25. combination, at least one expression cassette of E26, or at least one expression cassette of E27 Includes expression vectors.

[0229] E29: E28 cells, which are eukaryotic cells.

[0230] E30: E28 cells, which are mammalian cells.

[0231] E31: Any one of cells from E28 to E30, and at least one nucleus from E24 A combination of nucleic acid molecules or nucleic acid molecules, or said nucleic acid molecules or combinations of nucleic acid molecules. or at least one expression cassette comprising the nucleotide sequence of It comprises at least one expression vector.

[0232] E32: A cell of E31, comprising at least one RNA-TP segment selected from the group consisting of: E1 to E3 each including one EP and at least one EP selected from the O-RS segment and at least one AFP of any one of E7 to E15, or It comprises a nucleotide sequence which is complementary to the encoding nucleotide sequence.

[0233] E33: A cell of E31, comprising at least one RNA-TP segment selected from the group consisting of: At least one AFP from any one of E1 to E3 and E7 to E15, which contains one EP E1 to E3 and E7 to E8 each including at least one EP selected from the O-RS segment E15 and at least one AFP or The term "nucleotide sequence" includes a nucleotide sequence that is complementary to a nucleotide sequence that is

[0234] E34: Any one of cells from E28 to E30, and at least one nucleus from E25 A combination of nucleic acid molecules or nucleic acid molecules, or said nucleic acid molecules or combinations of nucleic acid molecules. or at least one expression cassette comprising the nucleotide sequence of It comprises at least one expression vector.

[0235] E35: A cell according to any one of E28 to E34, wherein the cell is at least at least one AFP, at least one AFP combination, or at least one RNA-T and expressing a P / O-RS fusion protein, which is a nucleic acid molecule or a combination of nucleic acid molecules. It is encoded by the nucleotide sequence

[0236] E36: Objective to include one or more non-canonical amino acid (ncAA) residues in its amino acid sequence A method for preparing a polypeptide (POI), comprising any one of E31 to E33. The POI is expressed by any one of the cells in the presence of the one or more ncAAs. and wherein the cells include: (i) a nucleotide sequence encoding the POI (CS POI ) (the one or more n of POIs cAA residues are encoded by a selector codon(s), (ii)CS POI and functionally linked to at least one R of AFP in a cell. a targeting nucleotide sequence (TN) capable of interacting with the NA-TP segment; (iii)CS POI an anticodon (single) complementary to the selector codon (single or multiple) of or more) carrying an orthogonal tRNA ncAA (O-tRNA ncAA )molecule( The O-tRNA ncAA The molecule comprises one or more of at least one AFP in a cell. The one or more ncA segments, together with the O-RS segment, are inserted onto the amino acid sequence of the POI. One or more orthogonal O-RS / O-tRNAs allowing the introduction of an A residue ncAA Form a pair ; Including, The method optionally further comprises recovering the expressed POI.

[0237] E37: Objectives that include one or more non-canonical amino acid (ncAA) residues in their amino acid sequence A method for preparing a polypeptide (POI) of the present invention, the method comprising the steps of: and expressing the POI in the presence of the one or more ncAAs, the cell comprising: (i) a nucleotide sequence encoding the POI (CS POI ) (the one or more n of POIs cAA residues are encoded by a selector codon(s), (ii)CS POI The RNA-TP / O-RS fusion protein is functionally linked to the A target capable of interacting with at least one RNA-TP segment of a protein nucleotide sequence (TN); (iii)CS POI an anticodon (single) complementary to the selector codon (single or multiple) of or more) carrying an orthogonal tRNA ncAA (O-tRNA ncAA )molecule( The O-tRNA ncAA The molecule is an RNA-TP / O-RS fusion protein in cells. and one or more O-RS segments of the same quality on the amino acid sequence of the POI. One or more orthogonal O-RS / O-tRNAs allowing the introduction of one or more ncAA residues ncA A form pairs); Including, The method optionally further comprises recovering the expressed POI.

[0238] E38: Objectives that contain one or more non-canonical amino acid (ncAA) residues in their amino acid sequence 1. A method for preparing a polypeptide (POI) of claim 1, comprising: (a) The cells are divided into E1 to E3 and E4, each of which contains at least one RNA-TP segment. At least one AFP from any one of 7 to E15 and at least one O-RS segment A stem cell expressing one or more AFPs from any one of E1 to E3 and E7 to E15, including Top and; (b) one or more orthogonal tRNAs ncAA (O-tRNA ncAA ) molecules by the cells And it is expressed, - the orthogonal tRNA ncAA One or more of the O-RS segments of the molecule and the AFP Within the , one or more orthogonal aminoacyl-tRNA synthetases / tRNA ncAA (O -RS / O-tRNA ncAA ) form a pair, - the O-RS / O-tRNA ncAA The pair is a 1 on the amino acid sequence of the POI. It allows the introduction of one or more ncAA residues, Steps (a) and (b) may be simultaneous or sequential in any order. and; (c) then expressing the POI by the cell in the presence of the one or more ncAAs. death, - the nucleotide sequence encoding the POI (CS POI ) is a nucleotide sequence of one or more ncAA residues one or more selector codons encoding a group, - the selector codon is selected from the one or more O-tRNAs ncAA The anticodon of the molecule Match; - Said CS POI is functionally linked to the targeting nucleotide sequence (TN), and thus CS POI / TN fusion sequence is formed, - Said CS POI The / TN fusion sequence is capable of directing at least one of the AFPs in cells via the TN. a step capable of interacting with one RNA-TP segment; (d) optionally recovering the expressed POI; Includes.

[0239] E39: Objectives that include one or more non-canonical amino acid (ncAA) residues in their amino acid sequence 1. A method for preparing a polypeptide (POI) of claim 1, comprising: (a) Any one of the RNA-TP / O-RS fusion proteins from E18 to E23 was transfected into cells. and expressing the resultant; (b) one or more orthogonal tRNAs ncAA (O-tRNA ncAA ) molecules by the cells And it is expressed, - one or more O-RS segments of the RNA-TP / O-RS fusion protein and said cross-tRNA ncAA The molecule is capable of catalyzing one or more orthogonal aminoacyl-tRNA synthases in a cell. Sesatase / tRNA ncAA (O-RS / O-tRNA ncAA ) form a pair, - the O-RS / O-tRNA ncAA The pair is a 1 on the amino acid sequence of the POI. It allows the introduction of one or more ncAA residues, Steps (a) and (b) may be simultaneous or sequential in any order. and; (c) then expressing the POI by the cell in the presence of the one or more ncAAs. death, - the nucleotide sequence encoding the POI (CS POI ) is a nucleotide sequence of one or more ncAA residues one or more selector codons encoding a group, - the selector codon is selected from the one or more O-tRNAs ncAA The anticodon of the molecule Match; - Said CS POI is functionally linked to the targeting nucleotide sequence (TN), and thus CS POI / TN fusion sequence is formed, - Said CS POI The / TN fusion sequence mediates RNA-TP expression in cells via the TN. 1. Interact with at least one RNA-TP segment of the / O-RS fusion protein and the steps that can be taken; (d) optionally recovering the expressed POI; Includes.

[0240] E40: Any one of the methods of E36 to E39, wherein TN is a virus-coated protein. The viral RNA motifs bound by the protein, as well as the viral coat protein The polypeptide is selected from among functional fragments and variants thereof which retain the ability to be bound by the polypeptide.

[0241] E41: The method of any one of E36 to E40, wherein TN is: - MS2 R comprising an RNA sequence encoded by the nucleotide sequence of SEQ ID NO: 17 NA stem loop, or an amino acid sequence of SEQ ID NO: 17 having at least 60%, at least 7 0%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% amino acid sequence identity a functional fragment or variant thereof having the structure - BoxB comprising an RNA sequence encoded by the nucleotide sequence of SEQ ID NO: 18, or a sequence which is at least 60%, at least 70%, or at least 80% identical to the amino acid sequence of SEQ ID NO: 18; 0%, at least 90%, at least 95%, at least 96%, at least 97%, and its functional fragments having at least 98% or at least 99% amino acid sequence identity. Fragments or mutants, and - exists in at least two different versions and in particular has SEQ ID NO: 289 or RNA corresponding to (encoded by) the nucleotide (DNA) sequence of sequence number 290 The RNA sequence encoded by the nucleotide sequence of the polynucleotide having the sequence pp7 RNA stem loop or the amino acid sequence of SEQ ID NO: 289 or 290 at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least a functional fragment or variant thereof having 99% amino acid sequence identity; is selected from.

[0242] E42. Any one of the methods of E36 to E41, further comprising: The selector codon(s) encoding the amber, ochre, and The opal stop codon is selected from the

[0243] E43: A nucleic acid molecule comprising: (i) a nucleotide sequence encoding a polypeptide of interest (POI) (CS POI )and( The POI is linked to a CS by a selector codon POI One or more non-standard addresses coded above containing non-amino acid (ncAA) residues), (ii) a targeting nucleotide sequence (TN) and (an RNA molecule comprising said TN) , which can interact with an RNA-targeting polypeptide (RNA-TP) via Includes.

[0244] E44: A nucleic acid molecule of E43, in which TN is bound by the viral coat protein The viral RNA motifs that are bound by the viral coat protein The polypeptides are selected from functional fragments and variants thereof that retain potency.

[0245] E45: A nucleic acid molecule of E43 or E44, wherein TN is: - MS2 R comprising an RNA sequence encoded by the nucleotide sequence of SEQ ID NO: 17 NA stem loop, or an amino acid sequence of SEQ ID NO: 17 having at least 60%, at least 7 0%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% amino acid sequence identity a functional fragment or variant thereof having the structure - BoxB comprising an RNA sequence encoded by the nucleotide sequence of SEQ ID NO: 18, or a sequence which is at least 60%, at least 70%, or at least 80% identical to the amino acid sequence of SEQ ID NO: 18; 0%, at least 90%, at least 95%, at least 96%, at least 97%, and its functional fragments having at least 98% or at least 99% amino acid sequence identity. Fragments or mutants; and - exists in at least two different versions and in particular has SEQ ID NO: 289 or RNA corresponding to (encoded by) the nucleotide (DNA) sequence of sequence number 290 The RNA sequence encoded by the nucleotide sequence of the polynucleotide having the sequence pp7 RNA stem loop or the amino acid sequence of SEQ ID NO: 289 or 290 at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least a functional fragment or variant thereof having 99% amino acid sequence identity; is selected from.

[0246] E46: any one of the nucleic acid molecules E43 to E45, wherein the ncAA residue of the POI ( The selector codon(s) encoding the amber, ochre, and the opal stop codon.

[0247] E47: A polypeptide of interest having at least one non-standard amino acid (ncAA) residue. 1. A kit for preparing a POI comprising: - at least one ncAA corresponding to at least one ncAA residue of the POI or Salt and - at least one expression vector for E27, Includes.

[0248] E48: The kit of E47, wherein the expression vector comprises at least one O-RS segment. The present invention encodes a fusion protein comprising an RNA-TP fragment and at least one RNA-TP segment.

[0249] E49: A kit of E47 or E48, further comprising an orthogonal tRNA ncAA (O-tR NA ncAA ) molecule.

[0250] E50: Any one of the kits E47 to E49, further comprising a multicloning At least one expression vector comprising a site and a targeting nucleotide sequence (TN), The RNA molecule containing the TN is targeted to an RNA targeting polypeptide (RNA-T P).

[0251] Any one of the above embodiments may also include the following modifications: (i.e., IC-TP and PSP) segments and / or EPs (RNA-TP or O- The RS segment further comprises a synthetic protein that induces and controls macromolecular interactions. Can be combined with segments such as 2, 3, 4, 5, 6, 7, 8, 9, or 10. Such one or more, preferably one, such protein segments form the AFP vector of the present invention. can be functionally fused to form heterodimeric coiled-coil protein structures Of particular interest in the context of the present invention are SYNZIPs that have the ability to IPs are pairs of synthetic peptides that can interact with each other to form macromolecular interactions. Used to induce and control. Non-limiting examples include SYNZIP 1:2; SYNZIP 3:4; and SYNZIP 5:6 pairs. nke, AW, Grant, RA, Keating, AE (2010)J Hetero specific compounds described by Am Chem Soc 132 6025-6031 Heterogeneous coiled coil pair SYNZIP2:SYNZIP1 (SYNZIP1: SEQ ID NO:3) 12, SYNZIP2: SEQ ID NO: 314, SYNZIP3: SEQ ID NO: 316, SYNZI P4: SEQ ID NO: 318, as well as functional fragments of the SYNZIP polypeptide thereof and variants are particularly preferred. The functional fragments and variants are At least 60%, at least 70%, at least 80%, or at least at least 90%, at least 95%, at least 96%, at least 97%, at least 9 It may contain 8%, or at least 99% amino acid sequence identity.

[0252] The present invention is further illustrated by the following non-limiting examples. EXAMPLES

[0253] method

[0254] (A) Cell culture, transfection, and ncAA feeding.

[0255] HEK293T cells (ATCC CRL-3216) and COS-7 cells (ATCC, CRL-1651) was cultured in Dulbecco's modified Eagle's medium (Life Technologies, Inc.) The cultures were maintained with 1% penicillin-streptomycin. (Sigma, 10,000U / ml penicillin, 10mg / ml streptomycin, 0.9% NaCl), 2 mM L-glutamine (Sigma), 1 mM sodium pyruvate (Life Technologies), and supplemented with 10% FBS (Sigma). The cells were incubated at 37°C with 5% CO 2 The cells were cultured in an ambient atmosphere for 15-20 passages every 2-3 days. The cells were passaged in .

[0256] In either case, 15–20 h before transfection, Cells were seeded at 70-80% confluency in the culture medium. Flow cytometry was performed on the plus 24-well plates with plastic bottom (Nunclon Delta Surface Immunofluorescence labeling and FISH were performed using a fluorochrome plated ... 24-well plates with lath bottom (Greiner Bio-One) or 4-well L Performed on ab-Tek #1.0 borosilicate coverglass (ThermoFisher) Ta.

[0257] HEK293T cell transfection was performed with 3 μg PEI per 1 μg DNA. This was performed with polyethyleneimine (PEI, Sigma-Aldrich) using COS-7 cells were cultured using JetPrime reagent (PeqLab) according to the manufacturer's recommendations. A 1:2 ratio was used for transfection.

[0258] For the amber suppression assay, cells were cultured at POI TAG Vector, tRNA Pyl , S The plasmid was transfected with either the ribosomal RNAi gene or the MCP or mock construct in a ratio of 1:1:1:1. Four to six hours after transfection, the medium was replaced with fresh ncAA-containing It became something like this.

[0259] All stock and working solutions of ncAA used were prepared according to the methods described in Nikic et al. As described in (Nat Protoc 10(5):780-791,2015) SCO (cyclooctyne lysine, SiChem SC-8000) was prepared at 250 3-Iodophenylalanine (Chem-Impex Intermediate) was used at a final concentration of 1 μM. (Rnational Inc.) was used at a final concentration of 1 mM. SCO was used as a PylRS AF (Y306A,Y384F) is efficiently recognized by Angew Chem 2011,50:3878-3881). Luaranine is PylRS AA (C346A,N348A) (Wan g et al., ACS Chem Biol 2013, 8:405-415) .

[0260] (B) Flow cytometry

[0261] HEK293T cells were harvested one day after transfection and resuspended in 1x PBS. The cells were then passed through a 100 μm nylon mesh. Infection was performed with 1.2 μg total DNA in a 1:1:1:1 ratio: - POI (a stop codon that encodes the amino acid position to be occupied by the ncAA) A reporter plasmid encoding - a matching (i.e., reverse complement) antisense codon in the POI coding sequence tRNA with codon Pyl(From now on, we will simply refer to tRNA Pyl The program that codes Rasmid; - a plasmid encoding PylRS or a functional variant thereof, respectively, and - either a plasmid encoding an MCP fusion polypeptide or a mock plasmid, This was done by. 4–6 h after transfection, the cell culture medium was diluted with ncDNA that should be incorporated into the POI. The medium was replaced with fresh medium containing AA and left until harvesting.

[0262] Data acquisition and analysis were performed by LSRFortessa SORP Cell Analysis. er (Becton, Dickinson and Company) and FlowJo The software (FlowJo) was used. First, forward scatter area (FSC-A) and Gate cells by cell type using side scatter area (SSC-A) and side scatter area (SSC-A) parameters. Single cells were then identified based on SSC-A and side scatter width (SSC-W). Each FFC plot shown is the sum of three independent biological replicates. From these, means and SEMs were calculated. At least 130,000 units per condition were collected. Cells were analyzed. GFP fluorescence was measured in the 488-530 / 30 channel, and mCherry fluorescence was measured in the 488-530 / 30 channel. was acquired on channel 561-610 / 20.

[0263] (C) PylRS immunostaining and imaging, fluorescence in situ hybridization (FISH)

[0264] For immunolabeling experiments, cells were rinsed with 1x PBS and resuspended in 2% paraformaldehyde in 1x PBS. The cells were fixed with methylaldehyde for 10 min at RT, rinsed again with 1x PBS, and then Then, permeabilized with 0.5% Triton X in 1x PBS for 15 min at RT. The permeabilized cell samples were rinsed twice with 1× PBS and then The cells were then incubated for 90 min in blocking solution (3% in 1x PBS for 90 min at room temperature). BSA), then 1 μg / ml primary antibody (Nikic et al. (Angew C hem Int Ed Engl 2016,55(52):16172-16176) Polyclonal rat anti-PylRS and / or polyclonal monoclonal rabbit anti-MCP (Merck, ABE76), and / or monoclonal rabbit Anti-RPL26L1 antibody (EPR8478, Abcam, ab137046) and blockade The cell samples were incubated in 1× PBS overnight at 4°C. Rinse with BS and cross-absorb with 2 μg / ml secondary antibody (chicken anti-rat IgG (H+L) Alexa Fluor 594-conjugated antibody (Thermo Fisher Scientific) Scientific, A-21471) and / or goat anti-rabbit IgG (H+L) conjugate Differentially adsorbed Alexa Fluor 647-conjugated F(ab') 2 (Ther mo Fisher Scientific, A-21246)) in blocking solution The DNA was purified by Hoechst chromatography and incubated at RT for 60 min. Stained with 33342 (1 μg / ml in 1× PBS) for 10 min at RT. If only DNA was stained, cells were fixed and permeabilized as described above. Then, Hoechst 33342 (1 μg / ml in 1× PBS) was added for 10 min. The staining was carried out at RT for 1 h. Finally, the cells were rinsed twice with 1× PBS.

[0265] FISH experiments were performed as described by Nikic et al. (Angew Chem Int Ed. Engl 2016,55(52):16172-16176) The experiment was performed 1 day after transfection, similar to that performed by Pierce et al. ethods Cell Biol 122:415-436, 2014) The reduction protocol was adapted for 24-well plates.

[0266] tRNA Pyl For imaging of only the 5' hybridization probe -CTAACCCGGCTGAACGGATTTAGAGTCCATTCGATC-3' (labeled with Cy5 at the 5' end; SEQ ID NO: 1) was used at 0.25 μM. Four washes with SC and TN buffer (0.1 M TrisHCl, 150 mM NaCl After one wash with l), the cells were incubated for 1 h prior to standard immunofluorescence labelling as described above. The sections were then incubated with 3% BSA in TN buffer at RT for 1 h.

[0267] tRNA Pyl For imaging of both the MS2 RNA stem-loop sequence, tRNA Pyl Hybridization probe (5'-terminated by digoxigenin) Labeled 5'-CTAACCCGGCTGAACGGATTTAGAGTCCATTC GATC-3'; SEQ ID NO:2) at 0.16 μM, and the MS2 RNA stem-loop sequence Hybridization probe (Alexa Fluor 647 at the 5' end) Labeled 5'-CTGCAGACATGGGTGATCCTCATGTTTTCTA -3'; SEQ ID NO: 3) was used at 0.75 μM. After four washes with SSC, the cells were Blocking buffer (0.1M TrisHCl, 150mM NaCl, 1x blocking The ink was then incubated in a 25 ml ethanol solution (Sigma 11096176001) for 1 h at RT. The cells were then incubated with fluorescein-conjugated sheep anti-digoxigenin. Nin Fab (Sigma 11207741910) in blocking buffer The plate was incubated at 4°C overnight at a 1:200 dilution. The next day, the plate was diluted with Tween buffer (0. 1M TrisHCl, 150mM NaCl, 0.5% Tween20) for 5 min. DNA was diluted with Hoechst 33342 (1 μg in 1× PBS) and washed three times in between. The cells were stained with 10 ml of 1:1 glycerol (1:1 / ml) for 10 min at RT.

[0268] Confocal images were taken using a 63x / 1.40 oil immersion objective, using the following laser lines for excitation: Acquired with a Leica SP8 STED 3X microscope equipped with Hoechst 33342 is 405 nm, fluorescein and GFP are 488 nm, mOrange is 548nm, Alexa Fluor 594 is 594nm, Alexa Fluor 647 and Cy5 at 647 nm. Emission light was measured by the HyD detector at 420–50 nm, respectively. The spectra were collected at 0 nm and 605-680 nm.

[0269] An Olympus 1000 sb / s 200 nm microscope equipped with a 60x / 1.40 oil immersion objective was used with the following laser lines for excitation: Ribosome immunofluorescence images were taken with an mpus Fluoroview FV3000 microscope. The fluorescence was measured using the following fluorescence spectrophotometers: GFP at 488 nm, Alexa Fluor 594 at 594 nm, and Alexa Fluor 647 is 640 nm.

[0270] Images were processed using FIJI software.

[0271] (D) Constructs, cloning, and mutagenesis.

[0272] Two different fluorescent protein reporters (dual-color reporters) are used to detect one protein. one reporter in the multicloning site and another reporter in the other multicloning site. It was cloned into the pBI-CMV1 vector (Clontech 631630). One CDS of the reporter is a sequence of two MS2 RNA stems fused to the 3' untranslated region. The authors generated mRNAs carrying the loop ("MS2 tag") but not other reporters. The mRNAs detected were not MS2 tagged.

[0273] GFP reporter for amber suppression studies 39TAG and mCherry 185 TAG was used as an N-terminal fusion with an NLS. Similar constructs were prepared for the expression of GFP and 39TAA and mCherry 18 5TAA , GFP 39TGA and mCherry 185TGA by).

[0274] NLS::GFP 39TAG ::MS2 tag reporter:NLS::GFP 39TAG as a reporter for successful amber suppression in imaging experiments, and pBI -Cloned with two copies of the MS2 RNA stem loop onto the CMV1 vector .

[0275] To study suppression of multiple amber codons, a second multiple cloning site was inserted. GFP without a reporter (e.g. mCherry) 39,149TAG and G.F. P 39,149,182TAG A pBI-CMV construct was prepared.

[0276] Further non-limiting examples of GFPs that are applicable in the context of the present invention are: GFP 66TAG GFP with amber site (SEQ ID NO: 238) GFP 66TCG GFP with serine site (SEQ ID NO: 240) GFP 66CCG GFP with proline site (SEQ ID NO: 242) GFP 66CTA GFP with leucine site (SEQ ID NO: 244) GFP 66TTA GFP with leucine site (SEQ ID NO: 246) GFP 66ATA GFP with isoleucine site (SEQ ID NO: 248) GFP 66CGG GFP with arginine site (SEQ ID NO: 250) GFP 39TCG GFP with serine site (SEQ ID NO: 252) GFP 39CCG GFP with proline site (SEQ ID NO: 254) GFP 39CTA GFP with leucine site (SEQ ID NO: 256) GFP 39CGG GFP with arginine site (SEQ ID NO: 258) GFP 39TCG LCK-GFP with serine site (SEQ ID NO: 278) GFP 39CCG LCK-GFP with proline site (SEQ ID NO: 280) GFP 39CTA LCK-GFP with leucine site (SEQ ID NO: 282) Extended GFP 39TCG GFP 66CCGA serine site at position 39 was genetically fused to GFP (SEQ ID NO: 284) Extended GFP 39CCG GFP 66TCG A proline site at position 39 fused to GFP (SEQ ID NO: 286) Extended GFP 39CTA GFP 66TCG A leucine site at position 39 fused to and GFP (SEQ ID NO: 288).

[0277] Further non-limiting examples of mCherry that are applicable in the context of the present invention are: mCherry 72TAG mCherry with amber site (SEQ ID NO: 260) mCherry 72TCG mCherry with serine site (SEQ ID NO: 262) mCherry 72CCG mCherry with a proline site (SEQ ID NO: 264) mCherry 72CTA mCherry with leucine site (SEQ ID NO: 266) mCherry 72TTA mCherry with leucine site (SEQ ID NO: 268) mCherry 72ATA mCherry with an isoleucine site (SEQ ID NO: 270 ) mCherry 185TCG mCherry with serine site (SEQ ID NO: 272) mCherry 185CCG mCherry with a proline site (SEQ ID NO: 274) mCherry 185CTA mCherry with leucine site (SEQ ID NO: 276) It is.

[0278] mCherry constructs containing different TN loops that are applicable in the context of the present invention Further non-limiting examples are: mCherry 190TAG -2xPP7 Contains an amber site and 2xpp7 loops mCherry (SEQ ID NO: 216) mCherry 190TAG -4xPP7 Contains an amber site and 4xpp7 loops mCherry (SEQ ID NO: 218) mCherry 190TAG -6xPP7 Contains an amber site and a 6xpp7 loop mCherry (SEQ ID NO: 220) H2B-mCherry 190TAG -2xMS2 amber sites and 2xms2 loops Human histone H2B1-J form (Uniprot:P 06899) ​​(sequence number 222).

[0279] They are particularly useful for the synthesis of AFP molecules that are free of any of the other polypeptide segments (AP and EP) of the AFP molecule. or one of the AFP molecules described herein at a position on the fusion molecule that does not inhibit the function of the AFP molecule. Such an epitope tag-containing AFP can be fused to either of the polypeptide chains. An example molecule is shown below.

[0280] The OT assembly construct was prepared as follows: tRNA Pyl is the human U6 promoter All other constructs were cloned under the control of pcDNA3.1 (Invit The gene was cloned under the CMV promoter in the vector V86020. The MCP protein was from addgene plasmid #31230, and FUS was from Addgene plasmid #31230. All FUS fusions were cloned from the ene plasmid #26374. The C-terminal NLS region was replaced with a Flag tag using amino acid 1-478 (S108N). All RS fusions expressed the efficient NES::PylRS gene as previously reported. AF (Y30 6A,Y384F) sequence was used (see, e.g., Nikic et al., Angew Chem Int Ed Engl 2016,55(52):16172-1617 6). PylRS mutant PylRS AA (N346A,C348A) is the wild-type Py Starting from the lRS, the SPD5 gene was cloned by site-directed mutagenesis. Order from enewiz and generate MCP and PylRS by restriction cloning. AF merged into . KIF13A 1-411 and KIF16B 1-400 Cloned from human cDNA and inserted into pcDNA3.1 by restriction cloning. 1-411 The P390 of MCP was deleted by site-directed mutagenesis. AF , E.W.S.R. 1::MCP, FUS::PylRS AF , FUS::PylRS AA , SPD5::M CP and SPD5::PylRS AF With KIF13A 1-411 , ΔP390 and K IF16B -400 Fusions were assembled by Gibson assembly (Gibson et al., Nat Methods 2009, 6:343-345).

[0281] Construct for differential imaging experiments: Nup153-EGFP 149TAG and Vi m 116TAG -To selectively express mOrange, one gene was first Inserted onto pBI-CMV1 together with the MS2 tag (Nikic et al., Ang ew Chem Int Ed Engl 2016,55(52):16172-16 (See 176 for comparison). Then, other genes were inserted without the MS2 tag. 3::EGFP 149TAG and Vim 116TAG ::mOrange::MS2 Tag Vim on a pBI vector with 116TAG By replacing -mOrange So, INSR 676TAG ::mOrange is merged with the MS2 tag, INSR 67 6TAG: ::mOrange on one cassette and Nup153::EGFP 149 TAG On the other cassette, a bicistronic vector was obtained.

[0282] Multicistronic amber suppression vector for COS-7 cell experiments: COS-7 cells Since OT has a lower transfection efficiency, we used the components of OT assembly. A multicistronic vector harboring the multicistronic amber suppression vector was generated. To construct the target gene, tRNA was placed under the control of the human U6 promoter. Pyl No. 1 One copy of was inserted onto the pBI-CMV1 vector by Gibson assembly. Then, first, AFP CDS KIF16B::FUS::PylRS AF And the most Later, AFP CDS KIF16B::EWSR1::MCP was used as a Gibson assembly. Alternatively, NES::PylRS under the CMV promoter was inserted by AF Reach and tRNA under the human U6 promoter Pyl The previously published pcDNA3.1 expressing Based on the construct (Nikic et al., Angew Chem Int Ed E ngl 2016,55(52):16172-16176) was used. , U6-tRNA Pyl , KIF16B::FUS::PylRS AF , and KIF16 B::EWSR1::MCP or NES::PylRS AF The construct was transformed into pDon or vector (GeneCopoeia). The sequence information for each AFP used in the following experiments is taken from the sequence listing below. can.

[0283] Example 1 - RNA-TP / O-RS fusion and AFP containing a single AP

[0284] Engineer an OT assembly ("OT organelle", Figure 1) that has the following components: Ta:

[0285] i) An mRNA targeting system comprising two MS2 RNA stem loops (MS2 tags) , fused to a selected mRNA encoding the POI to generate an mRNA::ms2 fusion. The MS2 tag specifically binds to the MS2 bacteriophage coat protein (MCP). (Bertrand et al., Mol Cell 1998, 2:437-44 5), which therefore represents a stable and specific mRNA::ms2-MC The MS2 tag is always located in the 3' untranslated region (3'UTR) of the mRNA. ) This ensured that the translation produced a final POI that was trace-free.

[0286] ii) tRNA / RS suppressor pair. Methanosarcina mazei pyrrolidin Orthogonal tRNA / RS pairs from the tRNA Pyl / PylRS) because , which has been shown to be effective in multiple cell types, including E. coli, mammalian cells, and even live mice. Using GCE, we have successfully synthesized over 200 n This is because it allows the coding of cAA (e.g., Liu et al., Annu Rev Biochem 2010,79:413-444;Lemke,C hemBioChem 2014,15:1691-1694;Chin,Nature 2017,550;53-60).

[0287] iii) The assembler (AP) is the main structure required to form the OT assembly. The purpose of the assembler was to create a membrane or other structure in the form of a dense phase, aggregates, droplets, or condensates. The aim of this study was to create a structure in which the mRNA::ms2-MCP complex was RNA Pyl / PylRS pair in close proximity.

[0288] The simplest strategy tested was a bimolecular fusion of MCP::PylRS (designated B). In addition, strategies expected to generate larger assemblies were tested. All of these assembly systems involve PylR coexpressed with assembler fusions to MCP. It was constructed by integrating assembler to S. CP was predicted to form large aggregates (coexpression is indicated herein by “·”). One tested assembly strategy was based on protein phase separation, and one was based on protein cleavage. Based on the assembly of necin, these are abbreviated herein as P and K, respectively (Figure 1). 2A). Furthermore, in each P and K approach, two different molecular designs (P1 and P2, respectively) were used. , and K1, K2) were tested. These are summarized as follows:

[0289] P1. Previous studies have demonstrated that fusion of sarcomas (FUS) and fusion of sarcomas (FUS) form mixed droplet-like structures by phase separation. and established the potential of Ewing sarcoma breakpoint region 1 (EWSR1) protein. Both of them exhibit prion-like denatured domains that facilitate phase separation into liquid, gel, and solid states. (e.g., Altmeyer et al., Nat Commun 2 015,6:8088;Patel et al.,Cell 2015,162:10 In the phase-separated state, these proteins are expressed in the cytoplasmic residues. FUS is highly enriched locally (by several orders of magnitude) compared to the soluble fraction of the nucleosome. This resulted in the synthesis of EWSR1, which was highly enriched for MCP and PylRS. It was predicted that this would lead to the formation of droplets. It is expressed as ::MCP.

[0290] P2. Caenorhabditis elegans protein spindle-defective protein Protein 5 (SPD5) in particular is capable of phase separation into large (several micron-sized) droplets. As shown (Woodruff et al., Cell 2017, 169:106 In the phase-separated state, SPD5 is expressed in the cytoplasmic nucleosomes. The SPD5 fused tag is highly enriched locally (by several orders of magnitude) compared to the soluble fraction. It was predicted that the proteins would condense into droplets similar to the FUS-EWSR1 droplets. PylRS fused to SPD5 and MCP fused to SPD5 were highly enriched. P2 was expressed as SPD5::PylRS·SPD5::MCP. .

[0291] K1. Kinesin shortening constitutively toward the microtubule plus ends in living cells Go to (Soppina et al., Proc Natl Acad Sci USA 2014, 111:5562-5567). One such truncated kinesin is KIF13A 1-411,ΔP390 and co-expressed with this kinesin truncation, respectively. The expressed PylRS and MCP are localized due to spatial targeting to the microtubule plus ends. It was predicted that K1 would be concentrated in KIF13A. 1-411,ΔP 390 ::PylRS·KIF13A 1-411,ΔP390 It is expressed as ::MCP.

[0292] K2. Similar to K1, a truncated kinesin, KIF16B 1-400 was also tested. 2 is KIF16B 1-400 ::PylRS·KIF16B 1-400 ::MCP and Tables will be done.

[0293] These assays are intended to facilitate functional orthogonal translation of MS2-tagged mRNAs. To assess the chromatin dynamics, a dual reporter construct was designed in which GFP and and mCherry variants were co-expressed from two different expression cassettes from a single plasmid. To ensure that the mRNA ratio between them is constant across all experiments, A stop codon was inserted at the permissive site to GFP at position 39 (GFP 39終止 ), The mCherry was introduced at position 185 (mCherry 185終止 ;Figure 2B). Only if the stop codon is successfully suppressed is the corresponding green or red fluorescent protein produced. Transfected cells (unless otherwise specified) will Py l and ncAA were always present) were analyzed by fluorescence flow cytometry (FFC). GFP and mCherry cannot distinguish between the mRNAs in the cytoplasm. When expressed from this plasmid using the system, an approximate diagonal line appears in the FFC plot. The settings were adjusted to allow selective and functional OT organelles to be produced. , only when the MS2 tag is fused to the 3'UTR of mCherry mRNA herry is selectively expressed, leading to the appearance of vertical lines in the cytometry plots. (Figure 2B). Unless otherwise reported, all experiments were performed using tRNA Pyl And widely used In the presence of the novel and well-characterized lysine derivative ncAA SCO The side chains of these amines are equipped with various cloning agents for installing various chemical groups onto proteins. The cyclooctyne supports a cyclooctyne that can be used in cyclohexyl chemistry. As previously reported, This ncAA was efficiently encoded by the Y306A, Y384F double mutant of PylRS. (For simplicity, this variant is referred to herein as P (Nikic et al., Angew Chem 2014) ,53:2245-2249;Plass,Angew Chem 2012,51:4 166-4170;Plass et al.,Angew Chem 2011,50 ncAA omission served as a standard negative control, and G This does not result in expression of FP or mCherry.

[0294] The performance of each OT system was evaluated according to its selectivity and relative efficiency. Selectivity was determined by the average G Defined as the ratio r of the average mCherry FFC signal divided by the FP signal The final value is expressed as ~-fold selectivity relative to that of cytoplasmic PylRS. The relative efficiency is calculated using a cell line that serves as a reference (defined here as 100%). The average mCherry signal of each line divided by the average mCherry signal of the quality PylRS line. The y signal is defined as the y-signal of the target protein. Selectivity (dark grey positive bars) and efficiency (light grey negative bars) are All results for the selected FFCs are summarized in the bar graph in Figure 2C. The data are also shown in Figure 2D.

[0295] The simplest strategy B (MCP fused to PylRS) showed a selectivity gain of about 1.5-fold. OT system P1 (based on the phase separation of FUS / EWSR1) showed somewhat lower The P2 system (based on SPD5) had approximately 2-fold selectivity gain (Fig. 2C, D). The results showed selectivity gain (Figure 2C). A two-fold increase in selectivity was observed for K1 (Figure 2C). The K2 system behaved similarly (Fig. 2C, D). Overall, the selectivity gain was relatively small. However, it was robustly detected and distinguishable from a simple drop in efficiency. The effect of amber suppression (data not shown) was due to the titration of amber suppression efficiency (specifically, 0.48ng, 2.4ng, 12ng, 60ng, or 300ng of tRNA, respectively. P ylThe ncAA aminoacylation activity (i.e., That is, tRNAPyl / PylRS in the presence of ncAA is brought into direct proximity to the target mRNA. These results suggest that the suppression of codons represents a route to more selective codon suppression.

[0296] Example 2 - AFP containing a combination of two APs

[0297] AFPs containing the AP combinations described in Example 1 were tested in a similar manner. : K1::P1= KIF13A 1-411,ΔP390 ::FUS::PylRS·KI F13A 1-411,ΔP390 ::EWSR1::MCP, K2::P1=KIF16B 1-400 ::FUS::PylRS·KIF16B 1-4 00 ::EWSR1::MCP, K1::P2=KIF13A 1-411,ΔP390 ::SPD5::PylRS·KI F13A 1-411,ΔP390 ::SPD5::MCP, K2::P2=KIF16B 1-400 ::SPD5::PylRS·KIF16B 1- 400 ::SPD5::MCP.

[0298] For all combinations, at least a 5-fold selectivity gain was observed, indicating orthogonal translation. The best performing of these lines was KIF16B 1-400 FUS / EW with Based on the SR1 fusion K2::P1, an 8-fold selectivity was observed (box in Figure 2C). This was also directly evident from the FFC data, where the bright mCher The ry-positive cell population remained distinct, with minimal GFP expression (arrow in Fig. 2D ).

[0299] Example 3 - AFPs containing combinations of APs including membrane-targeted APs

[0300] a phase separation polypeptide (PSP), FUS, optionally fused to a SYNZIP segment; and EWSR1 (also referred to herein as EWS)-derived APs, and membrane-targeted Different APs LcK, EB1, CG1, and EBAG9 act as activating signals 全長 , E.B. AG9 1-29 , CMP Sia Tr, P450 2C1 1-27 , and P450 2 C1 1-29 The AFPs containing the combinations were tested in a manner similar to that of Example 2.

[0301] LcK is a plasma membrane targeting signal (Resh, Bba-Mol Cell Res 1999,1451:1-16), which post-translationally adds an amphipathic helix to the POI In these experiments, AFP LcK::FUS::PylRS and LcK::EW SR1::MCP was co-expressed by HE293T cells (see Figures 3 and 6C). Testing of this system with a heavy reporter revealed that the MS2-tagged m This resulted in a dramatic shift in signal and strong selectivity for Cherry expression. See Figures 4 and 5, which show a 26-fold selectivity gain compared to the control. IF and FISH for tRNA and tRNA revealed the appearance of occasional droplet-like structures and all constitutive The results show a clear membrane signal with complete colocalization of the IL-1 and IL-2 components. Without wishing to be bound by theory, targeting the OT system to membranes is a 2D surface This results in the confinement of the components to the surface (i.e., a film), which is more stable than the cytoplasmic droplet. It is assumed that the two combined assemblies will provide high spatial isolation for the The choice of a combination strategy (LcK for membrane targeting and FUS / EWSR1 for droplet generation) According to the cumulative effect of the FUS / EWSR1 “assembler”, the presence of the FUS / EWSR1 “assembler” may lead to selective amber This indicates that LcK fusion (and hence the membrane anchor system) is not a requirement for obtaining repression. (data not shown). Nevertheless, the LcK targeting of FUS / EWSR1 The combination of MS2 on the fluorescent reporter resulted in higher selectivity of the system. Swapping tags results in swapped selectivities in the FFC data. was found, highlighting the alternative (orthogonal) translation of MS2-tagged mRNAs.

[0302] For further LcK-based experiments, the AFP construct LcK::FUS::SYNZI P1::PylRS and EWSR1::SYNZIP2::MCP were transfected into HE293T cells. Testing of this system with the same dual reporter was performed exclusively with MS 2 resulted in a dramatic shift in the tagged mCherry signal and strong selectivity of expression Upon expression, SYNZIP1 and 2 pair up and mediate the transport of MCPs to the plasma membrane-based OT domain. Recruitment to Luganella. AFP construct lacking SYNZIP1, LcK::FUS: Comparative approach to co-expressing PylRS and EWSR1::SYNZIP2::MCP No translational selectivity could be observed (see FIG. 8B).

[0303] EB1 is a microtubule plus end targeting signal (Nehlig A, Molina A, Rodrigues-Ferreira S, Honore S, Nahmias C.Regulation of end-binding protein EB1 in the control of microtubule dynamics. Cell Mol Life Sci.2017;74(13):2381-2393. doi:10.1007 / s00018-017-2476-2). In these experiments, The AFP construct EB1::PylRS was combined with EB1::MCP and EB1::FUS::Pyl RS with EB1::EWSR1::MCP or EB1::FUS::MCP with HE2 Testing of this system with the same dual reporter was performed using the control PylR The signal shift and expression of MS2-tagged mCherry were observed exclusively in comparison with S. This resulted in strong selectivity for the β-lactamase (Figure 6B).

[0304] CG1 is a nuclear envelope targeting signal (Kim SJ, Fernandez-Martin inez J, Nudelman I, et al.Integrative stru cture and functional anatomy of a nucleus r pore complex.Nature.2018;555(7697):475 -482.doi:10.1038 / nature26003). In these experiments, A The FP constructs CG1::FUS::PylRS and CG1::EWSR1::MCP were Testing of this system with the same dual reporter was performed using the control Py Compared to lRS, the signal shift was observed exclusively for expression of MS2-tagged mCherry. This resulted in strong selectivity and saturation of the nucleoside analogues (see Figure 6E).

[0305] EBAG9 全長 and EBAG9 1-29 is a Golgi membrane targeting signal (Enge lsberg A, Hermosilla R, Karsten U, Schulein R, Dorken B, Rehm A. The Golgi protein RC AS1 controls cell surface expression of tumor-associated O-linked glycan antigen sJ Biol Chem.2003;278(25):22998-23007.d oi:10.1074 / jbc.M301361200). In these experiments, the AFP structure Building EBAG9 1-29 ::FUS::PylRS and EBAG9 1-29 ::EWSR 1::MCP was co-expressed by HE293T cells. This system with the same dual reporter The study focused exclusively on the expression of MS2-tagged mCherry compared to the control PylRS. This resulted in a shift in the signal and strong selectivity, see Figure 6F (left).

[0306] CMP Sia Tr is a Golgi membrane targeting signal (Eckhardt M, G otza B,Gerardy-Schahn R.Membrane topolog y of the mammalian CMP-sialic acid trans porter.J Biol Chem. 1999;274(13):8779-87 87.doi:10.1074 / jbc.274.13.8779). In these experiments , AFP constructs CMP Sia Tr::FUS::PylRS and CMP Sia T r::MCP was co-expressed by HE293T cells. This system with the same dual reporter The study focused exclusively on the expression of MS2-tagged mCherry compared to the control PylRS. This resulted in a shift in the signal and strong selectivity, see Figure 6F (right).

[0307] P450 2C1 1-27 is an ER membrane targeting signal (Fazal FM, Han S, Parker KR, et al. Atlas of Subcellular RNA Localization Revealed by APEX-Seq.Ce ll.2019;178(2):473-490.e26.doi:10.1016 / j In these experiments, the AFP construct P450 2 C1 1-27 ::FUS::PylRS and P450 2C1 1-27 ::EWSR1: :MCP or P450 2C1 1-29 ::FUS::MCP::PylRS HE29 Testing of this system with the same dual reporter was performed using the control PylR The signal shift and expression of MS2-tagged mCherry were observed exclusively in comparison with S. This resulted in strong selectivity for the β-lactamase (Figure 6G).

[0308] Example 4 - Selectivity gain is specific to the interaction of mRNA MS2 tag with MCP Verification

[0309] The observed selectivity gain is specific to the interaction of the MCP segment with the MS2 tag of the mRNA. To verify that the OT systems are different, we performed RS analysis of each OT system without MCP. The assembler fusion was expressed and characterized. However, selective orthogonal translation of MS2-tagged mRNAs was not observed in these cases. In addition, the MS2 tag was transferred from mCherry to the dual-color reporter The reporter was inverted by moving it into a GFP cassette. As previously reported, this reversed the selectivity of the system to dominant GFP expression (data not shown). , we established that the OT system acts selectively on MS2-tagged RNA.

[0310] Example 5 - Introducing multiple ncAAs on the same POI

[0311] GCE can also be used to introduce multiple ncAAs onto the same POI. (For example, Liu et al., Annu Rev Biochem 2010,79 :413-444;Lemke,ChemBioChem 2014,15:1691- 1694; Chin, Nature 2017, 550; 53-60). However, only a very few publications have reported more than one on the same protein in eukaryotes. This is because, compared with single-codon suppression, The yields are typically poor (Xiao et al., Angew Chem. 2013,52:14080-14083;Schmied et al.,J Am Chem Soc 2014,136:15577-15583;Zhang et al.,Biochem Biophys Res Co 2017,489:490- 496). It is noteworthy that even double and triple amber proteins still express OT Luganella (data not shown).

[0312] Example 6-3-OT with iodophenylalanine

[0313] To ensure that other ncAAs can also be translated by OT assembly To determine whether this is the case, we tested another structurally distinct ncAA (3-iodophenylalanine). This is a phenylalanine derivative instead of a lysine derivative (e.g. SCO), and a different Encoded by the PylRS mutant (N346A, C348A) (Wang et al. (see ACS Chem Biol 2013, 8:405-415). Results were also observed for this system (Figure 2C).

[0314] Example 7 - OT with different selector codons

[0315] Opal and ochre codons are highly abundant in eukaryotic genomes (Human genome In eukaryotes, the amber codon is cleaved by GCE in 52% of cases and 28% of cases by opal. In addition, the removal of these codons throughout the eukaryotic genome Genomic approaches to orthogonal translation are even more challenging than amber codons, and current technology However, in the OT system of the present invention, tRNA Pyl Anticodon of in the loop and at each codon on the MS2-tagged POI-encoding mRNA Simple mutations in the nucleotide sequence allow for orthogonal translation of these codons. We demonstrated that T systems provide freedom of choice regarding stop (selector) codons. In fact, opal suppression was found to be the best performing system, with 11 Ochre suppression still showed a 5-fold increase in selectivity with 20% efficiency. Ta.

[0316] Example 8 - Orthogonal translation of proteins for different cellular compartments

[0317] OT goes beyond "simple" reporters K2::P1 The system (best in terms of selectivity and efficiency) To visualize the power of the amber suppression OT system, human nucleoporin 153 (N It was intended to show differential expression of Nup153) versus cytoskeletal vimentin. It is located in the nuclear pore complex and is more than 1500 amino acids long. This is approximately six times larger than the fluorescent protein reporter used in The previously described C-terminal GFP fusion (Nup153::EG FP 149TAG ) was used. This was only possible if amber suppression was successful. This produced characteristic nuclear envelope staining in the images (Nikic et al., ngew Chem 2016,55:16172-16276). up153::EGFP 149TAG , tagged at the mRNA level with MS2 tags (nup153::egfp 149TAG ::ms2), merged with mOrange Vimentin (cytoskeletal protein) containing an amber codon at position 116 (Vim 11 6TAG ::mOrange) from the same plasmid in HEK293T cells. Expression by the β-actin gene results in the production of both proteins in the presence of cytoplasmic PylRS. Each showed characteristic nuclear envelope and cytoskeleton staining. K2::P1 assembly With the use of GFP, only Nup153::GFP was visible (cotransfected H Confocal images of EK293T cells show selective nuclear rim staining. This effect was also observed in -7 cells. Swapping the MS tag for vimentin had no effect. Reverse it, and the result is Vim 116TAG Only ::mOrange was visible (CO This was observed in both S-7 and HEK293T cell experiments. K2::P1 The play We showed that these proteins act on specifically different mRNAs.

[0318] Example 9 - Orthogonal translation of a transmembrane protein

[0319] Transmembrane protein OT K2::P1 The assembly can be selectively expressed using It has also been shown that membrane protein expression represents another layer of translational complexity because During translation, ribosomes need to bind to the endoplasmic reticulum, where proteins are cotranslationally synthesized. In this experiment, mOrange and amber were inserted into the membrane at position 676. Fusion of insulin receptor 1 with codon (INSR 676TAG ::mOrange ), which is located at the plasma membrane and shows characteristic plasma membrane phenotypes in HEK293T cells, was used. It produces membrane staining (Nikic et al., Angew Chem 2014, 5 3:2245-2249). This construct was cloned using an MS2 tag in the 3'UTR. Tagged with Nup153::EGFP 149TAG With one dual cassette plus The construct was then transformed into a cytoplasmic PylRS system. OT K2::P1 In the presence of assembly by HEK293T cells either OT K2::P1In the presence of assembly, MS2-tagged proteins Selective expression and INSR 676TAG ::mOrange predicted plasma membrane localization and observed It has been suggested (data not shown) that the present invention involves a more complex membrane-associated translation process. It indicated the potential of the OT system.

[0320] Example 10 - Spatial distribution of elements of the OT system within cells

[0321] The spatial distribution of AFPs, particularly PylRS, within cells was assessed using immunofluorescence (IF). In addition, fluorescence in situ hybridization (FISH) was performed to identify tRNA Pyl Check In contrast to the two-color reporter used in the FFC experiments above, all Single-color NLS-GFP fused to MS2 tag in IF / FISH experiments 39TAG R Porter (nls-gfp 39TAG ::ms2) to activate amber suppression We identified cells that were amber-suppressed (which give rise to green nuclei when amber suppression is successful). (The authors helped optimize the color channels.) IF and FISH staining revealed cytoplasmic Pyl In contrast to RS, the P1 system forms small intracellular assembler::PylRS droplets. (Data not shown) This indicated the occurrence of phase separation. Pyl colocalized well with highly dispersed assembler::PylRS droplets, which suggested that assembler: : This indicates that the PylRS phase can be partitioned well into the PylRS phase. It showed further colocalization with cembla::MCP (data not shown). The P2 system showed larger but still multiple dispersed droplet-like structures (data not shown). A combination of both assembler strategies (K1::P1, K2::P1, K1::P2, K2 ::P2) induces the formation of large, micron-sized organelle-like structures in the cytoplasm. These structures were observed in most cases at a few or even one per cell. The combined assemblers were localized at the position of mRNA::ms2, tRN A Pyl , Assembler::PylRS, and Assembler::MCP all have organelle-like structures. The combination of two assembler strategies, i.e., spatial localization by kinesin shortening, Targeted and paired phase separation provides the best containment as determined by FISH and IF This resulted in the highest selectivity increase for tRNA Pyl , PylRS, and mR Higher spatial separation of NAs and therefore higher local concentration correlates with higher selectivity. This is consistent with the hypothesis.

[0322] Ribosomes were stained to show that they K2::P1 Colocalization in assemblies IF staining of the ribosomal protein RPL26L1 was observed in OT K2::P1 Olga revealed strong colocalization with Nera (data not shown) and, inconclusively, with translating mRNA We demonstrated ribosome recruitment due to binding to ms2. High ribosome mobility was also observed. Why is the membrane protein INSR (construct: INSR 676TAG ::mOrange:: This could explain why it was possible to successfully express ms2.

[0323] Without wishing to be bound by theory, the experimental results indicate that tRNA Pyl Concentration of by a set of ribosomes near or completely immersed in the pool of ribosomes. Selective orthogonal interactions in close proximity, and potentially even within, OT assemblies This strongly suggests that translation occurs in the tRNA Pyl That is, the assembler: OT due to its affinity for PylRS K2::P1 Mobilized in the assembly, Instead, they can be simultaneously distributed into the droplets and aminoacylated with their corresponding ncAAs. On the other hand, assembler::MCP recruits MS2-tagged mRNA. This attracts ribosomes and induces the double assembler system (K2::P1 = KIF16B::FU S::PylRS and KIF16B::EWSR1::MCP) co-partitioning into the nucleoids, which maintains access to other translation factors for translation to function tRNA Pyl Other ribosomes in the cytoplasm that are not exposed to Whenever they are encountered, they perform their standard function of terminating translation.

[0324] Example 11 - Further OT systems

[0325] In addition to the OT systems described in the previous examples, various other OT systems have been tested and These experiments have shown that the POI can be selectively orthogonal translated. A summary of the results is provided in Table 1 below. Unless otherwise indicated, the results are based on Nikic et al. gew Chem Int Ed Engl 2016,55(52):16172-1 6176) but with the corresponding AF, AA, or AAAF mutations The cytoplasmic NES-PylRS system carrying the nucleotide sequence 1441 was used as a non-specific reference (negative control). All experiments were performed using codon-specific tRNA Pyl and ncAA corresponding to PylRS mutants. was performed in the presence of

[0326] [Table 1-1]

[0327] [Table 1-2]

[0328] [Table 1-3]

[0329] [Table 1-4]

[0330] Example 12 - Further OT systems

[0331] In addition to the OT systems described in the previous examples, Various similar OT systems have been tested to demonstrate selective orthogonal translation of the reporter (i.e., POI). These experiments are summarized in Table 2 below. The results are shown in FIG. A, B, and C are shown in Nikic et al. (Angew Chem Int E d Engl 2016,55(52):16172-16176) The previously described cytoplasmic NES-PylRS system was used as a non-specific reference (negative control). All experiments were performed using codon-specific tRNA Pyl and ncAA corresponding to PylRS mutants The experiment was carried out in the presence of

[0332] Table 2: OT strains tested [Table 2] The results are shown in Figures 7A, B, and C.

[0333] Example 13 - Further OT fusion constructs tested

[0334] In addition to the OT system described in the previous examples, a variety of other OT fusion constructs have been prepared and tested. and found to allow for selective orthogonal translation of the reporter (i.e., POI). A summary of the constructs tested is provided in Table 3 below. et al.(Angew Chem Int Ed Engl 2016,55(52 ):16172-16176), but with the corresponding AF, A A or AAAF mutations, or Pyl RS mutants CpkRS, CbzRS, IFRS 1, and the cytoplasmic NES-PylRS system with OMeRS one was used as a nonspecific reference (negative was used as a control.

[0335] All experiments were performed using codon-specific tRNA Pyl and non-canonical addresses corresponding to PylRS mutants. The reaction was carried out in the presence of amino acids (e.g., cyclopropene-L-lysine and CpkRS, N(I) CbzRS, 3-iodo-L-phenyl phenylalanine and IFRS-1, 4-methoxy-L-phenylalanine and OMeRS).

[0336] All constructs were constructed with the ms2 loop, λ N22 Second For PCP, the BoxB loop was tested, and for PCP, the pp7 loop was tested.

[0337] In all fusion constructs, the synthetases should be freely interchangeable.

[0338] In the SYNZIP construct, SYNZIP1 is paired with SYNZIP2, and SYNZ It is important to note that IP3 pairs with SYNZIP4. As such, all other listed SYNZIPs should function similarly (https http: / / pubs.acs.org / doi / pdf / 10.1021 / ja907617 a).

[0339] [Table 3-1]

[0340] [Table 3-2]

[0341] [Table 3-3]

[0342] [Table 3-4]

[0343] Abbreviation

[0344] "-" or "::" symbol for peptidic linkage “·” Symbol for polypeptide combination AP Polypeptide segments acting as assemblers AFP assembler fusion protein BSA Bovine Serum Albumin BoxB λ N22 Specific binding site of lambda phage RNA stem loop P CbzRS Methanosarcina mazei PylRS(Y306M,L 309G,C348T) CDS Code Sequence CG1 CG1 (Nup42) nucleoporin protein for targeting to the nuclear envelope CMPSiaTr: A CMP sialic acid transporter for targeting to Golgi membranes. CpkRS Methanosarcina mazei PylRS(A302S EB1 Protein for targeting to microtubule plus ends Receptor-binding cancer antigen expressed on EBAG9 SiSo cells EBAG9 FL EBAG9 full-length protein for targeting to Golgi membranes EBAG9 1-29 EBAG9 amino acid residues 1-29 (N) for targeting to Golgi membranes terminal) EGFP 149TAG Amino acid position 149 is coded by an amber codon (TAG) Enhanced green fluorescent protein EP A polypeptide segment that acts as an effector ER endoplasmic reticulum EWSR1 Ewing sarcoma breakpoint region 1 (also referred to herein as EWS ) FBS Fetal Bovine Serum FFC Fluorescence Flow Cytometry FISH Fluorescence in situ hybridization FRB-CD28 Transmembrane proteins CD4, FRB (similar to mTOR), and CD28 Synthetic membrane targeting domain derived from FSC-A forward scatter region FUS sarcoma fusion FUS-CD28 is a synthetic membrane-targeted fusion polypeptide derived from CD4, FUS, and CD28. Chid GCE Genetic Code Extension GFP Green Fluorescent Protein GFP 39TAA Amino acid position 39 is encoded by the ochre codon (TAA). Green fluorescent protein GFP 39TAG Amino acid position 39 is encoded by an amber codon (TAG). Green fluorescent protein GFP 39TGA Amino acid position 39 is encoded by the opal codon (TGA) Green fluorescent protein GFP 39,149TAG The amino acid positions 39 and 149 each contain an amber codon ( Green fluorescent protein encoded by GFP 39,149,182TAG At amino acid positions 39, 149, and 182, respectively Green fluorescent protein encoded by amber codon (TAG) IC-TP Intracellular Targeting Polypeptide IDP Intrinsically Disordered Proteins IFRS1 Methanosarcina mazei PylRS(L305M,Y 306L, L309S, N346S, C348M) INSR Insulin receptor INS 676TAG Amino acid position 676 is encoded by an amber codon (TAG) Insulin receptor iRFP Near-infrared fluorescent protein KIF13A Kinesin family member 13A - not otherwise defined herein Unless otherwise specified, "KIF13A" refers specifically to the amide of KIF13A lacking P390. A fragment covering acid residues 1-411 (KIF13A 1-411,ΔP390 ) says. KIF16B Kinesin family member 16B - not otherwise defined herein Unless otherwise specified, "KIF16B" specifically refers to amino acid residues 1-400 of KIF16B. The fragment to be barred (KIF16B 1-400 ) λ N22 The 22-amino acid RNA binding of lambda phage antiterminator protein N domain Post-translational modifications for plasma membrane targeting of LcK lymphocyte-specific protein tyrosine kinase decorative part mCherry 185TAG The amino acid position 185 is amplified by the amber codon (TAG). Encoded mCherry MCP MS2 bacteriophage coat protein MLC Membraneless Compartment MS2 Enterobacteriaceae phage MS2 Two fused to the 3' untranslated region (or its coding sequence) of the MS2-tag mRNA MS2 RNA stem loop ms2 MS2 Tag ncAA non-standard amino acid NLS nuclear localization sequence Nup153 Nucleoporin 153 O-RS Orthogonal aminoacyl-tRNA synthetase OMeRS Methanosarcina mazei PyrRS(A302T,Y 384F, N346V, C348W, V401L) OT assembly can act as an artificial orthogonal translation (OT) organelle. Spatially Concentrated Components of the GCE Mechanism in a Membraneless Assembly that Can Be Used P450 2C1 1-27 P450 2C1 residues 1-27 (N terminal) PBS Phosphate Buffered Saline Bacteriophage coat protein for targeting to PCP pp7 loop tag PEI Polyethyleneimine POI Polypeptide of interest (= target polypeptide) POI TAG POI (or is its code sequence) pp7 loop tag from RNA bacteriophage pp7 PSP Phase-separating Polypeptide PylRS pyrrolysyl-tRNA synthetase PylRS AA Mutant M. mazei containing the amino acid substitutions N346A and C348A Lorisyl-tRNA synthetase PylRS AF Mutant M. mazei pi containing the amino acid substitutions Y306A and Y384F Lorisyl-tRNA synthetase PylRS AAAF Amino acid substitutions Y306A, N346A, C348A, and Y384 Mutant M. mazei pyrrolysyl-tRNA synthetase containing F RNA-TP RNA-targeting polypeptide RS aminoacyl-tRNA synthetase RT room temperature SCO Cyclooctyne Lysine SEM Standard error of the mean SSC Saline-Sodium Citrate (buffer solution) SSC-A side scatter area SSC-W side scatter width SPD5 Spindle-defective protein 5 SYNZIP1 Synthetic coiled-coil peptide 1 SYNZIP2 Synthetic Coiled Coil Peptide 2 SYNZIP3 Synthetic coiled-coil peptide 3 SYNZIP4 Synthetic Coiled-Coil Peptide 4 TOMM20 Translocase of the outer mitochondrial membrane 20 TOMM20 1-70 A fragment covering amino acid residues 1-70 of TOMM20 tRNA PylPyrrolidyl or another non-standard amino acid by wild-type or modified PylRS residues and site-specific incorporation of (non-standard) amino acid residues onto the POI For this purpose, a tRNA having an anticodon that is the reverse complement of the selector codon is preferably used. Inserted tRNA Pyl Which of these is the selector codon on the POI coding sequence? Depending on which tRNA was used, the stop codon amber Pyl,CUA ), Ochre (tRNA Pyl,UUA ), or Opal (tRNA Pyl,UCA ) against I had a codon. 3'UTR 3' untranslated region Vim 116TAG Amino acid position 116 is encoded by an amber codon (TAG). Vimentin

[0345] array

[0346] The following section describes the polypeptide and polynucleotide sequences described herein. show.

[0347] Nucleic acid sequences are written in the 5' to 3' orientation, protein sequences are written from the N to C terminus.

[0348] Array-Set 1 1. Hybridization probes tRNA labeled with Cy5 at the 5' end Pyl Hybridization probes CTAACCCGGCTGAACGGATTTAGAGTCCATTCGATC (sequence number No. 1) tRNA labeled with digoxigenin at the 5' end Pyl Hybridization probe CTAACCCGGCTGAACGGATTTAGAGTCCATTCGATC (sequence number No. 2) MS2 RNA stem fragment labeled with Alexa Fluor 647 at the 5' end Hybridization probe for loop sequence CTGCAGACATGGGTGATCCTCATGTTTTCTA (SEQ ID NO: 3)

[0349] 2. tRNA tRNA Pyl,CUA DNA sequence of Methanosarcina mazei pyrrolysyl-tRNA; anticodon is underlined) GGAAACCTGATCATGTAGATCGAATGGACT CTA AATCCGT TCAGCCGGGTTAGATTCCCGGGGTTTCCG (SEQ ID NO: 4) tRNA Pyl,UCA DNA sequence of Methanosarcina mazei pyrrolysyl-tRNA; anticodon is underlined) GGAAACCTGATCATGTAGATCGAATGGACT TCA AATCCGT TCAGCCGGGTTAGATTCCCGGGGTTTCCG (SEQ ID NO:5) tRNA Pyl,UUA DNA sequence of Methanosarcina mazei pyrrolysyl-tRNA; anticodon is underlined) GGAAACCTGATCATGTAGATCGAATGGACT TTA AATCCGT TCAGCCGGGTTAGATTCCCGGGGTTTCCG (SEQ ID NO: 6)

[0350] 3. O-RS PylRS AF(Methanosarcina mazei Pyrrolidyl tRNA Synthetase Double Mutant: Y306A, Y384F; Uniprot: Q8PWY1)

[0351] DNA: ATGGCGTGCCCGGTGCCGCTGCAGCTGCCGCCGCTGGAAC GCCTGACCCTGGATGATAAAAAACCGCTGAATACCCTGAT CTCTGCTACTGGTCTGTGGATGAGTCGTACCGGAACCATT CATAAAATCAAACACCACGAGGTTAGCCGTTCGAAAATCT ATATTGAGATGGCGTGTGGCGATCATCTGGTTGTGAACAA TAGCCGCTCTTCTCGTACAGCACGTGCACTGCGTCACCAC AAATATCGTAAAACCTGTAAACGTTGCCGTGTGTCCGATG AGGATCTGAACAAATTCCTGACAAAAGCCAATGAGGACCA AACAAGCGTGAAAGTGAAAGTCGTTAGCGCTCCTACCCGT ACTAAAAAAGCAATGCCGAAATCCGTTGCTCGTGCCCCTA AACCACTGGAAAACACTGAAGCAGCACAGGCACAGCCGTC TGGAAGCAAATTCTCTCCGGCCATTCCTGTTTCTACCCAG GAGTCCGTTTCTGTTCCAGCAAGTGTGAGCACCAGCATTA GCAGTATTAGCACCGGTGCCACCGCTAGCGCCCTGGTTAA AGGCAATACCAATCCGATTACAAGCATGTCTGCCCCGGTT CAAGCATCAGCTCCAGCACTGACAAAATCCCAAACCGATC GTCTGGAGGTTCTGCTGAATCCGAAAGACGAAATCAGCCT GAATTCCGGCAAACCGTTTCGTGAACTGGAGAGCGAACTG CTGTCACGTCGTAAAAAAGACCTGCAACAAATCTATGCCG AAGAACGTGAGAACTATCTGGGGAAACTGGAACGTGAAAT CACCCGCTTTTTCGTGGATCGTGGCTTTCTGGAGATCAAA TCCCCGATTCTGATTCCTCTGGAGTATATCGAGCGTATGG GCATCGACAATGATACCGAACTGAGCAAACAAATTTTCCG TGTGGATAAAAACTTCTGTCTGCGCCCTATGCTAGCACCA AATCTGGCTAACTATCTGCGCAAACTGGACCGTGCCCTGC CTGATCCTATCAAAATCTTCGAGATCGGCCCGTGTTATCG TAAAGAGTCCGACGGTAAAGAACATCTGGAGGAGTTTACC ATGCTGAACTTTTGCCAAATGGGTTCAGGTTGTACTCGTG AGAACCTGGAAAGCATCATCACCGATTTTCTGAACCACCT GGGCATTGACTTCAAAATTGTGGGCGACAGCTGTATGGTG TTTGGCGACACCCTGGATGTCATGCACGGCGACCTGGAAC TGTCTAGTGCCGTTGTGGGCCCAATCCCGCTGGATCGTGA GTGGGGTATCGACAAACCTTGGATCGGTGCGGGTTTTGGT CTGGAGCGTCTGCTGAAAGTAAAACACGACTTCAAGAACA TCAAACGTGCTGCACGTTCCGAGTCCTATTACAATGGTAT TTCTACTAACCTGTAA (SEQ ID NO: 7)

[0352] protein: MACPVPLQLPPLERLTLDDKKPLNTLISATGLWMSRTGTI HKIKHHEVSRSKIYIEMACGDHLVVNNSRSSRTARALRHH KYRKTCKRCRVSDEDLNKFLTKANEDQTSVKVKVVSAPTR TKKAMPKSVARAPKPLENTEAAQAQPSGSKFSPAIPVSTQ ESVSVPASVSTSISSISTGATASALVKGNTNPITSMSAPV QASAPALTKSQTDRLEVLLNPKDEISLNSGKPFRELESEL LSRRKKDLQQIYAEERENYLGKLEREITRFVDRGFLEIK SPILIPLEYIERMGIDNDTELSKQIFRVDKNFCLRPMLAP NLANYLRKLDRALPPDIKIFEIGPCYRKESDGKEHLEEFT MLNFCQMGSGCTRENLESIITDFLNHLGIDFKIVGDSCMV FGDTLDVMHGDLELSSAVVGPIPLDREWGIDKPWIGAGFG LERLLKVKHDFKNIKRAARSESYYNGISTNL (SEQ ID NO: 8) PylRS AA (Methanosarcina mazei pyrrolysyl-tRNA synthase tase double mutant: N346A, C348A; Uniprot: Q8PWY1)

[0353] DNA: ATGGCGTGCCCGGTGCCGCTGCAGCTGCCGCCGCTGGAAC GCCTGACCCTGGATGACAAAAAACCGCTGAATACCTGAT CTCTGCTACTGGTCTGTGGATGAGTCGTACCGGAACCATT CATAAAATCAAACACCACGAGGTTAGCCGTTCGAAAATCT ATATTGAGATGGCGTGTGGCGATCATCTGGTTGTGAACAA TAGCCGCTCTTCTCGTACAGCACGTGCACTGCGTCACCAC AAATATCGTAAAACCTGTAAACGTTGCCGTGTGTCCGATG AGGATCTGAACAAATTCCTGACAAAAGCCAATGAGGACCA AACAAGCGTGAAAGTGAAAGTCGTTAGCGCTCCTACCCGT ACTAAAAAAGCAATGCCGAAATCCGTTGCTCGTGCCCCTA AACCACTGGAAAACACTGAAGCAGCACAGGCACAGCCGTC TGGAAGCAAATTCTCTCCGGCCATTCCTGTTTCTACCCAG GAGTCCGTTTCTGTTCCAGCAAGTGTGAGCACCAGCATTA GCAGTATTAGCACCGGTGCCACCGCTAGCGCCCTGGTTAA AGGCAATACCAATCCGATTACAAGCATGTCTGCCCCGGTT CAAGCATCAGCTCCAGCACTGACAAAATCCCAAACCGATC GTCTGGAGGTTCTGCTGAATCCGAAAGACGAAATCAGCCT GAATTCCGGCAAACCGTTTCGTGAACTGGAGAGCGAACTG CTGTCACGTCGTAAAAAAGACCTGCAACAAATCTATGCCG AAGAACGTGAGAACTATCTGGGGAAACTGGAACGTGAAAT CACCCGCTTTTTCGTGGATCGTGGCTTTCTGGAGATCAAA TCCCCGATTCTGATTCCTCTGGAGTATATCGAGCGTATGG GCATCGACAATGATACCGAACTGAGCAAACAAATTTTCCG TGTGGATAAAAACTTCTGTCTGCGCCCTATGCTGGCACCA AATCTGTATAACTATCTGCGCAAACTGGACCGTGCCCTGC CTGATCCTATCAAAATCTTCGAGATCGGCCCGTGTTATCG TAAAGAGTCCGACGGTAAAGAACATCTGGAGGAGTTTACC ATGCTGGCCTTTGCCCAAATGGGTTCAGGTTGTACTCGTG AGAACCTGGAAAGCATCATCACCGATTTTCTGAACCACCT GGGCATTGACTTCAAAATTGTGGGCGACAGCTGTATGGTG TATGGCGACACCCTGGATGTCATGCACGGCGACCTGGAAC TGTCTAGTGCCGTTGTTGGACCAATTCCGCTGGACCGTGA GTGGGGTATCGACAAACCGTGGATCGGAGCAGGATTCGGT CTGGAACGCCTGCTGAAAGTGAAACACGACTTCAAAAACA TCAAACGTGCCGCCCGTTCTGAATCGTATTATAACGGGAT CTCTACGAACCTGTAA(SEQ ID NO:9)

[0354] protein: MACPVPLQLPPLERLTLDDKKPLNTLISATGLWMSRTGTI HKIKHHEVSRSKIYIEMACGDHLVVNNSRSSRTARALRHH KYRKTCKRCRVSDEDLNKFLTKANEDQTSVKVKVVSAPTR TKKAMPKSVARAPKPLENTEAAQAQPSGSKFSPAIPVSTQ ESVSVPASVSTSISSISTGATASALVKGNTNPITSMSAPV QASAPALTKSQTDRLEVLLNPKDEISLNSGKPFRELESEL LSRRKKDLQQIYAEERENYLGKLEREITRFFVDRGFLEIK SPILIPLEYIERMGIDNDTELSKQIFRVDKNFCLRPMLAP NLYNYLRKLDRALPDPIKIFEIGPCYRKESDGKEHLEEFT MLAFAQMGSGCTRENLESIITDFLNHLGIDFKIVGDSCMV YGDTLDVMHGDLELSSAVVGPIPLDREWGIDKPWIGAGFG LERLLKVKHDFKNIKRAARSESYYNGISTNL(SEQ ID NO:10) PylRS AAAF (Methanosarcina mazei pyrrolidyl tRNA synthetase quadruple mutant: Y306A, N346A, C348A, Y384F; Uniprot: Q8PWY1) ンセターゼ四重変異体:Y306A,N346A,C348A,Y384F; Unip rot:Q8PWY1)

[0355] DNA: GCGTGCCCGGTGCCGCTGCAGCTGCCGCCGCTGGAACGCC TGACCCTGGATGATAAAAAACCGCTGAATACCCTGATCTC TGCTACTGGTCTGTGGATGAGTCGTACCGGAACCATTCAT AAAATCAAACACCACGAGGTTAGCCGTTCGAAAATCTATA TTGAGATGGCGTGTGGCGATCATCTGGTTGTGAACAATAG CCGCTCTTCTCGTACAGCACGTGCACTGCGTCACCACAAA TATCGTAAAACCTGTAAACGTTGCCGTGTGTCCGATGAGG ATCTGAACAAATTCCTGACAAAAGCCAATGAGGACCAAAC AAGCGTGAAAGTGAAAGTCGTTAGCGCTCCTACCCGTACT AAAAAAGCAATGCCGAAATCCGTTGCTCGTGCCCCTAAAC CACTGGAAAACACTGAAGCAGCACAGGCACAGCCGTCTGG AAGCAAATTCTCTCCGGCCATTCCTGTTTCTACCCAGGAG TCCGTTTCTGTTCCAGCAAGTGTGAGCACCAGCATTAGCA GTATTAGCACCGGTGCCACCGCTAGCGCCCTGGTTAAAGG CAATACCAATCCGATTACAAGCATGTCTGCCCCGGTTCAA GCATCAGCTCCAGCACTGACAAAATCCCAAACCGATCGTC TGGAGGTTCTGCTGAATCCGAAAGACGAAATCAGCCTGAA TTCCGGCAAACCGTTTCGTGAACTGGAGAGCGAACTGCTG TCACGTCGTAAAAAAGACCTGCAACAAATCTATGCCGAAG AACGTGAGAACTATCTGGGGAAACTGGAACGTGAAATCAC CCGCTTTTTCGTGGATCGTGGCTTTCTGGAGATCAAATCC CCGATTCTGATTCCTCTGGAGTATATCGAGCGTATGGGCA TCGACAATGATACCGAACTGAGCAAACAAATTTTCCGTGT GGATAAAAACTTCTGTCTGCGCCCTATGCTAGCACCAAAT CTGGCTAACTATCTGCGCAAACTGGACCGTGCCCTGCCTG ATCCTATCAAAATCTTCGAGATCGGCCCGTGTTATCGTAA AGAGTCCGACGGTAAAGAACATCTGGAGGAGTTTACCATG CTGGCCTTTGCCCAAATGGGTTCAGGTTGTACTCGTGAGA ACCTGGAAAGCATCATCACCGATTTTCTGAACCACCTGGG CATTGACTTCAAAATTGTGGGCGACAGCTGTATGGTGTTT GGCGACACCCTGGATGTCATGCACGGCGACCTGGAACTGT CTAGTGCCGTTGTGGGCCCAATCCCGCTGGATCGTGAGTG GGGTATCGACAAACCTTGGATCGGTGCGGGTTTTGGTCTG GAGCGTCTGCTGAAAGTAAAACACGACTTCAAGAACATCA AACGTGCTGCACGTTCCGAGTCCTATTACAATGGTATTTC TACTAACCTGTAA(SEQ ID NO: 11)

[0356] protein: ACPVPLQLPPLERLTLDDKKPLNTLISATGLWMSRTGTIH KIKHHEVSRSKIYIEMACGDHLVVNNSRSSRTARALRHHK YRKTCKRCRVSDEDLNKFLTKANEDQTSVKVKVVSAPTRT KKAMPKSVARAPKPLENTEAAQAQPSGSKFSPAIPVSTQE SVSVPASVSTSISSISTGATASALVKGNTNPITSMSAPVQ ASAPALTKSQTDRLEVLLNPKDEISLNSGKPFRELESELL SRRKKDLQQIYAEERENYLGKLEREITRFVDRGFLEIKS PILIPLEYIERMGIDNDTELSKQIFRVDKNFCLRPMLAPN LANYLRKLDRALPPDIKIFEIGPCYRKESDGKEHLEEFTM LAFAQMGSGCTRENLESIITDFLNHLGIDFKIVGDSCMVF GDTLDVMHGDLELSSAVVGPIPLDREWGIDKPWIGAGFGL ERLLKVKHDFKNIKRAARSESYYNGISTNL (SEQ ID NO: 12)

[0357] 4. RNA-TP MCP (coat protein of enterobacteria phage MS2)

[0358] DNA: GCTTCTAACTTTACTCAGTTCGTTCTCGTCGACAATGGCG GAACTGGCGACGTGACTGTCGCCCCAAGCAACTTCGCTAA CGGGATCGCTGAATGGATCAGCTCTAACTCGCGTTCACAG GCTTACAAAGTAACCTGTAGCGTTCGTCAGAGCTCTGCGC AGAATCGCAAATACACCATCAAAGTCGAGGTGCCTAAAGG CGCCTGGCGTTCGTACTTAAATATGGAACTAACCATTCCA ATTTTCGCCACGAATTCCGACTGCGAGCTTATTGTTAAGG CAATGCAAGGTCTCCTAAAAGATGGAAACCCGATTCCCTC AGCAATCGCAGCAAACTCCGGCATCTAC (SEQ ID NO: 13)

[0359] protein: ASNFTQFVLVDNGGTGDVTVAPSNFANGIAEWISSNSRSQ AYKVTCSVRQSSAQNRKYTIKVEVPKGAWRSYLNMELTIP IFATNSDCELIVKAMQGLLKDGNPIPSAIAANSGIY (sequence number No. 14)

[0360] λ N22 (22 amino acid RNA binding of lambda phage antiterminator protein N) domain) DNA: ATGGACGCACAAACACGACGACGTGAGCGTCGCGCTGAGA AACAAGCTCAATGGAAAGCTGCAAAC (SEQ ID NO: 15) protein: MDAQTRRRERRAEKQAQWKAAN (SEQ ID NO: 16)

[0361] 5. TN DNA sequence of the enterobacterial phage MS2 RNA stem loop. ACATGAGGATCACCCATGT (SEQ ID NO: 17) The DNA sequence of BoxB (lambda phage RNA stem loop, λ N22 of specific binding site) GCCCTGAAAAAGGGC (SEQ ID NO: 18)

[0362] 6. IC-TP KIF16B 1-400 (Homo sapien covering amino acid residues 1-400) sKinesin family member 16B fragment; Uniprot:Q96L93)

[0363] DNA: ATGGCATCGGTCAAGGTGGCCGTGAGGGTCCGGCCCATGA ATCGCAGGGAAAAGGACTTGGAGGCCAAGTTCATTATTCA GATGGAGAAAAGCAAAACGACAATCACAAACTTAAAGATA CCAGAAGGAGGCACTGGGGACTCAGGAAGAGAACGGACCA AGACCTTCACCTATGACTTTTCTTTTTATTCTGCTGATAC AAAAAGCCCAGATTACGTTTCACAAGAAATGGTTTTCAAA ACCCTCGGCACAGATGTCGTGAAGTCTGCATTTGAAGGTT ATAATGCTTGTGTCTTTGCATATGGGCAAACTGGATCTGG AAAGTCATACACTATGATGGGAAATTCTGGAGATTCTGGC TTAATACCTCGGATCTGTGAAGGACTCTTCAGTCGGATAA ATGAAACCACCAGATGGGATGAAGCTTCTTTTCGAACTGA AGTCAGCTACTTAGAAATTTATAACGAACGTGTGAGAGAT CTACTTCGGCGGAAGTCATCTAAAACCTTCAATTTGAGAG TCCGTGAGCATCCCAAAGAAGGCCCTTATGTTGAGGATTT ATCCAAACATTTAGTACAGAATTATGGTGACGTAGAAGAA CTTATGGATGCGGGCAATATCAACCGGACCACCGCAGCGA CTGGGATGAACGACGTCAGTAGCAGGTCTCATGCCATCTT CACCATCAAGTTCACTCAGGCTAAATTTGATTCTGAAATG CCATGTGAAACCGTCAGTAAGATCCACTTGGTTGATCTTG CCGGAAGTGAGCGTGCAGATGCCACCGGAGCCACCGGGGT TAGGCTAAAGGAAGGGGGAAATATTAACAAGTCCCTCGTG ACTCTGGGGAACGTCATTTCTGCCTTAGCTGATTTATCTC AGGATGCTGCAAATACTCTTGCAAAGAAGAAGCAAGTTTT CGTGCCTTACAGGGATTCTGTGTTGACTTGGTTGTTAAAA GATAGCCTTGGAGGAAACTCTAAAACTATCATGATTGCCA CCATTTCACCTGCTGATGTCAATTATGGAGAAACCCTAAG TACTCTTCGCTATGCAAATAGAGCCAAAAACATCATCAAC AAGCCTACCATTAATGAGGATGCCAACGTCAAACTTATCC GTGAGCTGCGAGCTGAAATAGCCAGACTGAAAACGCTGCT TGCTCAAGGGAATCAGATTGCCCTCTTAGACTCCCCCACA (SEQ ID NO: 19)

[0364] protein: MASVKVAVRVRPMNRREKDLEAKFIIQMEKSKTTITNLKI PEGGTGDSGRERTKTFTYDFSFYSADTKSPDYVSQEMVFK TLGTDVVKSAFEGYNACVFAYGQTGSGKSYTMMGNSGDSG LIPRICEGLFSRINETTRWDEASFRTEVSYLEIYNERVRD LLRRKSSKTFNLRVREHPKEGPYVEDLSKHLVQNYGDVEE LMDAGNINRTTAATGMNDVSSRSHAIFTIKFTQAKFDSEM PCETVSKIHLVDLAGSERADATGATGVRLKEGGNINKSLV TLGNVISALADLSQDAANTLAKKKQVFVPYRDSVLTWLLK DSLGGNSKTIMIATISPADVNYGETLSTLRYANRAKNIIN KPTINEDANVKLIRELRAEIARLKTLLAQGNQIALLDSPT (SEQ ID NO:20)

[0365] KIF13A 1-411,ΔP390 (P390 is missing amino acid residues 1-411 Homo sapiens kinesin family member 13A fragment covering ; Uniprot:Q9H1H9)

[0366] DNA: ATGTCGGATACCAAGGTAAAAGTTGCCGTCCGGGTCCGGC CCATGAACCGACGAGAACTGGAACTGAACACCAGTGCGT GGTGGAGATGGAAGGGAATCAAACGGTCCTGCACCCTCCT CCTTCTAACACCAAACAGGGAGAAAGGAAACCTCCCAAGG TATTTGCCTTTGATTATTGCTTTTGGTCCATGGATGAATC TAACACTACAAAATACGCTGGTCAAGAAGTGGTTTCAAG TGCCTTGGGGAAGGAATTCTTGAAAAAGCCTTTCAGGGGT ATAATGCGTGTATTTTTGCATATGGACAGACAGGTTCGGG AAAATCCTTTTCCATGATGGGCCATGCTGAGCAGCTGGGC CTTATTCCAAGGCTCTGCTGTGCTTTATTTAAAAGGATCT CTTTGGAGCAAAATGAGTCACAGACCTTTAAAGTTGAAGT GTCCTATATGGAAATTTATAATGAGAAAGTTCGGGATCTT TTAGACCCCAAAGGGAGTAGACAGTCTCTTAAAGTTCGAG AACATAAAGTTTTGGGACCATATGTAGATGGTTTATCTCA ACTAGCTGTCACTAGTTTTGAGGATATTGAGTCATTGATG TCTGAGGGAAATAAGTCTCGAACGGTAGCTGCTACCAACA TGAACGAAGAAAGCAGCCGCTCCCATGCTGTGTTCAACAT CATAATCACACAGACACTTTATGACCTGCAGTCTGGGAAT TCCGGGGAGAAAGTCAGTAAGGTCAGCTTGGTAGACCTGG CGGGTAGCGAAAGAGTATCTAAAACAGGAGCTGCAGGAGA GCGACTGAAAGAAGGCAGCAACATTAACAAATCGCTTACA ACCTTGGGGTTGGTTATATCATCACTGGCTGACCAGGCAG CTGGCAAGGGTAAAAGCAAATTTGTGCCTTATCGAGATTC AGTCCTCACTTGGCTGCTTAAGGACAACTTGGGGGGCAAC AGCCAAACCTCTATGATAGCCACAATCAGCCCAGCCGCAG ACAACTATGAAGAGACCCTCTCCACATTAAGATATGCAGA CCGAGCCAAAAGGATTGTGAACCATGCTGTTGTGAATGAG GACCCCAACGCAAAAGTGATCCGAGAACTGCGGGAGGAAG TCGAGAAACTGAGAGAGCAGCTCTCTCAGGCAGAGGCCAT GAAGGCCGAACTGAAGGAGAAGCTCGAAGAGTCTGAAAAG CTGATAAAAGAACTAACAGTGACTTGGGAA (SEQ ID NO: 21)

[0367] protein: MSDTKVKVAVRVRPMNRRELELNTKCVVEMEGNQTVLHPP PSNTKQGERKPPKVFAFDYCFWSMDESNTTKYAGQEVVFK CLGEGILEKAFQGYNACIFAYGQTGSGKSFSMMGHAEQLG LIPRLCCALFKRISLEQNESQTFKVEVSYMEIYNEKVRDL LDPKGSRQSLKVREHKVLGPYVDGLSQLAVTSFEDIESLM SEGNKSRTVAATNMNEESSRSHAVFNIIITQTLYDLQSGN SGEKVSKVSLVDLAGSERVSKTGAAGERLKEGSNINKSLT TLGLVISSLADQAAGKGKSKFVPYRDSVLTWLLKDNLGGN SQTSMIATISPAADNYEETLSTLRYADRAKRIVNHAVVNE DPNAKVIRELREEVEKLREQLSQAEAMKAELKEKLEESEK LIKELTVTWE (SEQ ID NO:22)

[0368] TOMM20 1-70 (Homo sapiens M1 covering amino acid residues 1-70) Tochondriac outer membrane translocase 20 fragment; Uniprot:Q15388 ) DNA: ATGGTGGGTCGGAACAGCGCCATCGCCGCCGGTGTTATGCG GGGCCCTTTTCATTGGGTACTGCATCTACTTCGACCGCAA AAGACGAAGTGACCCCAACTTCAAGAACAGGCTTCGAGAA CGAAGAAAGAAACAGAAGCTTGCCAAGGAGAGAGCTGGGC TTTCCAAGTTACCTGACCTTAAAGATGCTGAAGCTGTTCA GAAATTCTTC (SEQ ID NO: 23) protein: MVGRNSAIAGVCGALFIGYCIYFDRKRRSDPNFKNRLRE RRKKQKLAKERAGLSKLPDLKDAEAVQKFF (SEQ ID NO: 24)

[0369] Plasma expression of LcK (Mus musculus lymphocyte-specific protein tyrosine kinase) Post-translational modification site for membrane targeting; Uniprot:P06240) DNA: GGCTGCGTGTGCAGCAGCAACCCCGAGGGTACCGAGCTC( SEQ ID NO:25) protein: (The same part is underlined P06240) GCVCSSNPE GTEL (SEQ ID NO:26)

[0370] FRB-CD28(Mus musculus CD4(Uniprot:P06332 ), FRB (similar to Homo sapiens mTOR; Uniprot:P423 45), and Mus musculus CD28 (Uniprot:P31041). Derived membrane-targeted synthetic fusion polypeptides

[0371] DNA: ATGTGCCGAGCCATCTCTCTTAGGCGCTTGCTGCTGCTGC TGCTGCAGCTGTCACAACTCCTAGCTGTCACTCAAGGGAT GCTCGAGATGTGGCATGAAGGCCTGGAAGAGGCATCTCGT TTGTACTTTGGGGAAAGGAACGTGAAAGGCATGTTTGAGG TGCTGGAGCCCTTGCATGCTATGATGGAACGGGGCCCCCA GACTCTGAAGGAAACATCCTTTAATCAGGCCTATGGTCGA GATTTAATGGAGGCCCAAGAGTGGTGCAGGAAGTACATGA AATCAGGGAATGTCAAGGACCTCCTCCAAGCCTGGGACCT CTATTATCATGTGTTCCGACGAATCTCAAAGACTAGAACC GGTAAGCTTTTTTGGGCACTGGTCGTGGTTGCTGGAGTCC TGTTTTGTTATGGCTTGCTAGTGACAGTGGCTCTTTGTGT T(SEQ ID NO:27)

[0372] protein: MCRAISLRRLLLLLLQLSQLLAVTQGMLEMWHEGLEEASR LYFGERNVKGMFEVLEPLHAMMERGPQTLKETSFNQAYGR DLMEAQEWCRKYMKSGNVKDLLQAWDLYYHVFRRISKTRT GKLFWALVVVAGVLFCYGLLVTVALCV(SEQ ID NO:28)

[0373] FUS-CD28 (Mus musculus CD4 (Uniprot: P06332 ), Homo sapiens sarcoma fusion (Uniprot: P35637), and Mus musculus CD28 (Uniprot: P31041)-derived membrane-targeted chimeric fusion polypeptide

[0374] DNA: ATGTGCCGAGCCATCTCTCTTAGGCGCTTGCTGCTGCTGC TGCTGCAGCTGTCACAACTCCTAGCTGTCACTCAAGGGAT GCTCATGGCCTCAAACGATTATACCCAACAAGCAACCCAA AGCTATGGGGCCTACCCCACCCAGCCCGGGCAGGGCTATT CCCAGCAGAGCAGTCAGCCCTACGGACAGCAGAGTTACAG TGGTTATAGCCAGTCCACGGACACTTCAGGATATGGCCAG AGCAGCTATTCTTCTTATGGCCAGAGCCAGAACACAGGCT ATGGAACTCAGTCAACTCCCCAGGGATATGGCTCGACTGG CGGCTATGGCAGTAGCCAGAGCTCCCAATCGTCTTACGGG CAGCAGTCCTCCTACCCTGGCTATGGCCAGCAGCCAGCTC CCAGCAGCACCTCGGGAAGTTACGGTAGCAGTTCTCAGAG CAGCAGCTATGGGCAGCCCCAGAGTGGGAGCTACAGCCAG CAGCCTAGCTATGGTGGACAGCAGCAAAGCTATGGACAGC AGCAAAGCTATAATCCCCCTCAGGGCTATGGACAGCAGAA CCAGTACAACAGCAGCAGTGGTGGTGGAGGTGGAGGTGGA GGTGGAGGTAACTATGGCCAAGATCAATCCTCCATGAGTA GTGGTGGTGGCAGTGGTGGCGGTTATGGCAATCAAGACCA GAGTGGTGGAGGTGGCAGCGGTGGCTATGGACAGCAGGAC CGTGGAGGCCGCGGCAGGGGTGGCAGTGGTGGCGGCGGCG GCGGCGGCGGTGGTGGTTACAACCGCAGCAGTGGTGGCTA TGAACCCAGAGGTCGTGGAGGTGGCCGTGGAGGCAGAGGT GGCATGGGCGGAAGTGACCGTGGTGGCTTCAATAAATTTG GTGGCCCTCGGGACCAAGGATCACGTCATGACTCCGAACA GGATAATTCAGACAACAACACCATCTTTGTGCAAGGCCTG GGTGAGAATGTTACAATTGAGTCTGTGGCTGATTACTTCA AGCAGATTGGTATTATTAAGACAAACAAGAAAACGGGACA GCCCATGATTAATTTGTACACAGACAGGGAAACTGGCAAG CTGAAGGGAGAGGCAACGGTCTCTTTTGATGACCCACCTT CAGCTAAAGCAGCTATTGACTGGTTTGATGGTAAAGAATT CTCCGGAAATCCTATCAAGGTCTCATTTGCTACTCGCCGG GCAGACTTTAATCGGGGTGGTGGCAATGGTCGTGGAGGCC GAGGGCGAGGAGGACCCATGGGCCGTGGAGGCTATGGAGG TGGTGGCAGTGGTGGTGGTGGCCGAGGAGGATTTCCCAGT GGAGGTGGTGGCGGTGGAGGACAGCAGCGAGCTGGTGACT GGAAGTGTCCTAATCCCACCTGTGAGAATATGAACTTCTC TTGGAGGAATGAATGCAACCAGTGTAAGGCCCCTAAACCA GATGGCCCAGGAGGGGGACCAGGTGGCTCTCACATGGGGG GTAACTACGGGGATGATCGTCGTGGTGGCAGAGGAGGCGG CACCGGTAAGCTTTTTTGGGCACTGGTCGTGGTTGCTGGA GTCCTGTTTTGTTATGGCTTGCTAGTGACAGTGGCTCTTT GTGTT (SEQ ID NO: 29)

[0375] protein: MCRAISLRRLLLLLLQLSQLLAVTQGMLMASNDYTQQATQ SYGAYPTQPGQGYSQQSSQPYGQQSYSGYSQSTDTSGYGQ SSYSSYGQSQNTGYGTQSTPQGYGSTGGYGSSQSSQSSYG QQSSYPGYGQQPAPSSTSGSYGSSSQSSSYGQPQSGSYSQ QPSYGGQQQSYGQQQSYNPPQGYGQQNQYNSSSGGGGGG GGGNYGQDQSSMSSGGGSGGGYGNQDQSGGGGSGGYGQQD RGGRGRGGSGGGGGGGGGGYNRSSGGYEPRGRGGGRGGRG GMGGSDRGGFNKFGGPRDQGSRHDSEQDNSDNNTIFVQGL GENVTIESVADYFKQIGIIKTNKKTGQPMINLYTDRETGK LKGEATVSFDDPPSAKAAIDWFDGKEFSGNPIKVSFATRR ADFNRGGGGNGRGGRGRGGPMGRGGYGGGGGSGGGGRGGFPS GGGGGGGQQRAGDWKCPNPTCENMNFSWRNECNQCKAPKP DGPGGGPGGSHMGGNYGDDRRGGRGGGTGKLFWALVVVAG VLFCYGLLVTVALCV (SEQ ID NO:30)

[0376] 7. PSP SPD5 (Caenorhabditis elegans spindle defect protein 5; Uniprot:P91349)

[0377] DNA: ATGGAGGACAACAGCGTGCTGAACGAGGACAGCAACCTGG AGCACGTGGAGGGCCAGCCCAGAAGAAGCATGAGCCAGCC CGTGCTGAACGTGGAGGGCGACAAGAGAACCAGCAGCACC AGCGCCACCCAGCAGCAGGTGCTGAGCGGCGCCTTCAGCA GCGCCGACGTGAGAAGCATCCCCATCATCCAGACCTGGGA GGAGAACAAGGCCCTGAAGACCAAGATCACCATCCTGAGA GGCGAGCTGCAGATGTACCAGAGAAGATACAGCGAGGCCA AGGAGGCCAGCCAGAAGAGAGTGAAGGAGGTGATGGACGA CTACGTGGACCTGAAGCTGGGCCAGGAGAACGTGCAGGAG AAGATGGAGCAGTACAAGCTGATGGAGGAGGACCTGCTGG CCATGCAGAGCAGAATCGAGACCAGCGAGGACAACTTCGC CAGACAGATGAAGGAGTTCGAGGCCCAGAAGCACGCCATG GAGGAGAGAATCAAGGAGCTGGAGCTGAGCGCCACCGACG CCAACAACACCACCGTGGGCAGCTTCAGAGGCACCCTGGA CGACATCCTGAAGAAGAACGACCCCGACTTCACCCTGACC AGCGGCTACGAGGAGAGAAAGATCAACGACCTGGAGGCCA AGCTGCTGAGCGAGATCGACAAGGTGGCCGAGCTGGAGGA CCACATCCAGCAGCTGAGACAGGAGCTGGACGACCAGAGC GCCAGACTGGCCGACAGCGAGAACGTGAGAGCCCAGCTGG AGGCCGCCACCGGCCAGGGCATCCTGGGCGCCGCCGGCAA CGCCATGGTGCCCAACAGCACCTTCATGATCGGCAACGGC AGAGAGAGCCAGACCAGAGACCAGCTGAACTACATCGACG ACCTGGAGACCAAGCTGGCCGACGCCAAGAAGGAGAACGA CAAGGCCAGACAGGCCCTGGTGGAGTACATGAACAAGTGC AGCAAGCTGGAGCACGAGATCAGAACCATGGTGAAGAACA GCACCTTCGACAGCAGCAGCATGCTGCTGGGCGGCCAGAC CAGCGACGAGCTGAAGATCCAGATCGGCAAGGTGAACGGC GAGCTGAACGTGCTGAGAGCCGAGAACAGAGAGCTGAGAA TCAGATGCGACCAGCTGACCGGCGGCGACGGCAACCTGAG CATCAGCCTGGGCCAGAGCAGACTGATGGCCGGCATCGCC ACCAACGACGTGGACAGCATCGGCCAGGGCAACGAGACCG GCGGCACCAGCATGAGAATCCTGCCCAGAGAGAGCCAGCT GGACGACCTGGAGGAGAGCAAGCTGCCCCTGATGGACACC AGCAGCGCCGTGAGAAACCAGCAGCAGTTCGCCAGCATGT GGGAGGACTTCGAGAGCGTGAAGGACAGCCTGCAGAACAA CCACAACGACACCCTGGAGGGCAGCTTCAACAGCAGCATG CCCCCCCCGGCAGAGACGCCACCCAGAGCTTCCTGAGCC AGAAGAGCTTCAAGAACAGCCCCCATCGTGATGCAGAGCC CAAGAGCCTGCACCTGCACCTGAAGAGCCACCAGAGCGAG GGCGCCGGCGAGCAGATCCAGAACAACAGCTTCAGCACCA AGACCGCCAGCCCCACGTGAGCCAGAGCCACATCCCCAT CCTGCACGACATGCAGCAGATCCTGGACAGCAGCGCCATG TTCCTGGAGGGCCAGCACGACGTGGCCGTGAACGTGGAGC AGATGCAGGAGAGATGAGCCAGATCAGAGAGGCCCTGGC CAGACTGTTCGAGAGACTGAAGAGCAGCGCGCCCTGTTC GAGGAGATCCTGGAGAGAATGGGCAGCAGCGACCCCAACG CCGACAAGATCAAGAAGATGAAGCTGGCCTTCGAGACCAG CATCAACGACAAGCTGAACGTGAGCGCCATCCTGGAGGCC GCCGAGAAGGACCTGCACAACATGAGCCTGAACTTCAGCA TCCTGGAGAAGAGCATCGTGAGCCAGGCCGCCGAGGCCAG CAGAAGATTCACCATCGCCCCGACGCCGAGGACGTGGCC AGCAGCAGCCTGCTGAACGCCAGCTACAGCCCCCTGTTCA AGTTCACCAGCAACAGCGACATCGTGGAGAAGCTGCAGAA CGAGGTGAGCGAGCTGAAGAACGAGCTGGAGATGGCCAGA ACCAGAGACATGAGAAGCCCCCTGAACGGCAGCAGCAGCGGCA GACTGAGCGACGTGCAGATCAAACACCAACAGAATGTTCGA GGACCTGGAGGTGAGCGAGGCCACCCTGCAGAAGGCCAAG GAGGAGAACAGCACCCTGAAGAGCCAGTTCGCCGAGCTGG AGGCCAACCTGCACCAGGTGAACAGCAAGCTGGGCGAGGT GAGATGCGAGCTGAACGAGGCCCTGGCCAGAGTGGACGGC GAGCAGGAGACCAGAGTGAAGGCCGAGAACGCCCTGGAGG AGGCCAGACAGCTGATCAGCAGCCTGAAGCACGAGGAAA CGAGCTGAAGAGACCATCACCGACATGGGCATGAGACTG AACGAGGCCAAGAAGAGCGACGAGTTCCTGAAGAGCGAGC TGAGCACCCGCCCTGGAGGAGGAGAAGAGAGCCCAGAACCT GGCCGACGAGCTGAGCGAGGAGCTGAACGGCTGGAGAATG AGAACCAAGGAGGCGAGAACAAGGTGGAGCACGCCAGCA GCGAGAAGAGCGAGATGCTGGAGAGAGAATCGTGCACCTGGA GACCGAGATGGAGAAGCTGAGCACCAGCGAGATCGCCGCC GACTACTGCAGCACCAAGATGACCGAGAGAAGAAGGAGA TCGAGCTGGCCAAGTACAGAGAGGACTTCGAGAACGCCGC CATCGTGGGCCTGGAGAGAATCAGCAAGGAGATCAGCGAG CTGACCAAGAAGACCCTGAAGGCCAAGATCATCCCCAGCA ACATCAGCAGCATCCAGCTGGTGTGCGACGAGCTGTGCAG AAGACTGAGCAGAGAGAGAGAGCAGCAGCACGAGTACGCC AAGGTGATGAGAGACGTGAACGAGAAGATCGAGAAGCTGC AGCTGGAGAAGGACGCCCTGGAGCACGAGCTGAAGATGAT GAGCAGCAACAACGAGAACGTGCCCCCCGTGGGCACCAGC GTGAGCGGCATGCCCACCAAGACCAGCAACCAGAAGTGCG CCCAGCCCCACTACACCAGCCCCACCAGACAGCTGCTGCA CGAGAGCACCATGGCCGTGGACGCCATCGTGCAGAAGCTG AAGAAGACCCACAACATGAGCGGCATGGGCCCCGAGCTGA AGGAGACCATCGGCAACGTGATCAACGAGAGCAGAGTGCT GAGAGACTTCCTGCACCAGAAGCTGATCCTGTTCAAGGGC ATCGACATGAGCAACTGGAAGAACGAGACCGTGGACCAGC TGATCACCGACCTGGGCCAGCTGCACCAGGACAACCTGAT GCTGGAGGAGCAGATCAAGAAGTACAAGAAGGAGCTGAAG CTGACCAAGAGCGCCATCCCCACCCTGGGCGTGGAGTTCC AGGACAGAATCAAGACCGAGATCGGCAAGATCGCCACCGA CATGGGCGGCGCCGTGAAGGAGATCAGAAAGAAG(SEQ ID NO: 3 1)

[0378] protein: MEDNSVLNEDSNLEHVEGQPRRSMSQPVNLVEGDKRTSST SATQQQVLSGAFSSADVRSIPIIQTWEENKALKTKITILR GELQMYQRRRYSEAKEASQKRVKEVMDDDYVDLKLGQENVQE KMEQYKLMEEDLLAMQSRIETSEDNFARQMKEFEAQKHAM EERIKEELSATDANNTTVGSFRGTLDDILKNDPDFTLT SGYEERKINDLEAKLLSEIDKVAELEDHIQQLRQELDDQS ARLADSENVRAQLEAATGQGILGAAGNAMVPNSTFMIGNG RESQTRDQLNYIDDLETKLADAKKENDKARQALVEYMNKC SKLEHEIRTMVKNSTFDSSSMLLGGQTSDELKIQIGKVNG ELNVLRAENRELRIRCDQLTGGDGNLSISLGQSRLMAGIA TNDVDSIGQGNETGGTSMRILPRESQLDDLEESKLPLMDT SSAVRNQQQFASMWEDFESVKDSLQNNNHNDTLEGSFNSSM PPPGRDATQSFLSQKSFKNSPIVMQKPKSLHLHLKSHQSE GAGEQIQNNSFSTKTASPHVSQSHIPILHDMQQILDSSAM FLEGQHDVAVNVEQMQEKMSQIREALARLFERLKSSAALF EEILERMGSSDPNADKIKKMKLAFETSINDKLNVSAILEA AEKDLHNMSLNFSILEKSIVSQAAEASRRFTIAPDAEDVA SSSLLNASYSPLFKFTSNDIVEKLQNEVSELKNELEMAR TRDMRSPLNGSSGRLSDVQINTNRMFEDLEVSEATLQKAK EENSTLKSQFAELEANLHQVNSKLGEVRCELNEALARVDG EQETRVKAENALEEARQLISSLKHEENELKKTITDMGMRL NEAKKSDEFLKSELSTALEEEKKSQNLADELSEELNGWRM RTKEAENKVEHASSEKSEMLERIVHLETEMEKLSTSEIAA DYCSTKMTERKKEIELAKYREDFENAAIVGLERISKEISE LTKKTLKAKIIPSNISSIQLVCDELCRRLSREREQQHEYA KVMRDVNEKIEKLQLEKDALEHELKMMSSNNENVPPVGTS VSGMPTKTSNQKCAQPHYTSPTRQLLHESTMAVDAIVQKL KKTHNMSGMGPELKETIGNVINESRVLRDFLHQKLILFKG IDMSNWKNETVDQLITDLGQLHQDNLMLEEQIKKYKKELK LTKSAIPTLGVEFQDRIKTEIGKIATDMGGAVKEIRKK(sequence number 32) 列番号32

[0379] FUS (Homo sapiens sarcoma fusion; Uniprot:P35637)

[0380] DNA: ATGGCCTCAAACGATTATACCCAACAAGCAACCCAAAGCT ATGGGGCCTACCCCACCCAGCCCGGGCAGGGCTATTCCCA GCAGAGCAGTCAGCCCTACGGACAGCAGAGTTACAGTGGT TATAGCCAGTCCACGGACACTTCAGGATATGGCCAGAGCA GCTATTCTTCTTATGGCCAGAGCCAGAACACAGGCTATGG AACTCAGTCAACTCCCCAGGGATATGGCTCGACTGGCGGC TATGGCAGTAGCCAGAGCTCCCAATCGTCTTACGGGCAGC AGTCCTCCTACCCTGGCTATGGCCAGCAGCCAGCTCCCAG CAGCACCTCGGGAAGTTACGGTAGCAGTTCTCAGAGCAGC AGCTATGGGCAGCCCCAGAGTGGGAGCTACAGCCAGCAGC CTAGCTATGGTGGACAGCAGCAAAGCTATGGACAGCAGCA AAGCTATAATCCCCCTCAGGGCTATGGACAGCAGAACCAG TACAACAGCAGCAGTGGTGGTGGAGGTGGAGGTGGAGGTG GAGGTAACTATGGCCAAGATCAATCCTCCATGAGTAGTGG TGGTGGCAGTGGTGGCGGTTATGGCAATCAAGACCAGAGT GGTGGAGGTGGCAGCGGTGGCTATGGACAGCAGGACCGTG GAGGCCGCGGCAGGGGTGGCAGTGGTGGCGGCGGCGGCGG CGGCGGTGGTGGTTACAACCGCAGCAGTGGTGGCTATGAA CCCAGAGGTCGTGGAGGTGGCCGTGGAGGCAGAGGTGGCA TGGGCGGAAGTGACCGTGGTGGCTTCAATAAATTTGGTGG CCCTCGGGACCAAGGATCACGTCATGACTCCGAACAGGAT AATTCAGACAACAACACCATCTTTGTGCAAGGCCTGGGTG AGAATGTTACAATTGAGTCTGTGGCTGATTACTTCAAGCA GATTGGTATTATTAAGACAAACAAGAAAACGGGACAGCCC ATGATTAATTTGTACACAGACAGGGAAACTGGCAAGCTGA AGGGAGAGGCAACGGTCTCTTTTGATGACCCACCTTCAGC TAAAGCAGCTATTGACTGGTTTGATGGTAAAGAATTCTCC GGAAATCCTATCAAGGTCTCATTTGCTACTCGCCGGGCAG ACTTTAATCGGGGTGGTGGCAATGGTCGTGGAGGCCGAGG GCGAGGAGGACCCATGGGCCGTGGAGGCTATGGAGGTGGT GGCAGTGGTGGTGGTGGCCGAGGAGGATTTCCCAGTGGAG GTGGTGGCGGTGGAGGACAGCAGCGAGCTGGTGACTGGAA GTGTCCTAATCCCACCTGTGAGAATATGAACTTCTCTTGG AGGAATGAATGCAACCAGTGTAAGGCCCCTAAACCAGATG GCCCAGGAGGGGGACCAGGTGGCTCTCACATGGGGGGTAA CTACGGGGATGATCGTCGTGGTGGCAGAGGAGGC(SEQ ID NO: 3 3)

[0381] protein: MASNDYTQQATQSYGAYPTQPGQGYSQQSSQPYGQQSYSG YSQSTDTSGYGQSSYSSYGQSQNTGYGTQSTPQGYGSTGG YGSSQSSQSSYGQQSSYPGYGQQPAPSSTSGSYGSSSQSS SYGQPQSGSYSQQPSYGGQQQSYGQQQSYNPPQGYGQQNQ YNSSSGGGGGGGGGGNYGQDQSSMSSGGGSGGGYGNQDQS GGGGSGGYGQQDRGGRGRGGSGGGGGGGGGGYNRSSGGYE PRGRGGGRGGRGGMGGSDRGGFNKFGGPRDQGSRHDSEQD NSDNNTIFVQGLGENVTIESVADYFKQIGIIKTNKKTGQP MINLYTDRETGKLKGEATVSFDDPPSAKAAIDWFDGKEFS GNPIKVSFATRRADFNRGGGNGRGGRGRGGPMGRGGYGGG GSGGGGRGGFPSGGGGGGGQQRAGDWKCPNPTCENMNFSW RNECNQCKAPKPDGPGGGPGGSHMGGNYGDDRRGGRGG(sequence number 34)

[0382] EWSR1 (Homo sapiens Ewing sarcoma breakpoint region 1; Un iprot:Q01844)

[0383] DNA: ATGGCGTCCACGGATTACAGTACCTATAGCCAAGCTGCAG CGCAGCAGGGCTACAGTGCTTACACCGCCCAGCCCACTCA AGGATATGCACAGACCACCCAGGCATATGGGCAACAAAGC TATGGAACCTATGGACAGCCCACTGATGTCAGCTATACCC AGGCTCAGACCACTGCAACCTATGGGCAGACCGCCTATGC AACTTCTTATGGACAGCCTCCCACTGGTTATACTACTCCA ACTGCCCCCCAGGCATACAGCCAGCCTGTCCAGGGGTATG GCACTGGTGCTTATGATACCACCACTGCTACAGTCACCAC CACCCAGGCCTCCTATGCAGCTCAGTCTGCATATGGCACT CAGCCTGCTTATCCAGCCTATGGGCAGCAGCCAGCAGCCA CTGCACCTACAAGACCGCAGGATGGAAACAAGCCCACTGA GACTAGTCAACCTCAATCTAGCACAGGGGGTTACAACCAG CCCAGCCTAGGATATGGACAGAGTAACTACAGTTATCCCC AGGTACCTGGGAGCTACCCCATGCAGCCAGTCACTGCACC TCCATCCTACCCTCCTACCAGCTATTCCTCTACACAGCCG ACTAGTTATGATCAGAGCAGTTACTCTCAGCAGAACACCT ATGGGCAACCGAGCAGCTATGGACAGCAGAGTAGCTATGG TCAACAAAGCAGCTATGGGCAGCAGCCTCCCACTAGTTAC CCACCCCAAACTGGATCCTACAGCCAAGCTCCAAGTCAAT ATAGCCAACAGAGCAGCAGCTACGGGCAGCAGAGTTCATT CCGACAGGACCACCCCAGTAGCATGGGTGTTTATGGGCAG GAGTCTGGAGGATTTTCCGGACCAGGAGAGAACCGGAGCA TGAGTGGCCCTGATAACCGGGGCAGGGGAAGAGGGGGATT TGATCGTGGAGGCATGAGCAGAGGTGGGCGGGGAGGAGGA CGCGGTGGAATGGGCAGCGCTGGAGAGCGAGGTGGCTTCA ATAAGCCTGGTGGACCCATGGATGAAGGACCAGATCTTGA TCTAGGCCCACCTGTAGATCCAGATGAAGACTCTGACAAC AGTGCAATTTATGTACAAGGATTAAATGACAGTGTGACTC TAGATGATCTGGCAGACTTCTTTAAGCAGTGTGGGGTTGT TAAGATGAACAAGAGAACTGGGCAACCCATGATCCACATC TACCTGGACAAGGAAACAGGAAAGCCCAAAGGCGATGCCA CAGTGTCCTATGAAGACCCACCCACTGCCAAGGCTGCCGT GGAATGGTTTGATGGGAAAGATTTTCAAGGGAGCAAACTT AAAGTCTCCCTTGCTCGGAAGAAGCCTCCAATGAACAGTA TGCGGGGTGGTCTGCCACCCCGTGAGGGCAGAGGCATGCC ACCACCACTCCGTGGAGGTCCAGGAGGCCCAGGAGGTCCT GGGGGACCCATGGGTCGCATGGGAGGCCGTGGAGGAGATA GAGGAGGCTTCCCTCCAAGAGGACCCCGGGGTTCCCGAGG GAACCCCTCTGGAGGAGGAAACGTCCAGCACCGAGCTGGA GACTGGCAGTGTCCCAATCCGGGTTGTGGAAACCAGAACT TCGCCTGGAGAACAGAGTGCAACCAGTGTAAGGCCCCAAA GCCTGAAGGCTTCCTCCCGCCACCCTTTCCGCCCCCGGGT GGTGATCGTGGCAGAGGTGGCCCTGGTGGCATGCGGGGAG GAAGAGGTGGCCTCATGGATCGTGGTGGTCCCGGTGGAAT GTTCAGAGGTGGCCGTGGTGGAGACAGAGGTGGCTTCCGT GGTGGCCGGGGCATGGACCGAGGTGGCTTTGGTGGAGGAA GACGAGGTGGCCCTGGGGGGCCCCCTGGACCTTTGATGGA ACAG (SEQ ID NO: 35)

[0384] protein: MASTDYSTYSQAAAQQGYSAYTAQPTQGYAQTTQAYGQQS YGTYGQPTDVSYTQAQTTATYGQTAYATSYGQPPTGYTTP TAPQAYSQPVQGYGTGAYDTTTATVTTTQASYAAQSAYGT QPAYPAYGQQPAATAPTRPQDGNKPTETSQPQSSTGGYNQ PSLGYGQSNYSYPQVPGSYPMQPVTAPPSYPPTSYSSTQP TSYDQSSYSQQNTYGQPSSYGQQSSYGQQSSYGQQPPTSY PPQTGSYSQAPSQYSQQSSSYGQQSSFRQDHPSSMGVYGQ ESGGFSGPGENRSMSGPDNRGRGRGGFDRGGMSRGGRGGG RGGMGSAGERGGFNKPGGPMDEGPDLDLGPPVDPDEDSDN SAIYVQGLNDSVTLDDLADFFKQCGVVKMNKRTGQPMIHI YLDKETGKPKGDATVSYEDPPTAKAAVEWFDGKDFQGSKL KVSLARKPPMNSMRGGLPPREGRGMPPPLRGGPGPGPGGP GGPMGRMGGRGGDRGGFPPRGPRGSRGNPSGGGNVQHRAG DWQCPNPGCGNQNFAWRTECNQCKAPKPEGFLPPPFPPPG GDRGRGGPGGMRGGRGGLMDRGGPGGMFRGGRGGDRGGFR GGRGMDRGGFGGGRRGGPGGPPGPLMEQ(SEQ ID NO: 36)

[0385] 8. AFP EWSR1-MCP

[0386] DNA: ATGGCGTCCACGGATTACAGTACCTATAGCCAAGCTGCAG CGCAGCAGGGCTACAGTGCTTACACCGCCCAGCCCACTCA AGGATATGCACAGACCACCCAGGCATATGGGCAACAAAGC TATGGAACCTATGGACAGCCCACTGATGTCAGCTATACCC AGGCTCAGACCACTGCAACCTATGGGCAGACCGCCTATGC AACTTCTTATGGACAGCCTCCCACTGGTTATACTACTCCA ACTGCCCCCCAGGCATACAGCCAGCCTGTCCAGGGGTATG GCACTGGTGCTTATGATACCACCACTGCTACAGTCACCAC CACCCAGGCCTCCTATGCAGCTCAGTCTGCATATGGCACT CAGCCTGCTTATCCAGCCTATGGGCAGCAGCCAGCAGCCA CTGCACCTACAAGACCGCAGGATGGAAACAAGCCCACTGA GACTAGTCAACCTCAATCTAGCACAGGGGGTTACAACCAG CCCAGCCTAGGATATGGACAGAGTAACTACAGTTATCCCC AGGTACCTGGGAGCTACCCCATGCAGCCAGTCACTGCACC TCCATCCTACCCTCCTACCAGCTATTCCTCTACACAGCCG ACTAGTTATGATCAGAGCAGTTACTCTCAGCAGAACACCT ATGGGCAACCGAGCAGCTATGGACAGCAGAGTAGCTATGG TCAACAAAGCAGCTATGGGCAGCAGCCTCCCACTAGTTAC CCACCCCAAACTGGATCCTACAGCCAAGCTCCAAGTCAAT ATAGCCAACAGAGCAGCAGCTACGGGCAGCAGAGTTCATT CCGACAGGACCACCCCAGTAGCATGGGTGTTTATGGGCAG GAGTCTGGAGGATTTTCCGGACCAGGAGAGAACCGGAGCA TGAGTGGCCCTGATAACCGGGGCAGGGGAAGAGGGGGATT TGATCGTGGAGGCATGAGCAGAGGTGGGCGGGGAGGAGGA CGCGGTGGAATGGGCAGCGCTGGAGAGCGAGGTGGCTTCA ATAAGCCTGGTGGACCCATGGATGAAGGACCAGATCTTGA TCTAGGCCCACCTGTAGATCCAGATGAAGACTCTGACAAC AGTGCAATTTATGTACAAGGATTAAATGACAGTGTGACTC TAGATGATCTGGCAGACTTCTTTAAGCAGTGTGGGGTTGT TAAGATGAACAAGAGAACTGGGCAACCCATGATCCACATC TACCTGGACAAGGAAACAGGAAAGCCCAAAGGCGATGCCA CAGTGTCCTATGAAGACCCACCCACTGCCAAGGCTGCCGT GGAATGGTTTGATGGGAAAGATTTTCAAGGGAGCAAACTT AAAGTCTCCCTTGCTCGGAAGAAGCCTCCAATGAACAGTA TGCGGGGTGGTCTGCCACCCCGTGAGGGCAGAGGCATGCC ACCACCACTCCGTGGAGGTCCAGGAGGCCCAGGAGGTCCT GGGGGACCCATGGGTCGCATGGGAGGCCGTGGAGGAGATA GAGGAGGCTTCCCTCCAAGAGGACCCCGGGGTTCCCGAGG GAACCCCTCTGGAGGAGGAAACGTCCAGCACCGAGCTGGA GACTGGCAGTGTCCCAATCCGGGTTGTGGAAACCAGAACT TCGCCTGGAGAACAGAGTGCAACCAGTGTAAGGCCCCAAA GCCTGAAGGCTTCCTCCCGCCACCCTTTCCGCCCCCGGGT GGTGATCGTGGCAGAGGTGGCCCTGGTGGCATGCGGGGAG GAAGAGGTGGCCTCATGGATCGTGGTGGTCCCGGTGGAAT GTTCAGAGGTGGCCGTGGTGGAGACAGAGGTGGCTTCCGT GGTGGCCGGGGCATGGACCGAGGTGGCTTTGGTGGAGGAA GACGAGGTGGCCCTGGGGGGCCCCCTGGACCTTTGATGGA ACAGGATTACAAGGATGACGACGATAAGGGTACCGAGCAG AAGCTGATCTCAGAGGAGGACCTGGGCGCCCCCGGCTCCG CCGGCTCCGCCGCCGGCTCCGGCGCTTCTAACTTTACTCA GTTCGTTCTCGTCGACAATGGCGGAACTGGCGACGTGACT GTCGCCCCAAGCAACTTCGCTAACGGGATCGCTGAATGGA TCAGCTCTAACTCGCGTTCACAGGCTTACAAAGTAACCTG TAGCGTTCGTCAGAGCTCTGCGCAGAATCGCAAATACACC ATCAAAGTCGAGGTGCCTAAAGGCGCCTGGCGTTCGTACT TAAATATGGAACTAACCATTCCAATTTTCGCCACGAATTC CGACTGCGAGCTTATTGTTAAGGCAATGCAAGGTCTCCTA AAAGATGGAAACCCGATTCCCTCAGCAATCGCAGCAAACT CCGGCATCTACTAA(SEQ ID NO: 37)

[0387] protein: MASTDYSTYSQAAAQQGYSAYTAQPTQGYAQTTQAYGQQS YGTYGQPTDVSYTQAQTTATYGQTAYATSYGQPPTGYTTP TAPQAYSQPVQGYGTGAYDTTTATVTTTQASYAAQSAYGT QPAYPAYGQQPAATAPTRPQDGNKPTETSQPQSSTGGYNQ PSLGYGQSNYSYPQVPGSYPMQPVTAPPSYPPTSYSSTQP TSYDQSSYSQQNTYGQPSSYGQQSSYGQQSSYGQQPPTSY PPQTGSYSQAPSQYSQQSSSYGQQSSFRQDHPSSMGVYGQ ESGGFSGPGENRSMSGPDNRGRGRGGFDRGGMSRGGRGGG RGGMGSAGERGGFNKPGGPMDEGPDLDLGPPVDPDEDSDN SAIYVQGLNDSVTLDDLADFFKQCGVVKMNKRTGQPMIHI YLDKETGKPKGDATVSYEDPPTAKAAVEWFDGKDFQGSKL KVSLARKKPPMNSMRGGLPPREGRGMPPPLRGGPGGPGGP GGPMGRMGGRGGDRGGFPPRGPRGSRGNPSGGGNVQHRAG DWQCPNPGCGNQNFAWRTECNQCKAPKPEGFLPPPFPPPG GDRGRGGPGGMRGGRGGLMDRGGPGGMFRGGRGGDRGGFR GGRGMDRGGFGGGRRGGPGGPPGPLMEQDYKDDDDKGTEQ KLISEEDLGAPGSAGSAAGSGASNFTQFVLVDNGGTGDVT VAPSNFANGIAEWISSNSRSQAYKVTCSVRQSSAQNRKYT IKVEVPKGAWRSYLNMELTIPIFATNSDCELIVKAMQGLL KDGNPIPSAIAANSGIY (SEQ ID NO: 38)

[0388] FUS-MCP

[0389] DNA: ATGGCCTCAAACGATTATACCCAACAAGCAACCCAAAGCT ATGGGGCCTACCCCACCCAGCCCGGGCAGGGCTATTCCCA GCAGAGCAGTCAGCCCTACGGACAGCAGAGTTACAGTGGT TATAGCCAGTCCACGGACACTTCAGGATATGGCCAGAGCA GCTATTCTTCTTATGGCCAGAGCCAGAACACAGGCTATGG AACTCAGTCAACTCCCCAGGGATATGGCTCGACTGGCGGC TATGGCAGTAGCCAGAGCTCCCAATCGTCTTACGGGCAGC AGTCCTCCTACCCTGGCTATGGCCAGCAGCCAGCTCCCAG CAGCACCTCGGGAAGTTACGGTAGCAGTTCTCAGAGCAGC AGCTATGGGCAGCCCCAGAGTGGGAGCTACAGCCAGCAGC CTAGCTATGGTGGACAGCAGCAAAGCTATGGACAGCAGCA AAGCTATAATCCCCCTCAGGGCTATGGACAGCAGAACCAG TACAACAGCAGCAGTGGTGGTGGAGGTGGAGGTGGAGGTG GAGGTAACTATGGCCAAGATCAATCCTCCATGAGTAGTGG TGGTGGCAGTGGTGGCGGTTATGGCAATCAAGACCAGAGT GGTGGAGGTGGCAGCGGTGGCTATGGACAGCAGGACCGTG GAGGCCGCGGCAGGGGTGGCAGTGGTGGCGGCGGCGGCGG CGGCGGTGGTGGTTACAACCGCAGCAGTGGTGGCTATGAA CCCAGAGGTCGTGGAGGTGGCCGTGGAGGCAGAGGTGGCA TGGGCGGAAGTGACCGTGGTGGCTTCAATAAATTTGGTGG CCCTCGGGACCAAGGATCACGTCATGACTCCGAACAGGAT AATTCAGACAACAACACCATCTTTGTGCAAGGCCTGGGTG AGAATGTTACAATTGAGTCTGTGGCTGATTACTTCAAGCA GATTGGTATTATTAAGACAAACAAGAAAACGGGACAGCCC ATGATTAATTTGTACACAGACAGGGAAACTGGCAAGCTGA AGGGAGAGGCAACGGTCTCTTTTGATGACCCACCTTCAGC TAAAGCAGCTATTGACTGGTTTGATGGTAAAGAATTCTCC GGAAATCCTATCAAGGTCTCATTTGCTACTCGCCGGGCAG ACTTTAATCGGGGTGGTGGCAATGGTCGTGGAGGCCGAGG GCGAGGAGGACCCATGGGCCGTGGAGGCTATGGAGGTGGT GGCAGTGGTGGTGGTGGCCGAGGAGGATTTCCCAGTGGAG GTGGTGGCGGTGGAGGACAGCAGCGAGCTGGTGACTGGAA GTGTCCTAATCCCACCTGTGAGAATATGAACTTCTCTTGG AGGAATGAATGCAACCAGTGTAAGGCCCCTAAACCAGATG GCCCAGGAGGGGGACCAGGTGGCTCTCACATGGGGGGTAA CTACGGGGATGATCGTCGTGGTGGCAGAGGAGGCGATTAC AAGGATGACGACGATAAGGGTACCGAGCAGAAGCTGATCT CAGAGGAGGACCTGGGCGCCCCCGGCTCCGCCGGCTCCGC CGCCGGCTCCGGCGCTTCTAACTTTACTCAGTTCGTTCTC GTCGACAATGGCGGAACTGGCGACGTGACTGTCGCCCCAA GCAACTTCGCTAACGGGATCGCTGAATGGATCAGCTCTAA CTCGCGTTCACAGGCTTACAAAGTAACCTGTAGCGTTCGT CAGAGCTCTGCGCAGAATCGCAAATACACCATCAAAGTCG AGGTGCCTAAAGGCGCCTGGCGTTCGTACTTAAATATGGA ACTAACCATTCCAATTTTCGCCACGAATTCCGACTGCGAG CTTATTGTTAAGGCAATGCAAGGTCTCCTAAAAGATGGAA ACCCGATTCCCCTCAGCAATCGCAGCAAACTCCGGCATCTA CTAA (SEQ ID NO: 39)

[0390] protein: MASNDYTQQATQSYGAYPTQPGQGYSQQSSQPYGQQSYSG YSQSTDTSGYGQSSYSSYGQSQNTGYGTQSTPQGYGSTGG YGSSQSSQSSYGQQSSYPGYGQQPAPSSTSGSYGSSSQSS SYGQPQSGSYSQQPSYGGQQQSYGQQQSYNPPQGYGQQNQ YNSSSGGGGGGGGGGNYGQDQSSMSSGGGSGGGYGNQDQS GGGGSGGYGQQDRGGRGRGGSGGGGGGGGGGYNRSSGGYE PRGRGGGRGGRGGMGGSDRGGFNKFGGPRDQGSRHDSEQD NSDNNTIFVQGLGENVTIESVADYFKQIGIIKTNKKTGQP MINLYTDRETGKLKGEATVSFDDPPSAKAAIDWFDGKEFS GNPIKVSFATRRADFNRGGGGNGRGGRGRGGPMGRGGYGGG GSGGGGRGGFPSGGGGGGGQQRAGDWKCPNPTCENMNFSW RNECNQCKAPKPDGPGGGPGGSHMGGNYGDDRRGGRGGDY KDDDDKGTEQKLISEEDLGAPGSAGSAAGSGASNFTQFVL VDNGGTGDVTVAPSNFANGIAEWISSNSRSQAYKVTCSVR QSSAQNRKYTIKVEVPKGAWRSYLNMELTIPIFATNSDCE LIVKAMQGLLKDGNPIPSAIAANSGIY (Sequence number 40)

[0391] FUS - PylRS AF

[0392] DNA: ATGGCCTCAAACGATTATACCCAACAAGCAACCCAAAGCT ATGGGGCCTACCCCACCCAGCCCGGGCAGGGCTATTCCCA GCAGAGCAGTCAGCCCTACGGACAGCAGAGTTACAGTGGT TATAGCCAGTCCACGGACACTTCAGGATATGGCCAGAGCA GCTATTCTTCTTATGGCCAGAGCCAGAACACAGGCTATGG AACTCAGTCAACTCCCCAGGGATATGGCTCGACTGGCGGC TATGGCAGTAGCCAGAGCTCCCAATCGTCTTACGGGCAGC AGTCCTCCTACCCTGGCTATGGCCAGCAGCCAGCTCCCAG CAGCACCTCGGGAAGTTACGGTAGCAGTTCTCAGAGCAGC AGCTATGGGCAGCCCCAGAGTGGGAGCTACAGCCAGCAGC CTAGCTATGGTGGACAGCAGCAAAGCTATGGACAGCAGCA AAGCTATAATCCCCCTCAGGGCTATGGACAGCAGAACCAG TACAACAGCAGCAGTGGTGGTGGAGGTGGAGGTGGAGGTG GAGGTAACTATGGCCAAGATCAATCCTCCATGAGTAGTGG TGGTGGCAGTGGTGGCGGTTATGGCAATCAAGACCAGAGT GGTGGAGGTGGCAGCGGTGGCTATGGACAGCAGGACCGTG GAGGCCGCGGCAGGGGTGGCAGTGGTGGCGGCGGCGGCGG CGGCGGTGGTGGTTACAACCGCAGCAGTGGTGGCTATGAA CCCAGAGGTCGTGGAGGTGGCCGTGGAGGCAGAGGTGGCA TGGGCGGAAGTGACCGTGGTGGCTTCAATAAATTTGGTGG CCCTCGGGACCAAGGATCACGTCATGACTCCGAACAGGAT AATTCAGACAACAACACCATCTTTGTGCAAGGCCTGGGTG AGAATGTTACAATTGAGTCTGTGGCTGATTACTTCAAGCA GATTGGTATTATTAAGACAAACAAGAAAACGGGACAGCCC ATGATTAATTTGTACACAGACAGGGAAACTGGCAAGCTGA AGGGAGAGGCAACGGTCTCTTTTGATGACCCACCTTCAGC TAAAGCAGCTATTGACTGGTTTGATGGTAAAGAATTCTCC GGAAATCCTATCAAGGTCTCATTTGCTACTCGCCGGGCAG ACTTTAATCGGGGTGGTGGCAATGGTCGTGGAGGCCGAGG GCGAGGAGGACCCATGGGCCGTGGAGGCTATGGAGGTGGT GGCAGTGGTGGTGGTGGCCGAGGAGGATTTCCCAGTGGAG GTGGTGGCGGTGGAGGACAGCAGCGAGCTGGTGACTGGAA GTGTCCTAATCCCACCTGTGAGAATATGAACTTCTCTTGG AGGAATGAATGCAACCAGTGTAAGGCCCCTAAACCAGATG GCCCAGGAGGGGGACCAGGTGGCTCTCACATGGGGGGTAA CTACGGGGATGATCGTCGTGGTGGCAGAGGAGGCGATTAC AAGGATGACGACGATAAGGGTACCGGCGCCCCCGGCTCCG CCGGCTCCGCCGCCGGCTCCGGCGCTTCTAACTTTACTCA GTTCGTTCTCGTCGACAATGGCGGAACTGGCGACGTGACT GTCGCCCCAAGCAACTTCGCTAACGGGATCGCTGAATGGA TCAGCTCTAACTCGCGTTCACAGGCTTACAAAGTAACCTG TAGCGTTCGTCAGAGCTCTGCGCAGAATCGCAAATACACC ATCAAAGTCGAGGTGCCTAAAGGCGCCTGGCGTTCGTACT TAAATATGGAACTAACCATTCCAATTTTCGCCACGAATTC CGACTGCGAGCTTATTGTTAAGGCAATGCAAGGTCTCCTA AAAGATGGAAACCCGATTCCCTCAGCAATCGCAGCAAACT CCGGCATCTACGGTACCGGCGCCCCCGGCTCCGCCGGCTC CGCCGCCGGCTCCGGCGCGTGCCCGGTGCCGCTGCAGCTG CCGCCGCTGGAACGCCTGACCCTGGATGATAAAAAACCGC TGAATACCCTGATCTCTGCTACTGGTCTGTGGATGAGTCG TACCGGAACCATTCATAAAATCAAACACCACGAGGTTAGC CGTTCGAAAATCTATATTGAGATGGCGTGTGGCGATCATC TGGTTGTGAACAATAGCCGCTCTTCTCGTACAGCACGTGC ACTGCGTCACCACAAATATCGTAAAACCTGTAAACGTTGC CGTGTGTCCGATGAGGATCTGAACAAATTCCTGACAAAAG CCAATGAGGACCAAACAAGCGTGAAAGTGAAAGTCGTTAG CGCTCCTACCCGTACTAAAAAAGCAATGCCGAAATCCGTT GCTCGTGCCCCTAAACCACTGGAAAACACTGAAGCAGCAC AGGCACAGCCGTCTGGAAGCAAATTCTCTCCGGCCATTCC TGTTTCTACCCAGGAGTCCGTTTCTGTTCCAGCAAGTGTG AGCACCAGCATTAGCAGTATTAGCACCGGTGCCACCGCTA GCGCCCTGGTTAAAGGCAATACCAATCCGATTACAAGCAT GTCTGCCCCGGTTCAAGCATCAGCTCCAGCACTGACAAAA TCCCAAACCGATCGTCTGGAGGTTCTGCTGAATCCGAAAG ACGAAATCAGCCTGAATTCCGGCAAACCGTTTCGTGAACT GGAGAGCGAACTGCTGTCACGTCGTAAAAAAGACCTGCAA CAAATCTATGCCGAAGAACGTGAGAACTATCTGGGGAAAC TGGAACGTGAAATCACCCGCTTTTTCGTGGATCGTGGCTT TCTGGAGATCAAATCCCCGATTCTGATTCCTCTGGAGTAT ATCGAGCGTATGGGCATCGACAATGATACCGAACTGAGCA AACAAATTTTCCGTGTGGATAAAAACTTCTGTCTGCGCCC TATGCTAGCACCAAATCTGGCTAACTATCTGCGCAAACTG GACCGTGCCCTGCCTGATCCTATCAAAATCTTCGAGATCG GCCCGTGTTATCGTAAAGAGTCCGACGGTAAAGAACATCT GGAGGAGTTTACCATGCTGAACTTTTGCCAAATGGGTTCA GGTTGTACTCGTGAGAACCTGGAAAGCATCATCACCGATT TTCTGAACCACCTGGGCATTGACTTCAAAATTGTGGGCGA CAGCTGTATGGTGTTTGGCGACACCCTGGATGTCATGCAC GGCGACCTGGAACTGTCTAGTGCCGTTGTGGGCCCAATCC CGCTGGATCGTGAGTGGGGTATCGACAAACCTTGGATCGG TGCGGGTTTTGGTCTGGAGCGTCTGCTGAAAGTAAAACAC GACTTCAAGAACATCAAACGTGCTGCACGTTCCGAGTCCT ATTACAATGGTATTTCTACTAACCTGTAA(SEQ ID NO: 41)

[0393] protein: MASNDYTQQATQSYGAYPTQPGQGYSQQSSQPYGQQSYSG YSQSTDTSGYGQSSYSSYGQSQNTGYGTQSTPQGYGSTGG YGSSQSSQSSYGQQSSYPGYGQQPAPSSTSGSYGSSSQSS SYGQPQSGSYSQQPSYGGQQQSYGQQQSYNPPQGYGQQNQ YNSSSGGGGGGGGGGNYGQDQSSMSSGGGSGGGYGNQDQS GGGGSGGYGQQDRGGGRGRGSGGGGGGGGGYNRSSGGYE PRGRGGGRGRGGMGGSDRGGFNKFGGPRDQGSRHDSEQD NSDNNTIFVQGLGENVTIESVADYFKQIGIIKTNKKTGQP MINLYTDRETGKLKGEATVSFDDPPSAKAAIDWFDGKEFS GNPIKVSFATRRADFNRGGGNGRGGRGRGGPMGRGGYGGG GSGGGGRGGFPSGGGGGGQQRAGDWKCPNPTCENMNFSW RNECNQCKAPKPDGPGGGPGGSHMGGNYGDDRRGGRGGDY KDDDDKGTGAPGSAGSAAGSGASNFTQFVLVDNGGTGDVT VAPSNFANGIAEWISSNSRSQAYKVTCSVRQSSAQNRKYT IKVEVPKGAWRSYLNMELTIPIFATNSDCELIVKAMQGLL KDGNPIPSAIAANSGIYGTGAPGSAGSAGSAGSGACPVPLQL PPLERLTLDDKKPLNTLISATGLWMSRTGTIHKIKHHEVS RSKIYIEMACGDHLVVNNSRSSTARARRHHKYRKTCKRC RVSDEDLDNKFLTKANEDQTSVKVKVVSAPTRTKKAMPKSV ARAPKPLENTEAAQAQPSGSKFSPAIPVSTQESVSVPASV STSISSISTGATASALVKGNTNPITSMSPAVQASAPALTK SQTDRLEVLLNPKDEISLNSGKPFRELESELLSRRKKDLQ QIYAEERENYLGKLEREITRFVDRGFLEIKSPILIPLEY IERMGIDNDTELSKQIFRVDKNFCLRPMLAPNLANYLRKL DRALPDPIKIFEIGPCYRKESDGKEHLEEFTMLNFCQMGS GCTRENLESIITDFLNHLGIDFKIVGDSCMVFGDTLDVMH GDLELSSAVVGPIPLDREWGIDKPWIGAGFGLERLLKVKH DFKNIKRAARSESYYNGISTNL (SEQ ID NO: 42)

[0394] MCP-PylRS AF

[0395] DNA: ATGGCTTCTAACTTTACTCAGTTCGTTCTCGTCGACAATG GCGGAACTGGCGACGTGACTGTCGCCCCAAGCAACTTCGC TAACGGGATCGCTGAATGGATCAGCTCTAACTCGCGTTCA CAGGCTTACAAAGTAACCTGTAGCGTTCGTCAGAGCTCTG CGCAGAATCGCAAATACACCATCAAAGTCGAGGTGCCTAA AGGCGCCTGGCGTTCGTACTTAAATATGGAACTAACCATT CCAATTTTCGCCACGAATTCCGACTGCGAGCTTATTGTTA AGGCAATGCAAGGTCTCCTAAAAGATGGAAACCCGATTCC CTCAGCAATCGCAGCAAACTCCGGCATCTACGGTACCGGC GCCCCCGGCTCCGCCGGCTCCGCCGCCGGCTCCGGCGCGT GCCCGGTGCCGCTGCAGCTGCCGCCGCTGGAACGCCTGAC CCTGGATGATAAAAAACCGCTGAATACCCTGATCTCTGCT ACTGGTCTGTGGATGAGTCGTACCGGAACCATTCATAAAA TCAAACACCACGAGGTTAGCCGTTCGAAAATCTATATTGA GATGGCGTGTGGCGATCATCTGGTTGTGAACAATAGCCGC TCTTCTCGTACAGCACGTGCACTGCGTCACCACAAATATC GTAAAACCTGTAAACGTTGCCGTGTGTCCGATGAGGATCT GAACAAATTCCTGACAAAAGCCAATGAGGACCAAACAAGC GTGAAAGTGAAAGTCGTTAGCGCTCCTACCCGTACTAAAA AAGCAATGCCGAAATCCGTTGCTCGTGCCCCTAAACCACT GGAAAACACTGAAGCAGCACAGGCACAGCCGTCTGGAAGC AAATTCTCTCCGGCCATTCCTGTTTCTACCCAGGAGTCCG TTTCTGTTCCAGCAAGTGTGAGCACCAGCATTAGCAGTAT TAGCACCGGTGCCACCGCTAGCGCCCTGGTTAAAGGCAAT ACCAATCCGATTACAAGCATGTCTGCCCCGGTTCAAGCAT CAGCTCCAGCACTGACAAAATCCCAAACCGATCGTCTGGA GGTTCTGCTGAATCCGAAAGACGAAATCAGCCTGAATTCC GGCAAACCGTTTCGTGAACTGGAGAGCGAACTGCTGTCAC GTCGTAAAAAAGACCTGCAACAAATCTATGCCGAAGAACG TGAGAACTATCTGGGGAAACTGGAACGTGAAATCACCCGC TTTTTCGTGGATCGTGGCTTTCTGGAGATCAAATCCCCGA TTCTGATTCCTCTGGAGTATATCGAGCGTATGGGCATCGA CAATGATACCGAACTGAGCAAACAAATTTTCCGTGTGGAT AAAAACTTCTGTCTGCGCCCTATGCTAGCACCAAATCTGG CTAACTATCTGCGCAAACTGGACCGTGCCCTGCCTGATCC TATCAAAATCTTCGAGATCGGCCCGTGTTATCGTAAAGAG TCCGACGGTAAAGAACATCTGGAGGAGTTTACCATGCTGA ACTTTTGCCAAATGGGTTCAGGTTGTACTCGTGAGAACCT GGAAAGCATCATCACCGATTTTCTGAACCACCTGGGCATT GACTTCAAAATTGTGGGCGACAGCTGTATGGTGTTTGGCG ACACCCTGGATGTCATGCACGGCGACCTGGAACTGTCTAG TGCCGTTGTGGGCCCAATCCCGCTGGATCGTGAGTGGGGT ATCGACAAACCTTGGATCGGTGCGGGTTTTGGTCTGGAGC GTCTGCTGAAAGTAAAACACGACTTCAAGAACATCAAACG TGCTGCACGTTCCGAGTCCTATTACAATGGTATTTCTACT AACCTGTAA(SEQ ID NO: 43)

[0396] protein: MASNFTQFVLVDNGGTGDVTVAPSNFANGIAEWISSNSRS QAYKVTCSVRQSSAQNRKYTIKVEVPKGAWRSYLNMELTI PIFATNSDCELIVKAMQGLLKDGNPIPSAIAANSGIYGTG APGSAGSAAGSGACPVPLQLPPLERLTLDDKKPLNTLISA TGLWMSRTGTIHKIKHHEVSRSKIYIEMACGDHLVVNNSR SSRTARALRHHKYRKTCKRCRVSDEDLNKFLTKANEDQTS VKVKVVSAPTRTKKAMPKSVARAPKPLENTEAAQAQPSGS KFSPAIPVSTQESVSVPASVSTSISSISTGATASALVKGN TNPITSMSAPVQASAPALTKSQTDRLEVLLNPKDEISLNS GKPFRELESELLSRRKKDLQQIYAEERENYLGKLEREITR FFVDRGFLEIKSPILIPLEYIERMGIDNDTELSKQIFRVD KNFCLRPMLAPNLANYLRKLDRALPDPIKIFEIGPCYRKE SDGKEHLEEFTMLNFCQMGSGCTRENLESIITDFLNHLGI DFKIVGDSCMVFGDTLDVMHGDLELSSAVVGPIPLDREWG IDKPWIGAGFGLERLLKVKHDFKNIKRAARSESYYNGIST NL (SEQ ID NO: 44)

[0397] SPD5-MCP

[0398] DNA: ATGGAGGACAACAGCGTGCTGAACGAGGACAGCAACCTGG AGCACGTGGAGGGCCAGCCCAGAAGAAGCATGAGCCAGCC CGTGCTGAACGTGGAGGGCGACAAGAGAACCAGCAGCACC AGCGCCACCCAGCAGCAGGTGCTGAGCGGCGCCTTCAGCA GCGCCGACGTGAGAAGCATCCCCATCATCCAGACCTGGGA GGAGAACAAGGCCCTGAAGACCAAGATCACCATCCTGAG GGCGAGCTGCAGATGTACCAGAGAAGATACAGCGAGGCCA AGGAGGCCAGCCAGCAAGAGAGTGAAGGAGGTGATGGACGA CTACGTGGACCTGAAGCTGGCCAGGAGAACGTGCAGGAG AAGATGGAGCAGTACAAGCTGATGGAGGAGGAGCACCTGCTGG CCATGCAGAGCAGAATCGAGACCAGCGAGAGACAACTTCGC CAGACAGATGAAGGAGTTCGAGGCCCAGAAGCACGCCATG GAGGAGAGAATCAAGGAGCTGGAGCTGAGCGCCACCGACG CCAACAACACCACCGTGGGCAGCTTCAGAGGCACCCTGGA CGACATCCTGAAAAGAACGACCCCGACTTCACCCTGACC AGCGGCTACGAGGAGAGAAAGATCAACGACCTGGAGGCCA AGCTGCTGAGCGAGATCGACAAGGTGGCCGAGCTGGAGGA CCACATCCAGCAGCTGAGACAGGAGCTGGACGACCAGAGC GCCAGACTGGCCGACAGCGAGAACGTGAGAGCCCAGCTGG AGGCCGCCACCGGCCAGGGCATCCTGGGCGCCGCCGGCAA CGCCATGGTGCCCAACAGCACCTTCCATGATCGGCAACGGC AGAGAGAGCCAGACCAGAGACCAGACCAGCTGAACTACATCGACG ACCTGGAGACCAAGCTGGCCGACGCCAAGAAGGAGAACGA CAAGGCCAGACAGGCCCTGGTGGAGTACATGAACAAGTGC AGCAAGCTGGAGCACGAGATCAGAACCATGGTGAAGAAACA GCACCTTCGACAGCAGCAGCATGCTGCTGGGCGGCCAGAC CAGCGACGAGCTGAAGATCCAGATCGGCAAGGTGAACGGC GAGCTGAACGTGCTGAGAGCCGAGAACAGAGAGCTGAGAA TCAGATGCGACCAGCTGACCGGCGGCGACGGCAACCTGAG CATCAGCCTGGGCCAGAGCAGACTGATGGCCGGCATCGCC ACCAACGACGTGGACAGCATCGGCCAGGGAACGAGACCG GCGGCACCAGCATGAGAATCCTGCCCAGAGAGAGCCAGCT GGACGACCTGGAGGAGAGCAAGCTGCCCCTGATGGACACCC AGCAGCGCCGTGAAACCAGCAGCAGTTCGCCAGCATGT GGGAGGACTTCGAGAGCGTGAAGGACAGCCTGCAGAACAA CCACAACGACACCCTGGAGGGCAGCTTCAACAGCAGCATG CCCCCCCCGGCAGAGACGCCACCCAGAGCTTCCTGAGCC AGAAGAGCTTCAAGAACAGCCCCCATCGTGATGCAGAGCC CAAGAGCCTGCACCTGCACCTGAAGAGCCACCAGAGCGAG GGCGCCGGCGAGCAGATCCAGAACAACAGCTTCAGCACCA AGACCGCCAGCCCCACGTGAGCCAGAGCCACATCCCCAT CCTGCACGACATGCAGCAGATCCTGGACAGCAGCGCCATG TTCCTGGAGGGCCAGCACGACGTGGCCGTGAACGTGGAGC AGATGCAGGAGAAGATGAGCCAGATCAGAGAGGCCCTGGC CAGACTGTTCGAGAGACTGAAGAGCAGCGCCGCCCTGTTC GAGGAGATCCTGGAGAGAATGGGCAGCAGCGACCCCAACG CCGACAAGATCAAGAAGATGAAGCTGGCCTTCGAGACCAG CATCAACGACAAGCTGAACGTGAGCGCCATCCTGGAGGCC GCCGAGAAGGACCTGCACAACATGAGCCTGAACTTCAGCA TCCTGGAGAAGAGCATCGTGAGCCAGGCCGCCGAGGCCAG CAGAAGATTCACCATCGCCCCCGACGCCGAGGACGTGGCC AGCAGCAGCCTGCTGAACGCCAGCTACAGCCCCCTGTTCA AGTTCACCAGCAACAGCGACATCGTGGAGAAGCTGCAGAA CGAGGTGAGCGAGCTGAAGAACGAGCTGGAGATGGCCAGA ACCAGAGACATGAGAAGCCCCCTGAACGGCAGCAGCGGCA GACTGAGCGACGTGCAGATCAACACCAACAGAATGTTCGA GGACCTGGAGGTGAGCGAGGCCACCCTGCAGAAGGCCAAG GAGGAGAACAGCACCCTGAAGAGCCAGTTCGCCGAGCTGG AGGCCAACCTGCACCAGGTGAACAGCAAGCTGGGCGAGGT GAGATGCGAGCTGAACGAGGCCCTGGCCAGAGTGGACGGC GAGCAGGAGACCAGAGTGAAGGCCGAGAACGCCCTGGAGG AGGCCAGACAGCTGATCAGCAGCCTGAAGCACGAGGAAA CGAGCTGAAGAGACCATCACCGACATGGGCATGAGACTG AACGAGGCCAAGAAGAGCGACGAGTTCCTGAAGAGCGAGC TGAGCACCCGCCCTGGAGGAGGAGAAGAGAGCCCAGAACCT GGCCGACGAGCTGAGCGAGGAGCTGAACGGCTGGAGAATG AGAACCAAGGAGGCGAGAACAAGGTGGAGCACGCCAGCA GCGAGAAGAGCGAGATGCTGGAGAGAGAATCGTGCACCTGGA GACCGAGATGGAGAAGCTGAGCACCAGCGAGATCGCCGCC GACTACTGCAGCACCAAGATGACCGAGAGAAGAAGGAGA TCGAGCTGGCCAAGTACAGAGAGGACTTCGAGAACGCCGC CATCGTGGGCCTGGAGAGAATCAGCAAGGAGATCAGCGAG CTGACCAAGAGACCCTGAAGGCCAAGATCATCCCCAGCA ACATCAGCAGCATCCAGCTGGTGTGCGACGAGCTGTGCAG AAGACTGAGCAGAGAGAGAGAGAGCAGCAGCACGAGTACGCC AAGGTGATGAGAGACAGTGAACGAGAAGATCGAGAGCTGC AGCTGGAGAAGGACGCCCTGGAGCACGAGCTGAAGATGAT GAGCAGCAACAACGAGAACGTGCCCCCCGTGGGCACCAGC GTGAGCGGCATGCCCACCAAGACCAGCAACCAGAAGTGCG CCCAGCCCCACTACACCAGCCCCACCAGACAGCTGCTGCA CGAGAGCACCATGGCCGTGGACGCCATCGTGCAGAAGCTG AAGAAGACCCACAACATGAGCGGCATGGGCCCCGAGCTGA AGGAGACCATCGGCAACGTGATCAACGAGAGCAGAGTGCT GAGAGACTTCCTGCACCAGAAGCTGATCCTGTTCAAGGGC ATCGACATGAGCAACTGGAAGAACGAGACCGTGGACCAGC TGATCACCGACCTGGGCCAGCTGCACCAGGACAACCTGAT GCTGGAGGAGCAGATCAAGAAGTACAAGAAGGAGCTGAAG CTGACCAAGAGCGCCATCCCCACCCTGGGCGTGGAGTTCC AGGACAGAATCAAGACCGAGATCGGCAAGATCGCCACCGA CATGGGCGGCGCCGTGAAGGAGATCAGAAAGAAGGGTACC GAGCAGAAGCTGATCTCAGAGGAGGACCTGGGCGCCCCCG GCTCCGCCGGCTCCGCCGCCGGCTCCGGCGCTTCTAACTT TACTCAGTTCGTTCTCGTCGACAATGGCGGAACTGGCGAC GTGACTGTCGCCCCAAGCAACTTCGCTAACGGGATCGCTG AATGGATCAGCTCTAACTCGCGTTCACAGGCTTACAAAGT AACCTGTAGCGTTCGTCAGAGCTCTGCGCAGAATCGCAAA TACACCATCAAAGTCGAGGTGCCTAAAGGCGCCTGGCGTT CGTACTTAAATATGGAACTAACCATTCCAATTTTCGCCAC GAATTCCGACTGCGAGCTTATTGTTAAGGCAATGCAAGGT CTCCTAAAAGATGGAAACCCGATTCCCTCAGCAATCGCAG CAAACTCCGGCATCTACTAA (SEQ ID NO: 45)

[0399] protein: MEDNSVLNEDSNLEHVEGQPRRSMSQPVLNVEGDKRTSST SATQQQVLSGAFSSADVRSIPIIQTWEENKALKTKITILR GELQMYQRRYSEAKEASQKRVKEVMDDYVDLKLGQENVQE KMEQYKLMEEEDLLAMQSRIETSEDNFARQMKEFEAQKHAM EERIKELELSATDANNTTVGSFRGTLDDILKKNDPDFTLT SGYEERKINDLEAKLLSEIDKVAELEDHIQQLRQELDDQS ARLADSENVRAQLEAATGQGILGAAGNAMVPNSTFMIGNG RESQTRDQLNYIDDLETKLADAKKENDKARQALVEYMNKC SKLEHEIRTMVKNSTFDSSSMLLGGQTSDELKIQIGKVNG ELNVLRAENRELRIRCDQLTGGDGNLSISLGQSRLMAGIA TNDVDSIGQGNETGGTSMRILPRESQLDDLEESKLPLMDT SSAVRNQQQFASMWEDFESVKDSLQNNHNDTLEGSFNSSM PPPGRDATQSFLSQKSFKNSPIVMQKPKSLHLHLKSHQSE GAGEQIQNNSFSTKTASPHVSQSHIPILHDMQQILDSSAM FLEGQHDVAVNVEQMQEKMSQIREALARLFERLKSSAALF EEILERMGSSDPNADKIKKMKLAFETSINDKLNVSAILEA AEKDLHNMSLNFSILEKSIVSQAAEASRRFTIAPDAEDVA SSSLLNASYSPLFKFTSNDIVEKLQNEVSELKNELEMAR TRDMRSPLNGSSGRLSDVQINTNRMFEDLEVSEATLQKAK EENSTLKSQFAELEANLHQVNSKLGEVRCELNEALARVDG EQETRVKAENALEEARQLISSLKHEENELKKTITDMGMRL NEAKKSDEFLKSELSTALEEEEKKSQNLADELSEELNGWRM RTKEAENKVEHASSEXEMLERIVHLETEMEKLSSEIAA DYCSTKMTERKKIELAKYREDFENAAIVGLERSKEISE LTKKTLKAKIIPSNISSIQLVCDELCRRLSREREQQHEYA KVMRDVNEKIEKLQLEKDALEHELKMMSSNENVPPVGTS VSGMPTKTSNQKCAQPHYTSPTRQLLHESTMAVDAIVQKL KKTHNMSGMGPELKETIGNVINESRVLRDFLHQKLILFKG IDMSNWKNETVDQLITDLGQLHQDNLMLEEQIKKKYLK LTKSAIPTLGVEFQDRIKTEIGKIATDMGGAVKEIRKKGT EQKLISEEDLGAPGSAGSAAGSGASNFTQFVLVDNGGTGD VTVAPSNFANGIAEWISSNSRSQAYKVTCSVRQSSAQNRK YTIKVEVPKGAWRSYLNMELTIPIFATNSDCELIVKAMQG LLKDGNPIPSAIAANSGIY (SEQ ID NO: 46)

[0400] SPD5-PylRS AF

[0401] DNA: ATGGAGGACAACAGCGTGCTGAACGAGGACAGCAACCTGG AGCACGTGGAGGGCCAGCCCAGAAGAAGCATGAGCCAGCC CGTGCTGAACGTGGAGGGCGACAAGAGAACCAGCAGCACC AGCGCCACCCAGCAGCAGGTGCTGAGCGGCGCCTTCAGCA GCGCCGACGTGAGAAGCATCCCCATCATCCAGACCTGGGA GGAGAACAAGGCCCTGAAGACCAAGATCACCATCCTGAGA GGCGAGCTGCAGATGTACCAGAGAAGATACAGCGAGGCCA AGGAGGCCAGCCAGAAGAGAGTGAAGGAGGTGATGGACGA CTACGTGGACCTGAAGCTGGGCCAGGAGAACGTGCAGGAG AAGATGGAGCAGTACAAGCTGATGGAGGAGGACCTGCTGG CCATGCAGAGCAGAATCGAGACCAGCGAGGACAACTTCGC CAGACAGATGAAGGAGTTCGAGGCCCAGAAGCACGCCATG GAGGAGAGAATCAAGGAGCTGGAGCTGAGCGCCACCGACG CCAACAACACCACCGTGGGCAGCTTCAGAGGCACCCTGGA CGACATCCTGAAGAAGAACGACCCCGACTTCACCCTGACC AGCGGCTACGAGGAGAGAAAGATCAACGACCTGGAGGCCA AGCTGCTGAGCGAGATCGACAAGGTGGCCGAGCTGGAGGA CCACATCCAGCAGCTGAGACAGGAGCTGGACGACCAGAGC GCCAGACTGGCCGACAGCGAGAACGTGAGAGCCCAGCTGG AGGCCGCCACCGGCCAGGGCATCCTGGGCGCCGCCGGCAA CGCCATGGTGCCCAACAGCACCTTCATGATCGGCAACGGC AGAGAGAGCCAGACCAGAGACCAGCTGAACTACATCGACG ACCTGGAGACCAAGCTGGCCGACGCCAAGAAGGAGAACGA CAAGGCCAGACAGGCCCTGGTGGAGTACATGAACAAGTGC AGCAAGCTGGAGCACGAGATCAGAACCATGGTGAAGAACA GCACCTTCGACAGCAGCAGCATGCTGCTGGGCGGCCAGAC CAGCGACGAGCTGAAGATCCAGATCGGCAAGGTGAACGGC GAGCTGAACGTGCTGAGAGCCGAGAACAGAGAGCTGAGAA TCAGATGCGACCAGCTGACCGGCGGCGACGGCAACCTGAG CATCAGCCTGGGCCAGAGCAGACTGATGGCCGGCATCGCC ACCAACGACGTGGACAGCATCGGCCAGGGCAACGAGACCG GCGGCACCAGCATGAGAATCCTGCCCAGAGAGAGCCAGCT GGACGACCTGGAGGAGAGCAAGCTGCCCCTGATGGACACC AGCAGCGCCGTGAGAAACCAGCAGCAGTTCGCCAGCATGT GGGAGGACTTCGAGAGCGTGAAGGACAGCCTGCAGAACAA CCACAACGACACCCTGGAGGGCAGCTTCAACAGCAGCATG CCCCCCCCGGCAGAGACGCCACCCAGAGCTTCCTGAGCC AGAAGAGCTTCAAGAACAGCCCCCATCGTGATGCAGAGCC CAAGAGCCTGCACCTGCACCTGAAGAGCCACCAGAGCGAG GGCGCCGGCGAGCAGATCCAGAACAACAGCTTCAGCACCA AGACCGCCAGCCCCACGTGAGCCAGAGCCACATCCCCAT CCTGCACGACATGCAGCAGATCCTGGACAGCAGCGCCATG TTCCTGGAGGGCCAGCACGACGTGGCCGTGAACGTGGAGC AGATGCAGGAGAGATGAGCCAGATCAGAGAGGCCCTGGC CAGACTGTTCGAGAGACTGAAGAGCAGCGCGCCCTGTTC GAGGAGATCCTGGAGAGAATGGGCAGCAGCGACCCCAACG CCGACAAGATCAAGAAGATGAAGCTGGCCTTCGAGACCAG CATCAACGACAAGCTGAACGTGAGCGCCATCCTGGAGGCC GCCGAGAAGGACCTGCACAACATGAGCCTGAACTTCAGCA TCCTGGAGAAGAGCATCGTGAGCCAGGCCGCCGAGGCCAG CAGAAGATTCACCATCGCCCCGACGCCGAGGACGTGGCC AGCAGCAGCCTGCTGAACGCCAGCTACAGCCCCCTGTTCA AGTTCACCAGCAACAGCGACATCGTGGAGAAGCTGCAGAA CGAGGTGAGCGAGCTGAAGAACGAGCTGGAGATGGCCAGA ACCAGAGACATGAGAAGCCCCCTGAACGGCAGCAGCAGCGGCA GACTGAGCGACGTGCAGATCAAACACCAACAGAATGTTCGA GGACCTGGAGGTGAGCGAGGCCACCCTGCAGAAGGCCAAG GAGGAGAACAGCACCCTGAAGAGCCAGTTCGCCGAGCTGG AGGCCAACCTGCACCAGGTGAACAGCAAGCTGGGCGAGGT GAGATGCGAGCTGAACGAGGCCCTGGCCAGAGTGGACGGC GAGCAGGAGACCAGAGTGAAGGCCGAGAACGCCCTGGAGG AGGCCAGACAGCTGATCAGCAGCCTGAAGCACGAGGAAA CGAGCTGAAGAGACCATCACCGACATGGGCATGAGACTG AACGAGGCCAAGAAGAGCGACGAGTTCCTGAAGAGCGAGC TGAGCACCCGCCCTGGAGGAGGAGAAGAGAGCCCAGAACCT GGCCGACGAGCTGAGCGAGGAGCTGAACGGCTGGAGAATG AGAACCAAGGAGGCGAGAACAAGGTGGAGCACGCCAGCA GCGAGAAGAGCGAGATGCTGGAGAGAGAATCGTGCACCTGGA GACCGAGATGGAGAAGCTGAGCACCAGCGAGATCGCCGCC GACTACTGCAGCACCAAGATGACCGAGAGAAGAAGGAGA TCGAGCTGGCCAAGTACAGAGAGGACTTCGAGAACGCCGC CATCGTGGGCCTGGAGAGAATCAGCAAGGAGATCAGCGAG CTGACCAAGAGACCCTGAAGGCCAAGATCATCCCCAGCA ACATCAGCAGCATCCAGCTGGTGTGCGACGAGCTGTGCAG AAGACTGAGCAGAGAGAGAGAGAGCAGCAGCACGAGTACGCC AAGGTGATGAGAGACAGTGAACGAGAAGATCGAGAGCTGC AGCTGGAGAAGGACGCCCTGGAGCACGAGCTGAAGATGAT GAGCAGCAACAACGAGAACGTGCCCCCCGTGGGCACCAGC GTGAGCGGCATGCCCACCAAGACCAGCAACCAGAAGTGCG CCCAGCCCCACTACACCAGCCCCACCAGACAGCTGCTGCA CGAGAGCACCATGGCCGTGGACGCCATCGTGCAGAAGCTG AAGAAGACCCCACAACATGAGCGGCATGGGCCCCGAGCTGA AGGAGACCATCGGCAACGTGATCAACGAGAGCAGAGTGCT GAGAGACTTCCTGCACCAGAAGCTGATCCTTGTTCAAGGGC ATCGACATGAGCAACTGGAAGAACGAGCGTGGACCAGC TGATCACCGACCTGGGCCAGCTGCACCAGGCAACCTGAT GCTGGAGGAGCAGATCAAGAAGTACAAGAAGGAGCTGAG CTGACCAAGAGCGCCATCCCCACCCTGGGCGTGGAGTTCC AGGACAGAATCAAGACCGAGATCGGCAAGATCGCCACCGA CATGGGCGGCGCCGTGAAGGAGATCAGAAAAGAAGGGTACC GGCGCCCCCGGCTCCGCCGGCTCCCGCCGCCGGCTCCGGCG CGTGCCCGGTGCCGCTGCAGCTGCCGCCGCTGGAACGCCT GACCCTGGATGATAAAAAACCGCTGAATACCCTGATCCT GCTACTGGTCTGTGGATGAGTCGTACCGGAACCATTCATA AAATCAAACACCACGAGGTTAGCCGTTCGAAAATCTATAT TGAGATGGCGTGTGGCGATCATCTGGTTGTGAACAATAGC CGCTCTTCTCGTACAGCACGTGCACTGCGTCACCACAAAT ATCGTAAAACCTGTAAACGTTGCCGTGTGTCCGATGAGGA TCTGAACAAATTCCTGACAAAAGCCAATGAGGACCAAACA AGCGTGAAAGTGAAAGTCGTTAGCGCTCCTACCCGTACTA AAAAAGCAATGCCGAAATCCGTTGCTCGTGCCCCTAAACC ACTGGAAAACACTGAAGCAGCACAGGCACAGCCGTCTGGA AGCAAATTCTCTCCGGCCATTCCTGTTTCTACCCAGGAGT CCGTTTCTGTTCCAGCAAGTGTGAGCACCAGCATTAGCAG TATTAGCACCGGTGCCACCGCTAGCGCCCTGGTTAAAGGC AATACCAATCCGATTACAAGCATGTCTGCCCGGTTCAAG CATCAGCTCCAGCACTGACAAAATCCCAAACCGATCGTCT GGAGGTTCTGCTGAATCCGAAAGACGAAATCAGCCTGAAT TCCGGCAAACCGTTTCGTGAACTGGAGAGCGAACTGCTGT CACGTCGTAAAAAAGACCTGCAACAAATCTATGCCGAAGA ACGTGAGAACTATCTGGGGAAACTGGAACGTGAAATCACC CGCTTTTTCGTGGATCGTGGCTTTCTGGAGATCAAATCCC CGATTCTGATTCCTCTGGAGTATATCGAGCGTATGGGCAT CGACAATGATACCGAACTGAGCAAACAAATTTTCCGTGTG GATAAAAACTTCTGTCTGCGCCCTATGCTAGCACCAAATC TGGCTAACTATCTGCGCAAACTGGACCGTGCCCTGCCTGA TCCTATCAAAATCTTCGAGATCGGCCCGTGTTATCGTAAA GAGTCCGACGGTAAAGAACATCTGGAGGAGTTTACCATGC TGAACTTTTGCCAAATGGGTTCAGGTTGTACTCGTGAGAA CCTGGAAAGCATCATCACCGATTTTCTGAACCACCTGGGC ATTGACTTCAAAATTGTGGGCGACAGCTGTATGGTGTTTG GCGACACCCTGGATGTCATGCACGGCGACCTGGAACTGTC TAGTGCCGTTGTGGGCCCAATCCCGCTGGATCGTGAGTGG GGTATCGACAAACCTTGGATCGGTGCGGGTTTTGGTCTGG AGCGTCTGCTGAAAGTAAAACACGACTTCAAGAACATCAA ACGTGCTGCACGTTCCGAGTCCTATTACAATGGTATTTCT ACTAACCTGTAA (SEQ ID NO: 47)

[0402] protein: MEDNSVLNEDSNLEHVEGQPRRSMSQPVNLVEGDKRTSST SATQQQVLSGAFSSADVRSIPIIQTWEENKALKTKITILR GELQMYQRRRYSEAKEASQKRVKEVMDDDYVDLKLGQENVQE KMEQYKLMEEDLLAMQSRIETSEDNFARQMKEFEAQKHAM EERIKEELSATDANNTTVGSFRGTLDDILKNDPDFTLT SGYEERKINDLEAKLLSEIDKVAELEDHIQQLRQELDDQS ARLADSENVRAQLEAATGQGILGAAGNAMVPNSTFMIGNG RESQTRDQLNYIDDLETKLADAKKENDKARQALVEYMNKC SKLEHEIRTMVKNSTFDSSSMLLGGQTSDELKIQIGKVNG ELNVLRAENRELRIRCDQLTGGDGNLSISLGQSRLMAGIA TNDVDSIGQGNETGGTSMRILPRESQLDDLEESKLPLMDT SSAVRNQQQFASMWEDFESVKDSLQNNNHNDTLEGSFNSSM PPPGRDATQSFLSQKSFKNSPIVMQKPKSLHLHLKSHQSE GAGEQIQNNSFSTKTASPHVSQSHIPILHDMQQILDSSAM FLEGQHDVAVNVEQMQEKMSQIREALARLFERLKSSAALF EEILERMGSSDPNADKIKKMKLAFETSINDKLNVSAILEA AEKDLHNMSLNFSILEKSIVSQAAEASRRFTIAPDAEDVA SSSLLNASYSPLFKFTSNDIVEKLQNEVSELKNELEMAR TRDMRSPLNGSSGRLSDVQINTNRMFEDLEVSEATLQKAK EENSTLKSQFAELEANLHQVNSKLGEVRCELNEALARVDG EQETRVKAENALEEARQLISSLKHEENELKKTITDMGMRL NEAKKSDEFLKSELSTALEEEEKKSQNLADELSEELNGWRM RTKEAENKVEHASSEXEMLERIVHLETEMEKLSSEIAA DYCSTKMTERKKIELAKYREDFENAAIVGLERSKEISE LTKKTLKAKIIPSNISSIQLVCDELCRRLSREREQQHEYA KVMRDVNEKIEKLQLEKDALEHELKMMSSNENVPPVGTS VSGMPTKTSNQKCAQPHYTSPTRQLLHESTMAVDAIVQKL KKTHNMSGMGPELKETIGNVINESRVLRDFLHQKLILFKG IDMSNWKNETVDQLITDLGQLHQDNLMLEEQIKKKYLK LTKSAIPTLGVEFQDRIKTEIGKIATDMGGAVKEIRKKGT GAPGSAGSAAGSGACPVPLQLPPLERLTLDDKKPLNTLIS ATGLWMSRTGTIHKIKHHEVSRSKIYIEMACGDHLVVNNS RSSRTARALRHHKYRKTCKRCRVSDEDLDNKFLTKANEDQT SVKVKVVSAPTRTKKAMPKSVARAKPPLENTEAAQAQPSG SKFSPAIPVSTQESVSVPASVSTSISISTGATASALVKG NTNPITSMSAPVQASAPALTKSQTDRLEVLLNPKDEISLN SGKPFRELESELLSRRKKDLQQIYAEERENYLGKLEREIT RFFVDRGFLEIKSPILIPLEYIERMGIDNDTELSKQIFRV DKNFCLRPMLAPNLANYLRKLDRALPDPIKIFEIGPCYRK ESDGKEHLEEFTMLNFCQMGSGCTRENLESIITDFLNHLG IDFKIVGDSCMVFGDTLDVMHGDLELSSAVVGPIPLDREW GIDKPWIGAGFGLERLLKVKHDFKNIKRAARSESYYNGIS TNL

[0403] KIF16B-FUS-PylRS AF

[0404] DNA: ATGGCATCGGTCAAGGTGGCCGTGAGGGTCCGGCCATGA ATCGCAGGGAAAAGGACTTGGAGGCCAAGTTCATTATTCA GATGGAGAAAAGCAAAACGACAATCACAAACTTAAAGATA CCAGAAGGAGGCACTGGGGACTCAGGAAGAGAACGGACCA AGACCTTCACCTATGACTTTTCTTTTATTCTGCTGATAC AAAAAGCCCAGATTACGTTTCACAAGAAATGGTTTTCAAA ACCCTCGGCACAGATGTCGTGAAGTCTGCATTTGAAGGTT ATAATGCTTGTGTCTTTGCATATGGGCAAACTGGATCTGG AAAGTCATACACTATGATGGGAAATTCTGGAGATTCTGGC TTAATACCTCGGATCTGTGAAGGACTCTTCAGTCGGATAA ATGAAACCACCAGATGGGATGAAGCTTCTTTTCGAACTGA AGTCAGCTACTTAGAAATTTATAACGAACGTGTGAGAGAT CTACTTCGGCGGAAGTCATCTAAAACCTTCAATTTGAGAG TCCGTGAGCATCCCAAAGAAGGCCCTTATGTTGAGGATTT ATCCAAACATTTAGTACAGAATTATGGTGACGTAGAAGAA CTTATGGATGCGGGCAATATCAACCGGACCACCGCAGCGA CTGGGATGAACGACGTCAGTAGCAGGTCTCATGCCATCTT CACCATCAAGTTCACTCAGGCTAAATTTGATTCTGAAATG CCATGTGAAACCGTCAGTAAGATCCACTTGGTTGATCTTG CCGGAAGTGAGCGTGCAGATGCCACCGGAGCCACCGGGGT TAGGCTAAAGGAAGGGGGAAATATTAACAAGTCCCTCGTG ACTCTGGGGAACGTCATTTCTGCCTTAGCTGATTTATCTC AGGATGCTGCAAATACTCTTGCAAAGAAGAAGCAAGTTTT CGTGCCTTACAGGGATTCTGTGTTGACTTGGTTGTTAAAA GATAGCCTTGGAGGAAACTCTAAAACTATCATGATTGCCA CCATTTCACCTGCTGATGTCAATTATGGAGAAACCCTAAG TACTCTTCGCTATGCAAATAGAGCCAAAAACATCATCAAC AAGCCTACCATTAATGAGGATGCCAACGTCAAACTTATCC GTGAGCTGCGAGCTGAAATAGCCAGACTGAAAACGCTGCT TGCTCAAGGGAATCAGATTGCCCTCTTAGACTCCCCCACA TATACAGATATTGAAATGAACAGATTGGGAAAGGGCGCCC CCGGCTCCCGCCGGCTCCGCCGCCGGCTCCGGCATGGCCTC AAACGATTATACCCAACAAGCAACCCAAAGCTATGGGGCC TACCCCACCCAGCCCGGGCAGGGCTATTCCCAGCAGAGCA GTCAGCCCTACGGACAGCAGAGTTACAGTGGTTATAGCCA GTCCACGGACACTTCAGGATATGGGCCAGAGCAGCTATTCT TCTTATGGCCAGAGCCAGAACACAGGCTATGGAACTCAGT CAACTCCCCAGGGATATGGCTCGACTGGCGGCTATGGCAG TAGCCAGAGCTCCCAATCGTCTTACGGGCAGCAGTCCTCC TACCCTGGCTATGGCCAGCAGCCAGCTCCCAGCAGCACCT CGGGAAGTTACGGTAGCAGTTCTCAGAGCAGCAGCTATGG GCAGCCCCAGAGTGGGAGCTACAGCCAGCAGCCTAGCTAT GGTGGACAGCAGCAAAGCTATGGACAGCAGCAAAGCTATA ATCCCCCTCAGGGCTATGGACAGCAGAACCAGTACAACAG CAGCAGTGGTGGTGGAGGTGGAGGTGGAGGTGGAGGTAAC TATGGCCAAGATCAATCCTCCATGAGTAGTGGTGGTGGCA GTGGTGGCGGTTATGGCAATCAAGACCAGAGTGGTGGAGG TGGCAGCGGTGGCTATGGACAGCAGGACCGTGGAGGCCGC GGCAGGGGTGGCAGTGGTGGCGGCGGCGGCGGCGGCGGTG GTGGTTACAACCGCAGCAGTGGTGGCTATGAACCCAGAGG TCGTGGAGGTGGCCGTGGAGGCAGAGGTGGCATGGGCGGA AGTGACCGTGGTGGCTTCAATAAATTTGGTGGCCCTCGGG ACCAAGGATCACGTCATGACTCCGAACAGGATAATTCAGA CAACAACACCATCTTTGTGCAAGGCCTGGGTGAGAATGTT ACAATTGAGTCTGTGGCTGATTACTTCAAGCAGATTGGTA TTATTAAGACAAACAAGAAAACGGGACAGCCCATGATTAA TTTGTACACAGACAGGGAAACTGGCAAGCTGAAGGGAGAG GCAACGGTCTCTTTTGATGACCCACCTTCAGCTAAAGCAG CTATTGACTGGTTTGATGGTAAAGAATTCTCCGGAAATCC TATCAAGGTCTCATTTGCTACTCGCCGGGCAGACTTTAAT CGGGGTGGTGGCAATGGTCGTGGAGGCCGAGGGCGAGGAG GACCCATGGGCCGTGGAGGCTATGGAGGTGGTGGCAGTGG TGGTGGTGGCCGAGGAGGATTTCCCAGTGGAGGTGGTGGC GGTGGAGGACAGCAGCGAGCTGGTGACTGGAAGTGTCCTA ATCCCACCTGTGAGAATATGAACTTCTCTTGGAGGAATGA ATGCAACCAGTGTAAGGCCCCTAAACCAGATGGCCCAGGA GGGGGACCAGGTGGCTCTCACATGGGGGGTAACTACGGGG ATGATCGTCGTGGTGGCAGAGGAGGCGATTACAAGGATGA CGACGATAAGGGTACCGGCGCCCCCGGCTCCGCCGGCTCC GCCGCCGGCTCCGGCGCGTGCCCGGTGCCGCTGCAGCTGC CGCCGCTGGAACGCCTGACCCTGGATGATAAAAAACCGCT GAATACCCTGATCTCTGCTACTGGTCTGTGGATGAGTCGT ACCGGAACCATTCATAAAATCAAACACCACGAGGTTAGCC GTTCGAAAATCTATATTGAGATGGCGTGTGGCGATCATCT GGTTGTGAACAATAGCCGCTCTTCTCGTACAGCACGTGCA CTGCGTCACCACAAATATCGTAAAACCTGTAAACGTTGCC GTGTGTCCGATGAGGATCTGAACAAATTCCTGACAAAAGC CAATGAGGACCAAACAAGCGTGAAAGTGAAAGTCGTTAGC GCTCCTACCCGTACTAAAAAAGCAATGCCGAAATCCGTTG CTCGTGCCCCTAAACCACTGGAAAACACTGAAGCAGCACA GGCACAGCCGTCTGGAAGCAAATTCTCTCCGGCCATTCCT GTTTCTACCCAGGAGTCCGTTTCTGTTCCAGCAAGTGTGA GCACCAGCATTAGCAGTATTAGCACCGGTGCCACCGCTAG CGCCCTGGTTAAAGGCAATACCAATCCGATTACAAGCATG TCTGCCCCGGTTCAAGCATCAGCTCCAGCACTGACAAAAT CCCAAACCGATCGTCTGGAGGTTCTGCTGAATCCGAAAGA CGAAATCAGCCTGAATTCCGGCAAACCGTTTCGTGAACTG GAGAGCGAACTGCTGTCACGTCGTAAAAAAGACCTGCAAC AAATCTATGCCGAAGAACGTGAGAACTATCTGGGGAAACT GGAACGTGAAATCACCCGCTTTTTCGTGGATCGTGGCTTT CTGGAGATCAAATCCCCGATTCTGATTCCTCTGGAGTATA TCGAGCGTATGGGCATCGACAATGATACCGAACTGAGCAA ACAAATTTTCCGTGTGGATAAAAACTTCTGTCTGCGCCCT ATGCTAGCACCAAATCTGGCTAACTATCTGCGCAAACTGG ACCGTGCCCTGCCTGATCCTATCAAAATCTTCGAGATCGG CCCGTGTTATCGTAAAGAGTCCGACGGTAAAGAACATCTG GAGGAGTTTACCATGCTGAACTTTTGCCAAATGGGTTCAG GTTGTACTCGTGAGAACCTGGAAAGCATCATCACCGATTT TCTGAACCACCTGGGCATTGACTTCAAAATTGTGGGCGAC AGCTGTATGGTGTTTGGCGACACCCTGGATGTCATGCACG GCGACCTGGAACTGTCTAGTGCCGTTGTGGGCCCAATCCC GCTGGATCGTGAGTGGGGTATCGACAAACCTTGGATCGGT GCGGGTTTTGGTCTGGAGCGTCTGCTGAAAGTAAAACACG ACTTCAAGAACATCAAACGTGCTGCACGTTCCGAGTCCTA TTACAATGGTATTTCTACTAACCTGTAA (SEQ ID NO: 49)

[0405] protein: MASVKVAVRVRPMNRREKDLEAKFIIQMEKSKTTITNLKI PEGGTGDSGRERTKTFTYDFSFYSADTKSPDYVSQEMVFK TLGTDVVKSAFEGYNACVFAYGQTGSGKSYTMMGNSGDSG LIPRICEGLFSRINETTRWDEASFRTEVSYLEIYNERVRD LLRRKSSKTFNLRVREHPKEGPYVEDLSKHLVQNYGDVEE LMDAGNINRTTAATGMNDVSSRSHAIFTIKFTQAKFDSEM PCETVSKIHLVDLAGSERADATGATGVRLKEGGNINKSLV TLGNVISALADLSQDAANTLAKKKQVFVPYRDSVLTWLLK DSLGGNSKTIMIATISPADVNYGETLSTLRYANRAKNIIN KPTINEDANVKLIRELRAEIARLKTLLAQGNQIALLDSPT YTDIEMNRLGKGAPGSAGSAAGSGMASNDYTQQATQSYGA YPTQPGQGYSQQSSQPYGQQSYSGYSQSTDTSGYGQSSYS SYGQSQNTGYGTQSTPQGYGSTGGYGSSQSSQSSYGQQSS YPGYGQQPAPSSTSGSYGSSSQSSSYGQPQSGSYSQQPSY GGQQQSYGQQQSYNPPQGYGQQNQYNSSSGGGGGGGGGN YGQDQSSMSSGGGSGGGYGNQDQSGGGGSGGYGQQDRGGR GRGGSGGGGGGGGGGYNRSSGGYEPRGRGGGRGGRGGMGG SDRGGFNKFGGPRDQGSRHDSEQDNSDNNTIFVQGLGENV TIESVADYFKQIGIIKTNKKTGQPMINLYTDRETGKLKGE ATVSFDDPPSAKAAIDWFDGKEFSGNPIKVSFATRRADFN RGGGNGRGGRGRGGPMGRGGYGGGGSGGGGRGGFPSGGGG GGGQQRAGDWKCPNPTCENMNFSWRNECNQCKAPKPDGPG GGPGGSHMGGNYGDDRRGGRGGDYKDDDDKGTGAPGSAGS AAGSGACPVPLQLPPLERLTLDDKKPLNTLISATGLWMSR TGTIHKIKHHEVSRSKIYIEMACGDHLVVNNSRSSRTARA LRHHKYRKTCKRCRVSDEDLNKFLTKANEDQTSVKVKVVS APTRTKKAMPKSVARAPKPLENTEAAQAQPSGSKFSPAIP VSTQESVSVPASVSTSISSISTGATASALVKGNTNPITSM SAPVQASAPALTKSQTDRLEVLLNPKDEISLNSGKPFREL ESELLSRRKKDLQQIYAEERENYLGKLEREITRFFVDRGF LEIKSPILIPLEYIERMGIDNDTELSKQIFRVDKNFCLRP MLAPNLANYLRKLDRALPDPIKIFEIGPCYRKESDGKEHL EEFTMLNFCQMGSGCTRENLESIITDFLNHLGIDFKIVGD SCMVFGDTLDVMHGDLELSSAVVGPIPLDREWGIDKPWIG AGFGLERLLKVKHDFKNIKRAARSESYYNGISTNL (SEQ ID NO: 50)

[0406] KIF16B-VSV-G-FUS-PylRS AF

[0407] DNA: ATGGCATCGGTCAAGGTGGCCGTGAGGGTCCGGCCCATGA ATCGCAGGGAAAAGGACTTGGAGGCCAAGTTCATTATTCA GATGGAGAAAAGCAAAACGACAATCACAAACTTAAAGATA CCAGAAGGAGGCACTGGGGACTCAGGAAGAGAACGGACCA AGACCTTCACCTATGACTTTTCTTTTTATTCTGCTGATAC AAAAAGCCCAGATTACGTTTCACAAGAAATGGTTTTCAAA ACCCTCGGCACAGATGTCGTGAAGTCTGCATTTGAAGGTT ATAATGCTTGTGTCTTTGCATATGGGCAAACTGGATCTGG AAAGTCATACACTATGATGGGAAATTCTGGAGATTCTGGC TTAATACCTCGGATCTGTGAAGGACTCTTCAGTCGGATAA ATGAAACCACCAGATGGGATGAAGCTTCTTTTCGAACTGA AGTCAGCTACTTAGAAATTTATAACGAACGTGTGAGAGAT CTACTTCGGCGGAAGTCATCTAAAACCTTCAATTTGAGAG TCCGTGAGCATCCCAAAGAAGGCCCTTATGTTGAGGATTT ATCCAAACATTTAGTACAGAATTATGGTGACGTAGAAGAA CTTATGGATGCGGGCAATATCAACCGGACCACCGCAGCGA CTGGGATGAACGACGTCAGTAGCAGGTCTCATGCCATCTT CACCATCAAGTTCACTCAGGCTAAATTTGATTCTGAAATG CCATGTGAAACCGTCAGTAAGATCCACTTGGTTGATCTTG CCGGAAGTGAGCGTGCAGATGCCACCGGAGCCACCGGGGT TAGGCTAAAGGAAGGGGGAAATATTAACAAGTCCCTCGTG ACTCTGGGGAACGTCATTTCTGCCTTAGCTGATTTATCTC AGGATGCTGCAAATACTCTTGCAAAGAAGAAGCAAGTTTT CGTGCCTTACAGGGATTCTGTGTTGACTTGGTTGTTAAAA GATAGCCTTGGAGGAAACTCTAAAACTATCATGATTGCCA CCATTTCACCTGCTGATGTCAATTATGGAGAAACCCTAAG TACTCTTCGCTATGCAAATAGAGCCAAAAACATCATCAAC AAGCCTACCATTAATGAGGATGCCAACGTCAAACTTATCC GTGAGCTGCGAGCTGAAATAGCCAGACTGAAAACGCTGCT TGCTCAAGGGAATCAGATTGCCCTCTTAGACTCCCCCACA TATACAGATATTGAAATGAACAGATTGGGAAAGGGCGCCC CCGGCTCCGCCGGCTCCGCCGCCGGCTCCGGCATGGCCTC AAACGATTATACCCAACAAGCAACCCAAAGCTATGGGGCC TACCCCACCCAGCCCGGGCAGGGCTATTCCCAGCAGAGCA GTCAGCCCTACGGACAGCAGAGTTACAGTGGTTATAGCCA GTCCACGGACACTTCAGGATATGGCCAGAGCAGCTATTCT TCTTATGGCCAGAGCCAGAACACAGGCTATGGAACTCAGT CAACTCCCCAGGGATATGGCTCGACTGGCGGCTATGGCAG TAGCCAGAGCTCCCAATCGTCTTACGGGCAGCAGTCCTCC TACCCTGGCTATGGCCAGCAGCCAGCTCCCAGCAGCACCT CGGGAAGTTACGGTAGCAGTTCTCAGAGCAGCAGCTATGG GCAGCCCCAGAGTGGGAGCTACAGCCAGCAGCCTAGCTAT GGTGGACAGCAGCAAAGCTATGGACAGCAGCAAAGCTATA ATCCCCCTCAGGGCTATGGACAGCAGAACCAGTACAACAG CAGCAGTGGTGGTGGAGGTGGAGGTGGAGGTGGAGGTAAC TATGGCCAAGATCAATCCTCCATGAGTAGTGGTGGTGGCA GTGGTGGCGGTTATGGCAATCAAGACCAGAGTGGTGGAGG TGGCAGCGGTGGCTATGGACAGCAGGACCGTGGAGGCCGC GGCAGGGGTGGCAGTGGTGGCGGCGGCGGCGGCGGCGGTG GTGGTTACAACCGCAGCAGTGGTGGCTATGAACCCAGAGG TCGTGGAGGTGGCCGTGGAGGCAGAGGTGGCATGGGCGGA AGTGACCGTGGTGGCTTCAATAAATTTGGTGGCCCTCGGG ACCAAGGATCACGTCATGACTCCGAACAGGATAATTCAGA CAACAACACCATCTTTGTGCAAGGCCTGGGTGAGAATGTT ACAATTGAGTCTGTGGCTGATTACTTCAAGCAGATTGGTA TTATTAAGACAAACAAGAAAACGGGACAGCCCATGATTAA TTTGTACACAGACAGGGAAACTGGCAAGCTGAAGGGAGAG GCAACGGTCTCTTTTGATGACCCACCTTCAGCTAAAGCAG CTATTGACTGGTTTGATGGTAAAGAATTCTCCGGAAATCC TATCAAGGTCTCATTTGCTACTCGCCGGGCAGACTTTAAT CGGGGTGGTGGCAATGGTCGTGGAGGCCGAGGGCGAGGAG GACCCATGGGCCGTGGAGGCTATGGAGGTGGTGGCAGTGG TGGTGGTGGCCGAGGAGGATTTCCCAGTGGAGGTGGTGGC GGTGGAGGACAGCAGCGAGCTGGTGACTGGAAGTGTCCTA ATCCCACCTGTGAGAATATGAACTTCTCTTGGAGGAATGA ATGCAACCAGTGTAAGGCCCCTAAACCAGATGGCCCAGGA GGGGGACCAGGTGGCTCTCACATGGGGGGTAACTACGGGG ATGATCGTCGTGGTGGCAGAGGTGGTGCGATCGCAGGAGC ACCAGGAAGTGCTGGTTCTGCTGCTGGTAGTGGAGCGTGC CCGGTGCCGCTGCAGCTGCCGCCGCTGGAACGCCTGACCC TGGATGATAAAAAACCGCTGAATACCCTGATCTCTGCTAC TGGTCTGTGGATGAGTCGTACCGGAACCATTCATAAAATC AAACACCACGAGGTTAGCCGTTCGAAAATCTATATTGAGA TGGCGTGTGGCGATCATCTGGTTGTGAACAATAGCCGCTC TTCTCGTACAGCACGTGCACTGCGTCACCACAAATATCGT AAAACCTGTAAACGTTGCCGTGTGTCCGATGAGGATCTGA ACAAATTCCTGACAAAAGCCAATGAGGACCAAACAAGCGT GAAAGTGAAAGTCGTTAGCGCTCCTACCCGTACTAAAAAA GCAATGCCGAAATCCGTTGCTCGTGCCCCTAAACCACTGG AAAACACTGAAGCAGCACAGGCACAGCCGTCTGGAAGCAA ATTCTCTCCGGCCATTCCTGTTTCTACCCAGGAGTCCGTT TCTGTTCCAGCAAGTGTGAGCACCAGCATTAGCAGTATTA GCACCGGTGCCACCGCTAGCGCCCTGGTTAAAGGCAATAC CAATCCGATTACAAGCATGTCTGCCCCGGTTCAAGCATCA GCTCCAGCACTGACAAAATCCCAAACCGATCGTCTGGAGG TTCTGCTGAATCCGAAAGACGAAATCAGCCTGAATTCCGG CAAACCGTTTCGTGAACTGGAGAGCGAACTGCTGTCACGT CGTAAAAAAGACCTGCAACAAATCTATGCCGAAGAACGTG AGAACTATCTGGGGAAACTGGAACGTGAAATCACCCGCTT TTTCGTGGATCGTGGCTTTCTGGAGATCAAATCCCCGATT CTGATTCCTCTGGAGTATATCGAGCGTATGGGCATCGACA ATGATACCGAACTGAGCAAACAAATTTTCCGTGTGGATAA AAACTTCTGTCTGCGCCCTATGCTAGCACCAAATCTGGCT AACTATCTGCGCAAACTGGACCGTGCCCTGCCTGATCCTA TCAAAATCTTCGAGATCGGCCCGTGTTATCGTAAAGAGTC CGACGGTAAAGAACATCTGGAGGAGTTTACCATGCTGAAC TTTTGCCAAATGGGTTCAGGTTGTACTCGTGAGAACCTGG AAAGCATCATCACCGATTTTCTGAACCACCTGGGCATTGA CTTCAAAATTGTGGGCGACAGCTGTATGGTGTTTGGCGAC ACCCTGGATGTCATGCACGGCGACCTGGAACTGTCTAGTG CCGTTGTGGGCCCAATCCCGCTGGATCGTGAGTGGGGTAT CGACAAACCTTGGATCGGTGCGGGTTTTGGTCTGGAGCGT CTGCTGAAAGTAAAACACGACTTCAAGAACATCAAACGTG CTGCACGTTCCGAGTCCTATTACAATGGTATTTCTACTAA CCTGTAA(SEQ ID NO: 51)

[0408] protein: MASVKVAVRVRPMNRREKDLEAKFIIQMEKSKTTITNLKI PEGGTGDSGRERTKTFTYDFSFYSADTKSPDYVSQEMVFK TLGTDVVKSAFEGYNACVFAYGQTGSGKSYTMMGNSGDSG LIPRICEGLFSRINETTRWDEASFRTEVSYLEIYNERVRD LLRRKSSKTFNLRVREHPKEGPYVEDLSKHLVQNYGDVEE LMDAGNINRTTAATGMNDVSSRSHAIFTIKFTQAKFDSEM PCETVSKIHLVDLAGSERADATGATGVRLKEGGNINKSLV TLGNVISALADLSQDAANTLAKKKQVFVPYRDSVLTWLLK DSLGGNSKTIMIATISPADVNYGETLSTLRYANRAKNIIN KPTINEDANVKLIRELRAEIARLKTLLAQGNQIALLDSPT YTDIEMNRLGKGAPGSAGSAAGSGMASNDYTQQATQSYGA YPTQPGQGYSQQSSQPYGQQSYSGYSQSTDTSGYGQSSYS SYGQSQNTGYGTQSTPQGYGSTGGYGSSQSSQSSYGQQSS YPGYGQQPAPSSTSGSYGSSSQSSSYGQPQSGSYSQQPSY GGQQQSYGQQQSYNPPQGYGQQNQYNSSSGGGGGGGGGGN YGQDQSSMSSGGGSGGGYGNQDQSGGGGSGGYGQQDRGGR GRGGSGGGGGGGGGGYNRSSGGYEPRGRGGGRGGRGGMGG SDRGGFNKFGGPRDQGSRHDSEQDNSDNNTIFVQGLGENV TIESVADYFKQIGIIKTNKKTGQPMINLYTDRETGKLKGE ATVSFDDPPSAKAAIDWFDGKEFSGNPIKVSFATRRADFN RGGGNGRGGRGRGGPMGRGGYGGGGSGGGGRGGFPSGGGG GGGQQRAGDWKCPNPTCENMNFSWRNECNQCKAPKPDPGG GGPGGSHMGGNYGDDRRGGRGGAIAGAPGSAGSAGSGAC PVPLQLPPLERLTLDDKKPLNTLISATGLWMSRTGTIHKI KHHEVSRSKIYIEMACGDHLVVNNSRSSTARARRHHKYR KTCKRCRVSDEDLNKFLTKANEDQTSVKVKVVSAPTRTKK AMPKSVARAKPPLENTEAAQAQPSGSKFSPAIPVSTQESV SVPASVSTSISISTGATASALVKGNTNPITSMSPAVQAS APALTKSQTDRLEVLLNPKDEISLNSGKPFRELESELLSR RKKDLQQIYAEERENYLGKLEREITRFFVDRGFLEIKSPI LIPLEYIERMGIDNDTELSKQIFRVDKNFCLRPMLAPNLA NYLRKLDRALPDPIKIFEIGPCYRKESDGKEHLEEFTMLN FCQMGSGCTRENLESIITDFLNHLGIDFKIVGDSCMVFGD TLDVMHGDLELSSAVVGPIPLDREWGIDKPWIGAGFGLER LLKVKHDFKNIKRAARSESYYNGISTNL(SEQ ID NO:52)

[0409] KIF16B-FUS-PylRS AA

[0410] DNA: ATGGCATCGGTCAAGGTGGCGTGAGGGTCCGGCCCATGA ATCGCAGGGAAAAGGACTTGGAGGCCAAGTTCATTATTCA GATGGAGAAAAGCAAAACGACAATCACAAACTTAAAGATA CCAGAAGGAGGCACTGGGGACTCAGGAAGAGAACGGACCA AGACCTTCACCTATGACTTTTCTTTTTATTCTGCTGATAC AAAAAGCCCAGATTACGTTTCACAAGAAATGGTTTTCAAA ACCCTCGGCACAGATGTCGTGAAGTCTGCATTTGAAGGTT ATAATGCTTGTGTCTTTGCATATGGGCAAACTGGATCTGG AAAGTCATACACTATGATGGGAAATTCTGGAGATTCTGGC TTAATACCTCGGATCTGTGAAGGACTCTTCAGTCGGATAA ATGAAACCACCAGATGGGATGAAGCTTCTTTTCGAACTGA AGTCAGCTACTTAGAAATTTATAACGAACGTGTGAGAGAT CTACTTCGGCGGAAGTCATCTAAAACCTTCAATTTGAGAG TCCGTGAGCATCCCAAAGAAGGCCCTTATGTTGAGGATTT ATCCAAACATTTAGTACAGAATTATGGTGACGTAGAAGAA CTTATGGATGCGGGCAATATCAACCGGACCACCGCAGCGA CTGGGATGAACGACGTCAGTAGCAGGTCTCATGCCATCTT CACCATCAAGTTCACTCAGGCTAAATTTGATTCTGAAATG CCATGTGAAACCGTCAGTAAGATCCACTTGGTTGATCTTG CCGGAAGTGAGCGTGCAGATGCCACCGGAGCCACCGGGGT TAGGCTAAAGGAAGGGGGAAATATTAACAAGTCCCTCGTG ACTCTGGGGAACGTCATTTCTGCCTTAGCTGATTTATCTC AGGATGCTGCAAATACTCTTGCAAAGAAGAAGCAAGTTTT CGTGCCTTACAGGGATTCTGTGTTGACTTGGTTGTTAAAA GATAGCCTTGGAGGAAACTCTAAAACTATCATGATTGCCA CCATTTCACCTGCTGATGTCAATTATGGAGAAACCCTAAG TACTCTTCGCTATGCAAATAGAGCCAAAAACATCATCAAC AAGCCTACCATTAATGAGGATGCCAACGTCAAACTTATCC GTGAGCTGCGAGCTGAAATAGCCAGACTGAAAACGCTGCT TGCTCAAGGGAATCAGATTGCCCTCTTAGACTCCCCCACA TATACAGATATTGAAATGAACAGATTGGGAAAGGGCGCCC CCGGCTCCGCCGGCTCCGCCGCCGGCTCCGGCATGGCCTC AAACGATTATACCCAACAAGCAACCCAAAGCTATGGGGCC TACCCCACCCAGCCCGGGCAGGGCTATTCCCAGCAGAGCA GTCAGCCCTACGGACAGCAGAGTTACAGTGGTTATAGCCA GTCCACGGACACTTCAGGATATGGCCAGAGCAGCTATTCT TCTTATGGCCAGAGCCAGAACACAGGCTATGGAACTCAGT CAACTCCCCAGGGATATGGCTCGACTGGCGGCTATGGCAG TAGCCAGAGCTCCCAATCGTCTTACGGGCAGCAGTCCTCC TACCCTGGCTATGGCCAGCAGCCAGCTCCCAGCAGCACCT CGGGAAGTTACGGTAGCAGTTCTCAGAGCAGCAGCTATGG GCAGCCCCAGAGTGGGAGCTACAGCCAGCAGCCTAGCTAT GGTGGACAGCAGCAAAGCTATGGACAGCAGCAAAGCTATA ATCCCCCTCAGGGCTATGGACAGCAGAACCAGTACAACAG CAGCAGTGGTGGTGGAGGTGGAGGTGGAGGTGGAGGTAAC TATGGCCAAGATCAATCCTCCATGAGTAGTGGTGGTGGCA GTGGTGGCGGTTATGGCAATCAAGACCAGAGTGGTGGAGG TGGCAGCGGTGGCTATGGACAGCAGGACCGTGGAGGCCGC GGCAGGGGTGGCAGTGGTGGCGGCGGCGGCGGCGGCGGTG GTGGTTACAACCGCAGCAGTGGTGGCTATGAACCCAGAGG TCGTGGAGGTGGCCGTGGAGGCAGAGGTGGCATGGGCGGA AGTGACCGTGGTGGCTTCAATAAATTTGGTGGCCCTCGGG ACCAAGGATCACGTCATGACTCCGAACAGGATAATTCAGA CAACAACACCATCTTTGTGCAAGGCCTGGGTGAGAATGTT ACAATTGAGTCTGTGGCTGATTACTTCAAGCAGATTGGTA TTATTAAGACAAACAAGAAAACGGGACAGCCCATGATTAA TTTGTACACAGACAGGGAAACTGGCAAGCTGAAGGGAGAG GCAACGGTCTCTTTTGATGACCCACCTTCAGCTAAAGCAG CTATTGACTGGTTTGATGGTAAAGAATTCTCCGGAAATCC TATCAAGGTCTCATTTGCTACTCGCCGGGCAGACTTTAAT CGGGGTGGTGGCAATGGTCGTGGAGGCCGAGGGCGAGGAG GACCCATGGGCCGTGGAGGCTATGGAGGTGGTGGCAGTGG TGGTGGTGGCCGAGGAGGATTTCCCAGTGGAGGTGGTGGC GGTGGAGGACAGCAGCGAGCTGGTGACTGGAAGTGTCCTA ATCCCACCTGTGAGAATATGAACTTCTCTTGGAGGAATGA ATGCAACCAGTGTAAGGCCCCTAAACCAGATGGCCCAGGA GGGGGACCAGGTGGCTCTCACATGGGGGGTAACTACGGGG ATGATCGTCGTGGTGGCAGAGGAGGCGGCGCCCCCGGCTC CGCCGGCTCCGCCGCCGGCTCCGGCATGGCGTGCCCGGTG CCGCTGCAGCTGCCGCCGCTGGAACGCCTGACCCTGGATG ACAAAAAACCGCTGAATACCCTGATCTCTGCTACTGGTCT GTGGATGAGTCGTACCGGAACCATTCATAAAATCAAACAC CACGAGGTTAGCCGTTCGAAAATCTATATTGAGATGGCGT GTGGCGATCATCTGGTTGTGAACAATAGCCGCTCTTCTCG TACAGCACGTGCACTGCGTCACCACAAATATCGTAAAACC TGTAAACGTTGCCGTGTGTCCGATGAGGATCTGAACAAAT TCCTGACAAAAGCCAATGAGGACCAAACAAGCGTGAAAGT GAAAGTCGTTAGCGCTCCTACCCGTACTAAAAAAGCAATG CCGAAATCCGTTGCTCGTGCCCCTAAACCACTGGAAAACA CTGAAGCAGCACAGGCACAGCCGTCTGGAAGCAAATTCTC TCCGGCCATTCCTGTTTCTACCCAGGAGTCCGTTTCTGTT CCAGCAAGTGTGAGCACCAGCATTAGCAGTATTAGCACCG GTGCCACCGCTAGCGCCCTGGTTAAAGGCAATACCAATCC GATTACAAGCATGTCTGCCCCGGTTCAAGCATCAGCTCCA GCACTGACAAAATCCCAAACCGATCGTCTGGAGGTTCTGC TGAATCCGAAAGACGAAATCAGCCTGAATTCCGGCAAACC GTTTCGTGAACTGGAGAGCGAACTGCTGTCACGTCGTAAA AAAGACCTGCAACAAATCTATGCCGAAGAACGTGAGAACT ATCTGGGGAAACTGGAACGTGAAATCACCCGCTTTTTCGT GGATCGTGGCTTTCTGGAGATCAAATCCCCGATTCTGATT CCTCTGGAGTATATCGAGCGTATGGGCATCGACAATGATA CCGAACTGAGCAAACAAATTTTCCGTGTGGATAAAAACTT CTGTCTGCGCCCTATGCTGGCACCAAATCTGTATAACTAT CTGCGCAAACTGGACCGTGCCCTGCCTGATCCTATCAAAA TCTTCGAGATCGGCCCGTGTTATCGTAAAGAGTCCGACGG TAAAGAACATCTGGAGGAGTTTACCATGCTGGCCTTTGCC CAAATGGGTTCAGGTTGTACTCGTGAGAACCTGGAAAGCA TCATCACCGATTTTCTGAACCACCTGGGCATTGACTTCAA AATTGTGGGCGACAGCTGTATGGTGTATGGCGACACCCTG GATGTCATGCACGGCGACCTGGAACTGTCTAGTGCCGTTG TTGGACCAATTCCGCTGGACCGTGAGTGGGGTATCGACAA ACCGTGGATCGGAGCAGGATTCGGTCTGGAACGCCTGCTG AAAGTGAAACACGACTTCAAAAACATCAAACGTGCCGCCC GTTCTGAATCGTATTATAACGGGATCTCTACGAACCTGTA A(SEQ ID NO:53)

[0411] protein: MASVKVAVRVRPMNRREKDLEAKFIIQMEKSKTTITNLKI PEGGTGDSGRERTKTFTYDFSFYSADTKSPDYVSQEMVFK TLGTDVVKSAFEGYNACVFAYGQTGSGKSYTMMGNSGDSG LIPRICEGLFSRINETTRWDEASFRTEVSYLEIYNERVRD LLRRKSSKTFNLRVREHPKEGPYVEDLSKHLVQNYGDVEE LMDAGNINRTTAATGMNDVSSRSHAIFTIKFTQAKFDSEM PCETVSKIHLVDLAGSERADATGATGVRLKEGGNINKSLV TLGNVISALADLSQDAANTLAKKKQVFVPYRDSVLTWLLK DSLGGNSKTIMIATISPADVNYGETLSTLRYANRAKNIIN KPTINEDANVKLIRELRAEIARLKTLLAQGNQIALLDSPT YTDIEMNRLGKGAPGSAGSAAGSGMASNDYTQQATQSYGA YPTQPGQGYSQQSSQPYGQQSYSGYSQSTDTSGYGQSSYS SYGQSQNTGYGTQSTPQGYGSTGGYGSSQSSQSSYGQQSS YPGYGQQPAPSSTSGSYGSSSQSSSYGQPQSGSYSQQPSY GGQQQSYGQQQSYNPPQGYGQQNQYNSSSGGGGGGGGGGN YGQDQSSMSSGGGSGGGYGNQDQSGGGGSGGYGQQDRGGR GRGGSGGGGGGGGGGYNRSSGGYEPRGRGGGRGGRGGMGG SDRGGFNKFGGPRDQGSRHDSEQDNSDNNTIFVQGLGENV TIESVADYFKQIGIIKTNKKTGQPMINLYTDRETGKLKGE ATVSFDDPPSAKAAIDWFDGKEFSGNPIKVSFATRRADFN RGGGNGRGGRGRGGPMGRGGYGGGGSGGGGRGGFPSGGGG GGGQQRAGDWKCPNPTCENMNFSWRNECNQCKAPKPDGPG GGPGGSHMGGNYGDDRRGGRGGGAPGSAGSAAGSGMACPV PLQLPPLERLTLDDKKPLNTLISATGLWMSRTGTIHKIKH HEVSRSKIYIEMACGDHLVVNNSRSSRTARALRHHKYRKT CKRCRVSDEDLNKFLTKANEDQTSVKVKVVSAPTRTKKAM PKSVARAPKPLENTEAAQAQPSGSKFSPAIPVSTQESVSV PASVSTSISSISGATASALVKGNTNPITSMSAPVQASAP ALTKSQTDRLEVLLNPKDEISLNSGKPFRELESELLSRRK KDLQQIYAEERENYLGKLEREITRFFVDRGFLEIKSPILI PLEYIERMGIDNDTELSKQIFRVDKNFCLRPMLAPNLYNY LRKLDRALPDPIKIFEIGPCYRKESDGKEHLEEFTMLAFA QMGSGCTRENLESIITDFLNHLGIDFKIVGDSCMVYGDTL DVMHGDLELSSAVVGPIPLDREWGIDKPWIGAGFGLERLL KVKHDFKNIKRAARSESYYNGISTNL(SEQ ID NO:54)

[0412] KIF16B-FUS-PylRS AAAF

[0413] DNA: ATGGCATCGGTCAAGGTGGCGTGAGGGTCCGGCCCATGA ATCGCAGGGAAAAGGACTTGGAGGCCAAGTTCATTATTCA GATGGAGAAAAGCAAAACGACAATCACAAACTTAAAGATA CCAGAAGGAGGCACTGGGGACTCAGGAAGAGAACGGACCA AGACCTTCACCTATGACTTTTCTTTTATTCTGCTGATAC AAAAAGCCCAGATTACGTTTCACAAGAAATGGTTTTCAAA ACCCTCGGCACAGATGTCGTGAAGTCTGCATTTGAAGGTT ATAATGCTTGTGTCTTTGCATATGGGCAAACTGGATCTGG AAAGTCATACACTATGATGGGAAATTCTGGAGATTCTGGC TTAATACCTCGGATCTGTGAAGGACTCTTCAGTCGGATAA ATGAAACCACCAGATGGGATGAAGCTTCTTTTCGAACTGA AGTCAGCTACTTAGAAATTTATAACGAACGTGTGAGAGAT CTACTTCGGCGGAAGTCATCTAAAACCTTCAATTTGAGAG TCCGTGAGCATCCCAAAGAAGGCCCTTATGTTGAGGATTT ATCCAAACATTTAGTACAGAATTATGGTGACGTAGAAGAA CTTATGGATGCGGGCAATATCAACCGGACCACCGCAGCGA CTGGGATGAACGACGTCAGTAGCAGGTCTCATGCCATCTT CACCATCAAGTTCACTCAGGCTAAATTTGATTCTGAAATG CCATGTGAAACCGTCAGTAAGATCCACTTGGTTGATCTTG CCGGAAGTGAGCGTGCAGATGCCACCGGAGCCACCGGGGT TAGGCTAAAGGAAGGGGGAAATATTAACAAGTCCCTCGTG ACTCTGGGGAACGTCATTTCTGCCTTAGCTGATTTATCTC AGGATGCTGCAAATACTCTTGCAAAGAAGAAGCAAGTTTT CGTGCCTTACAGGGATTCTGTGTTGACTTGGTTGTTAAAA GATAGCCTTGGAGGAAACTCTAAAACTATCATGATTGCCA CCATTTCACCTGCTGATGTCAATTATGGAGAAACCCTAAG TACTCTTCGCTATGCAAATAGAGCCAAAAACATCATCAAC AAGCCTACCATTAATGAGGATGCCAACGTCAAACTTATCC GTGAGCTGCGAGCTGAAATAGCCAGACTGAAAACGCTGCT TGCTCAAGGGAATCAGATTGCCCTCTTAGACTCCCCCACA TATACAGATATTGAAATGAACAGATTGGGAAAGGGCGCCC CCGGCTCCCGCCGGCTCCGCCGCCGGCTCCGGCATGGCCTC AAACGATTATACCCAACAAGCAACCCAAAGCTATGGGGCC TACCCCACCCAGCCCGGGCAGGGCTATTCCCAGCAGAGCA GTCAGCCCTACGGACAGCAGAGTTACAGTGGTTATAGCCA GTCCACGGACACTTCAGGATATGGGCCAGAGCAGCTATTCT TCTTATGGCCAGAGCCAGAACACAGGCTATGGAACTCAGT CAACTCCCCAGGGATATGGCTCGACTGGCGGCTATGGCAG TAGCCAGAGCTCCCAATCGTCTTACGGGCAGCAGTCCTCC TACCCTGGCTATGGCCAGCAGCCAGCTCCCAGCAGCACCT CGGGAAGTTACGGTAGCAGTTCTCAGAGCAGCAGCTATGG GCAGCCCCAGAGTGGGAGCTACAGCCAGCAGCCTAGCTAT GGTGGACAGCAGCAAAGCTATGGACAGCAGCAAAGCTATA ATCCCCCTCAGGGCTATGGACAGCAGAACCAGTACAACAG CAGCAGTGGTGGTGGAGGTGGAGGTGGAGGTGGAGGTAAC TATGGCCAAGATCAATCCTCCATGAGTAGTGGTGGTGGCA GTGGTGGCGGTTATGGCAATCAAGACCAGAGTGGTGGAGG TGGCAGCGGTGGCTATGGACAGCAGGACCGTGGAGGCCGC GGCAGGGGTGGCAGTGGTGGCGGCGGCGGCGGCGGCGGTG GTGGTTACAACCGCAGCAGTGGTGGCTATGAACCCAGAGG TCGTGGAGGTGGCCGTGGAGGCAGAGGTGGCATGGGCGGA AGTGACCGTGGTGGCTTCAATAAATTTGGTGGCCCTCGGG ACCAAGGATCACGTCATGACTCCGAACAGGATAATTCAGA CAACAACACCATCTTTGTGCAAGGCCTGGGTGAGAATGTT ACAATTGAGTCTGTGGCTGATTACTTCAAGCAGATTGGTA TTATTAAGACAAACAAGAAAACGGGACAGCCCATGATTAA TTTGTACACAGACAGGGAAACTGGCAAGCTGAAGGGAGAG GCAACGGTCTCTTTTGATGACCCACCTTCAGCTAAAGCAG CTATTGACTGGTTTGATGGTAAAGAATTCTCCGGAAATCC TATCAAGGTCTCATTTGCTACTCGCCGGGCAGACTTTAAT CGGGGTGGTGGCAATGGTCGTGGAGGCCGAGGGCGAGGAG GACCCATGGGCCGTGGAGGCTATGGAGGTGGTGGCAGTGG TGGTGGTGGCCGAGGAGGATTTCCCAGTGGAGGTGGTGGC GGTGGAGGACAGCAGCGAGCTGGTGACTGGAAGTGTCCTA ATCCCACCTGTGAGAATATGAACTTCTCTTGGAGGAATGA ATGCAACCAGTGTAAGGCCCCTAAACCAGATGGCCCAGGA GGGGGACCAGGTGGCTCTCACATGGGGGGTAACTACGGGG ATGATCGTCGTGGTGGCAGAGGAGGCGATTACAAGGATGA CGACGATAAGGGTACCGGCGCCCCCGGCTCCGCCGGCTCC GCCGCCGGCTCCGGCGCGTGCCCGGTGCCGCTGCAGCTGC CGCCGCTGGAACGCCTGACCCTGGATGATAAAAAACCGCT GAATACCCTGATCTCTGCTACTGGTCTGTGGATGAGTCGT ACCGGAACCATTCATAAAATCAAACACCACGAGGTTAGCC GTTCGAAAATCTATATTGAGATGGCGTGTGGCGATCATCT GGTTGTGAACAATAGCCGCTCTTCTCGTACAGCACGTGCA CTGCGTCACCACAAATATCGTAAAACCTGTAAACGTTGCC GTGTGTCCGATGAGGATCTGAACAAATTCCTGACAAAAGC CAATGAGGACCAAACAAGCGTGAAAGTGAAAGTCGTTAGC GCTCCTACCCGTACTAAAAAAGCAATGCCGAAATCCGTTG CTCGTGCCCCTAAACCACTGGAAAACACTGAAGCAGCACA GGCACAGCCGTCTGGAAGCAAATTCTCTCCGGCCATTCCT GTTTCTACCCAGGAGTCCGTTTCTGTTCCAGCAAGTGTGA GCACCAGCATTAGCAGTATTAGCACCGGTGCCACCGCTAG CGCCCTGGTTAAAGGCAATACCAATCCGATTACAAGCATG TCTGCCCCGGTTCAAGCATCAGCTCCAGCACTGACAAAAT CCCAAACCGATCGTCTGGAGGTTCTGCTGAATCCGAAAGA CGAAATCAGCCTGAATTCCGGCAAACCGTTTCGTGAACTG GAGAGCGAACTGCTGTCACGTCGTAAAAAAGACCTGCAAC AAATCTATGCCGAAGAACGTGAGAACTATCTGGGGAAACT GGAACGTGAAATCACCCGCTTTTTCGTGGATCGTGGCTTT CTGGAGATCAAATCCCCGATTCTGATTCCTCTGGAGTATA TCGAGCGTATGGGCATCGACAATGATACCGAACTGAGCAA ACAAATTTTCCGTGTGGATAAAAACTTCTGTCTGCGCCCT ATGCTAGCACCAAATCTGGCTAACTATCTGCGCAAACTGG ACCGTGCCCTGCCTGATCCTATCAAAATCTTCGAGATCGG CCCGTGTTATCGTAAAGAGTCCGACGGTAAAGAACATCTG GAGGAGTTTACCATGCTGGCCTTTGCCCAAATGGGTTCAG GTTGTACTCGTGAGAACCTGGAAAGCATCATCACCGATTT TCTGAACCACCTGGGCATTGACTTCAAAATTGTGGGCGAC AGCTGTATGGTGTTTGGCGACACCCTGGATGTCATGCACG GCGACCTGGAACTGTCTAGTGCCGTTGTGGGCCCAATCCC GCTGGATCGTGAGTGGGGTATCGACAAACCTTGGATCGGT GCGGGTTTTGGTCTGGAGCGTCTGCTGAAAGTAAAACACG ACTTCAAGAACATCAAACGTGCTGCACGTTCCGAGTCCTA TTACAATGGTATTTCTACTAACCTGTAA (SEQ ID NO: 55)

[0414] protein: MASVKVAVRVRPMNRREKDLEAKFIIQMEKSKTTITNLKI PEGGTGDSGRERTKTFTYDFSFYSADTKSPDYVSQEMVFK TLGTDVVKSAFEGYNACVFAYGQTGSGKSYTMMGNSGDSG LIPRICEGLFSRINETTRWDEASFRTEVSYLEIYNERVRD LLRRKSSKTFNLRVREHPKEGPYVEDLSKHLVQNYGDVEE LMDAGNINRTTAATGMNDVSSRSHAIFTIKFTQAKFDSEM PCETVSKIHLVDLAGSERADATGATGVRLKEGGNINKSLV TLGNVISALADLSQDAANTLAKKKQVFVPYRDSVLTWLLK DSLGGNSKTIMIATISPADVNYGETLSTLRYANRAKNIIN KPTINEDANVKLIRELRAEIARLKTLLAQGNQIALLDSPT YTDIEMNRLGKGAPGSAGSAAGSGMASNDYTQQATQSYGA YPTQPGQGYSQQSSQPYGQQSYSGYSQSTDTSGYGQSSYS SYGQSQNTGYGTQSTPQGYGSTGGYGSSQSSQSSYGQQSS YPGYGQQPAPSSTSGSYGSSSQSSSYGQPQSGSYSQQPSY GGQQQSYGQQQSYNPPQGYGQQNQYNSSSGGGGGGGGGGN YGQDQSSMSSGGGSGGGYGNQDQSGGGGSGGYGQQDRGGR GRGGSGGGGGGGGGGYNRSSGGYEPRGRGGGRGGRGGMGG SDRGGFNKFGGPRDQGSRHDSEQDNSDNNTIFVQGLGENV TIESVADYFKQIGIIKTNKKTGQPMINLYTDRETGKLKGE ATVSFDDPPSAKAAIDWFDGKEFSGNPIKVSFATRRADFN RGGGNGRGGRGRGGPMGRGGYGGGGSGGGGRGGFPSGGGG GGGQQRAGDWKCPNPTCENMNFSWRNECNQCKAPKPDGPG GGPGGSHMGGNYGDDRRGGRGGDYKDDDDKGTGAPGSAGS AAGSGACPVPLQLPPLERLTLDDKKPLNTLISATGLWMSR TGTIHKIKHHEVSRSKIYIEMACGDHLVVNNSRSSRTARA LRHHKYRKTCKRCRVSDEDLNKFLTKANEDQTSVKVKVVS APTRTKKAMPKSVARAPKPLENTEAAQAQPSGSKFSPAIP VSTQESVSVPASVSTSISSISTGATASALVKGNTNPITSM SAPVQASAPALTKSQTDRLEVLLNPKDEISLNSGKPFREL ESELLSRRKKDLQQIYAEERENYLGKLEREITRFFVDRGF LEIKSPILIPLEYIERMGIDNDTELSKQIFRVDKNFCLRP MLAPNLANYLRKLDRALPDPIKIFEIGPCYRKESDGKEHL EEFTMLAFAQMGSGCTRENLESIITDFLNHLGIDFKIVGD SCMVFGDTLDVMHGDLELSSAVVGPIPLDREWGIDKPWIG AGFGLERLLKVKHDFKNIKRAARSESYYNGISTNL(SEQ ID NO: 56)

[0415] KIF16B-EWSR1-MCP

[0416] DNA: ATGGCATCGGTCAAGGTGGCCGTGAGGGTCCGGCCCATGA ATCGCAGGGAAAAGGACTTGGAGGCCAAGTTCATTATTCA GATGGAGAAAAGCAAAACGACAATCACAAACTTAAAGATA CCAGAAGGAGGCACTGGGGACTCAGGAAGAGAACGGACCA AGACCTTCACCTATGACTTTTCTTTTTATTCTGCTGATAC AAAAAGCCCAGATTACGTTTCACAAGAAATGGTTTTCAAA ACCCTCGGCACAGATGTCGTGAAGTCTGCATTTGAAGGTT ATAATGCTTGTGTCTTTGCATATGGGCAAACTGGATCTGG AAAGTCATACACTATGATGGGAAATTCTGGAGATTCTGGC TTAATACCTCGGATCTGTGAAGGACTCTTCAGTCGGATAA ATGAAACCACCAGATGGGATGAAGCTTCTTTTCGAACTGA AGTCAGCTACTTAGAAATTTATAACGAACGTGTGAGAGAT CTACTTCGGCGGAAGTCATCTAAAACCTTCAATTTGAGAG TCCGTGAGCATCCCAAAGAAGGCCCTTATGTTGAGGATTT ATCCAAACATTTAGTACAGAATTATGGTGACGTAGAAGAA CTTATGGATGCGGGCAATATCAACCGGACCACCGCAGCGA CTGGGATGAACGACGTCAGTAGCAGGTCTCATGCCATCTT CACCATCAAGTTCACTCAGGCTAAATTTGATTCTGAAATG CCATGTGAAACCGTCAGTAAGATCCACTTGGTTGATCTTG CCGGAAGTGAGCGTGCAGATGCCACCGGAGCCACCGGGGT TAGGCTAAAGGAAGGGGGAAATATTAACAAGTCCCTCGTG ACTCTGGGGAACGTCATTTCTGCCTTAGCTGATTTATCTC AGGATGCTGCAAATACTCTTGCAAAGAAGAAGCAAGTTTT CGTGCCTTACAGGGATTCTGTGTTGACTTGGTTGTTAAAA GATAGCCTTGGAGGAAACTCTAAAACTATCATGATTGCCA CCATTTCACCTGCTGATGTCAATTATGGAGAAACCCTAAG TACTCTTCGCTATGCAAATAGAGCCAAAAACATCATCAAC AAGCCTACCATTAATGAGGATGCCAACGTCAAACTTATCC GTGAGCTGCGAGCTGAAATAGCCAGACTGAAAACGCTGCT TGCTCAAGGGAATCAGATTGCCCTCTTAGACTCCCCCACA ATGGCGTCCACGGATTACAGTACCTATAGCCAAGCTGCAG CGCAGCAGGGCTACAGTGCTTACACCGCCCAGCCCACTCA AGGATATGCACAGACCACCCAGGCATATGGGCAACAAAGC TATGGAACCTATGGACAGCCCACTGATGTCAGCTATACCC AGGCTCAGACCACTGCAACCTATGGGCAGACCGCCTATGC AACTTCTTATGGACAGCCTCCCACTGGTTATACTACTCCA ACTGCCCCCCAGGCATACAGCCAGCCTGTCCAGGGGTATG GCACTGGTGCTTATGATACCACCACTGCTACAGTCACCAC CACCCAGGCCTCCTATGCAGCTCAGTCTGCATATGGCACT CAGCCTGCTTATCCAGCCTATGGGCAGCAGCCAGCAGCCA CTGCACCTACAAGACCGCAGGATGGAAACAAGCCCACTGA GACTAGTCAACCTCAATCTAGCACAGGGGGTTACAACCAG CCCAGCCTAGGATATGGACAGAGTAACTACAGTTATCCCC AGGTACCTGGGAGCTACCCCATGCAGCCAGTCACTGCACC TCCATCCTACCCTCCTACCAGCTATTCCTCTACACAGCCG ACTAGTTATGATCAGAGCAGTTACTCTCAGCAGAACACCT ATGGGCAACCGAGCAGCTATGGACAGCAGAGTAGCTATGG TCAACAAAGCAGCTATGGGCAGCAGCCTCCCACTAGTTAC CCACCCCAAACTGGATCCTACAGCCAAGCTCCAAGTCAAT ATAGCCAACAGAGCAGCAGCTACGGGCAGCAGAGTTCATT CCGACAGGACCACCCCAGTAGCATGGGTGTTTATGGGCAG GAGTCTGGAGGATTTTCCGGACCAGGAGAGAACCGGAGCA TGAGTGGCCCTGATAACCGGGGCAGGGGAAGAGGGGGATT TGATCGTGGAGGCATGAGCAGAGGTGGGCGGGGAGGAGGA CGCGGTGGAATGGGCAGCGCTGGAGAGCGAGGTGGCTTCA ATAAGCCTGGTGGACCCATGGATGAAGGACCAGATCTTGA TCTAGGCCCACCTGTAGATCCAGATGAAGACTCTGACAAC AGTGCAATTTATGTACAAGGATTAAATGACAGTGTGACTC TAGATGATCTGGCAGACTTCTTTAAGCAGTGTGGGGTTGT TAAGATGAACAAGAGAACTGGGCAACCCATGATCCACATC TACCTGGACAAGGAAACAGGAAAGCCCAAAGGCGATGCCA CAGTGTCCTATGAAGACCCACCTACTGCCAAGGCTGCCGT GGAATGGTTTGATGGGAAAGATTTTCAAGGGAGCAAACTT AAAGTCTCCCTTGCTCGGAAGAAGCCTCCAATGAACAGTA TGCGGGGTGGTCTGCCACCCCGTGAGGGCAGAGGCATGCC ACCACCACTCCGTGGAGGTCCAGGAGGCCCAGGAGGTCCT GGGGGACCCATGGGTCGCATGGGAGGCCGTGGAGGAGATA GAGGAGGCTTCCCTCCAAGAGGACCCCGGGGTTCCCGAGG GAACCCCTCTGGAGGAGGAAACGTCCAGCACCGAGCTGGA GACTGGCAGTGTCCCAATCCGGGTTGTGGAAACCAGAACT TCGCCTGGAGAACAGAGTGCAACCAGTGTAAGGCCCCAAA GCCTGAAGGCTTCCTCCCGCCACCCTTTCCGCCCCCGGGT GGTGATCGTGGCAGAGGTGGCCCTGGTGGCATGCGGGGAG GAAGAGGTGGCCTCATGGATCGTGGTGGTCCCGGTGGAAT GTTCAGAGGTGGCCGTGGTGGAGACAGAGGTGGCTTCCGT GGTGGCCGGGGCATGGACCGAGGTGGCTTTGGTGGAGGAA GACGAGGTGGCCCTGGGGGGCCCCCTGGACCTTTGATGGA ACAGGATTACAAGGATGACGACGATAAGGGTACCGAGCAG AAGCTGATCTCAGAGGAGGACCTGGGCGCCCCCGGCTCCG CCGGCTCCGCCGCCGGCTCCGGCGCTTCTAACTTTACTCA GTTCGTTCTCGTCGACAATGGCGGAACTGGCGACGTGACT GTCGCCCCAAGCAACTTCGCTAACGGGATCGCTGAATGGA TCAGCTCTAACTCGCGTTCACAGGCTTACAAAGTAACCTG TAGCGTTCGTCAGAGCTCTGCGCAGAATCGCAAATACACC ATCAAAGTCGAGGTGCCTAAAGGCGCCTGGCGTTCGTACT TAAATATGGAACTAACCATTCCAATTTTCGCCACGAATTC CGACTGCGAGCTTATTGTTAAGGCAATGCAAGGTCTCCTA AAAGATGGAAACCCGATTCCCTCAGCAATCGCAGCAAACT CCGGCATCTACTAA (SEQ ID NO:57)

[0417] protein: MASVKVAVRVRPMNRREKDLEAKFIIQMEKSKTTITNLKI PEGGTGDSGRERTKTFTYDFSFYSADTKSPDYVSQEMVFK TLGTDVVKSAFEGYNACVFAYGQTGSGKSYTMMGNSGDSG LIPRICEGLFSRINETTRWDEASFRTEVSYLEIYNERVRD LLRRKSSKTFNLRVREHPKEGPYVEDLSKHLVQNYGDVEE LMDAGNINRTTAATGMNDVSSRSHAIFTIKFTQAKFDSEM PCETVSKIHLVDLAGSERADATGATGVRLKEGGNINKSLV TLGNVISALADLSQDAANTLAKKKQVFVPYRDSVLTWLLK DSLGGNSKTIMIATISPADVNYGETLSTLRYANRAKNIIN KPTINEDANVKLIRELRAEIARLKTLLAQGNQIALLDSPT MASTDYSTYSQAAAQQGYSAYTAQPTQGYAQTTQAYGQQS YGTYGQPTDVSYTQAQTTATYGQTAYATSYGQPPTGYTTP TAPQAYSQPVQGYGTGAYDTTTATVTTTQASYAAQSAYGT QPAYPAYGQQPAATAPTRPQDGNKPTETSQPQSSTGGYNQ PSLGYGQSNYSYPQVPGSYPMQPVTAPPSYPPTSYSSTQP TSYDQSSYSQQNTYGQPSSYGQQSSYGQQSSYGQQPPTSY PPQTGSYSQAPSQYSQQSSSYGQQSSFRQDHPSSMGVYGQ ESGGFSGPGENRSMSGPDNRGRGRGGFDRGGMSRGGRGGG RGGMGSAGERGGFNKPGGPMDEGPDLDLGPPVDPDEDSDN SAIYVQGLNDSVTLDDLADFFKQCGVVKMNKRTGQPMIHI YLDKETGKPKGDATVSYEDPPTAKAAVEWFDGKDFQGSKL KVSLARKKPPMNSMRGGLPPREGRGMPPPLRGGPGGPGGP GGPMGRMGGRGGDRGGFPPRGPRGSRGNPSGGGNVQHRAG DWQCPNPGCGNQNFAWRTECNQCKAPKPEGFLPPPFPPPG GDRGRGGPGGMRGGRGGLMDRGGPGGMFRGGRGGDRGGFR GGRGMDRGGFGGGRRGGPGGPPGPLMEQDYKDDDDKGTEQ KLISEEDLGAPGSAGSAAGSGASNFTQFVLVDNGGTGDVT VAPSNFANGIAEWISSNSRSQAYKVTCSVRQSSAQNRKYT IKVEVPKGAWRSYLNMELTIPIFATNSDCELIVKAMQGLL KDGNPIPSAIAANSGIY (SEQ ID NO: 58)

[0418] KIF16B-FUS-4xλ N22 -PylRS AF

[0419] DNA: ATGGCATCGGTCAAGGTGGCCGTGAGGGTCCGGCCCATGA ATCGCAGGGAAAAGGACTTGGAGGCCAAGTTCATTATTCA GATGGAGAAAAGCAAAACGACAATCACAAACTTAAAGATA CCAGAAGGAGGCACTGGGGACTCAGGAAGAGAACGGACCA AGACCTTCACCTATGACTTTTCTTTTTATTCTGCTGATAC AAAAAGCCCAGATTACGTTTCACAAGAAATGGTTTTCAAA ACCCTCGGCACAGATGTCGTGAAGTCTGCATTTGAAGGTT ATAATGCTTGTGTCTTTGCATATGGGCAAACTGGATCTGG AAAGTCATACACTATGATGGGAAATTCTGGAGATTCTGGC TTAATACCTCGGATCTGTGAAGGACTCTTCAGTCGGATAA ATGAAACCACCAGATGGGATGAAGCTTCTTTTCGAACTGA AGTCAGCTACTTAGAAATTTATAACGAACGTGTGAGAGAT CTACTTCGGCGGAAGTCATCTAAAACCTTCAATTTGAGAG TCCGTGAGCATCCCAAAGAAGGCCCTTATGTTGAGGATTT ATCCAAACATTTAGTACAGAATTATGGTGACGTAGAAGAA CTTATGGATGCGGGCAATATCAACCGGACCACCGCAGCGA CTGGGATGAACGACGTCAGTAGCAGGTCTCATGCCATCTT CACCATCAAGTTCACTCAGGCTAAATTTGATTCTGAAATG CCATGTGAAACCGTCAGTAAGATCCACTTGGTTGATCTTG CCGGAAGTGAGCGTGCAGATGCCACCGGAGCCACCGGGGT TAGGCTAAAGGAAGGGGGAAATATTAACAAGTCCCTCGTG ACTCTGGGGAACGTCATTTCTGCCTTAGCTGATTTATCTC AGGATGCTGCAAATACTCTTGCAAAGAAGAAGCAAGTTTT CGTGCCTTACAGGGATTCTGTGTTGACTTGGTTGTTAAAA GATAGCCTTGGAGGAAACTCTAAAACTATCATGATTGCCA CCATTTCACCTGCTGATGTCAATTATGGAGAAACCCTAAG TACTCTTCGCTATGCAAATAGAGCCAAAAACATCATCAAC AAGCCTACCATTAATGAGGATGCCAACGTCAAACTTATCC GTGAGCTGCGAGCTGAAATAGCCAGACTGAAAACGCTGCT TGCTCAAGGGAATCAGATTGCCCTCTTAGACTCCCCCACA TATACAGATATTGAAATGAACAGATTGGGAAAGGGCGCCC CCGGCTCCGCCGGCTCCGCCGCCGGCTCCGGCATGGCCTC AAACGATTATACCCAACAAGCAACCCAAAGCTATGGGGCC TACCCCACCCAGCCCGGGCAGGGCTATTCCCAGCAGAGCA GTCAGCCCTACGGACAGCAGAGTTACAGTGGTTATAGCCA GTCCACGGACACTTCAGGATATGGCCAGAGCAGCTATTCT TCTTATGGCCAGAGCCAGAACACAGGCTATGGAACTCAGT CAACTCCCCAGGGATATGGCTCGACTGGCGGCTATGGCAG TAGCCAGAGCTCCCAATCGTCTTACGGGCAGCAGTCCTCC TACCCTGGCTATGGCCAGCAGCCAGCTCCCAGCAGCACCT CGGGAAGTTACGGTAGCAGTTCTCAGAGCAGCAGCTATGG GCAGCCCCAGAGTGGGAGCTACAGCCAGCAGCCTAGCTAT GGTGGACAGCAGCAAAGCTATGGACAGCAGCAAAGCTATA ATCCCCCTCAGGGCTATGGACAGCAGAACCAGTACAACAG CAGCAGTGGTGGTGGAGGTGGAGGTGGAGGTGGAGGTAAC TATGGCCAAGATCAATCCTCCATGAGTAGTGGTGGTGGCA GTGGTGGCGGTTATGGCAATCAAGACCAGAGTGGTGGAGG TGGCAGCGGTGGCTATGGACAGCAGGACCGTGGAGGCCGC GGCAGGGGTGGCAGTGGTGGCGGCGGCGGCGGCGGCGGTG GTGGTTACAACCGCAGCAGTGGTGGCTATGAACCCAGAGG TCGTGGAGGTGGCCGTGGAGGCAGAGGTGGCATGGGCGGA AGTGACCGTGGTGGCTTCAATAAATTTGGTGGCCCTCGGG ACCAAGGATCACGTCATGACTCCGAACAGGATAATTCAGA CAACAACACCATCTTTGTGCAAGGCCTGGGTGAGAATGTT ACAATTGAGTCTGTGGCTGATTACTTCAAGCAGATTGGTA TTATTAAGACAAACAAGAAAACGGGACAGCCCATGATTAA TTTGTACACAGACAGGGAAACTGGCAAGCTGAAGGGAGAG GCAACGGTCTCTTTTGATGACCCACCTTCAGCTAAAGCAG CTATTGACTGGTTTGATGGTAAAGAATTCTCCGGAAATCC TATCAAGGTCTCATTTGCTACTCGCCGGGCAGACTTTAAT CGGGGTGGTGGCAATGGTCGTGGAGGCCGAGGGCGAGGAG GACCCATGGGCCGTGGAGGCTATGGAGGTGGTGGCAGTGG TGGTGGTGGCCGAGGAGGATTTCCCAGTGGAGGTGGTGGC GGTGGAGGACAGCAGCGAGCTGGTGACTGGAAGTGTCCTA ATCCCACCTGTGAGAATATGAACTTCTCTTGGAGGAATGA ATGCAACCAGTGTAAGGCCCCTAAACCAGATGGCCCAGGA GGGGGACCAGGTGGCTCTCACATGGGGGGTAACTACGGGG ATGATCGTCGTGGTGGCAGAGGAGGCGCCACCATGGACGC ACAAACACGACGACGTGAGCGTCGCGCTGAGAAACAAGCT CAATGGAAAGCTGCAAACCCACCGCTCGACGGAGCCGGAG CTGGCGCTGGAGCTGGAGCCGGAGCTGGCGGTCTAGCCAC CATGGACGCACAAACACGACGACGTGAGCGTCGCGCTGAG AAACAAGCTCAATGGAAAGCTGCAAACCCACCGCTCGACG GAGCCGGAGCTGGCGCTGGAGCTGGAGCCGGAGCTGGCGG TCTAGCCACCATGGACGCACAAACACGACGACGTGAGCGT CGCGCTGAGAAACAAGCTCAATGGAAAGCTGCAAACCCAC CGCTCGACGGAGCCGGAGCTGGCGCTGGAGCTGGAGCCGG AGCTGGCGGTCTAGCCACCATGGACGCACAAACACGACGA CGTGAGCGTCGCGCTGAGAAACAAGCTCAATGGAAAGCTG CAAACCCACCGCTCGATTACAAGGATGACGACGATAAGGG TACCGGCGCCCCCGGCTCCGCCGGCTCCGCCGCCGGCTCC GGCGCGTGCCCGGTGCCGCTGCAGCTGCCGCCGCTGGAAC GCCTGACCCTGGATGATAAAAAACCGCTGAATACCCTGAT CTCTGCTACTGGTCTGTGGATGAGTCGTACCGGAACCATT CATAAAATCAAACACCACGAGGTTAGCCGTTCGAAAATCT ATATTGAGATGGCGTGTGGCGATCATCTGGTTGTGAACAA TAGCCGCTCTTCTCGTACAGCACGTGCACTGCGTCACCAC AAATATCGTAAAACCTGTAAACGTTGCCGTGTGTCCGATG AGGATCTGAACAAATTCCTGACAAAAGCCAATGAGGACCA AACAAGCGTGAAAGTGAAAGTCGTTAGCGCTCCTACCCGT ACTAAAAAAGCAATGCCGAAATCCGTTGCTCGTGCCCCTA AACCACTGGAAAACACTGAAGCAGCACAGGCACAGCCGTC TGGAAGCAAATTCTCTCCGGCCATTCCTGTTTCTACCCAG GAGTCCGTTTCTGTTCCAGCAAGTGTGAGCACCAGCATTA GCAGTATTAGCACCGGTGCCACCGCTAGCGCCCTGGTTAA AGGCAATACCAATCCGATTACAAGCATGTCTGCCCCGGTT CAAGCATCAGCTCCAGCACTGACAAAATCCCAAACCGATC GTCTGGAGGTTCTGCTGAATCCGAAAGACGAAATCAGCCT GAATTCCGGCAAACCGTTTCGTGAACTGGAGAGCGAACTG CTGTCACGTCGTAAAAAAGACCTGCAACAAATCTATGCCG AAGAACGTGAGAACTATCTGGGGAAACTGGAACGTGAAAT CACCCGCTTTTTCGTGGATCGTGGCTTTCTGGAGATCAAA TCCCCGATTCTGATTCCTCTGGAGTATATCGAGCGTATGG GCATCGACAATGATACCGAACTGAGCAAACAAATTTTCCG TGTGGATAAAAACTTCTGTCTGCGCCCTATGCTAGCACCA AATCTGGCTAACTATCTGCGCAAACTGGACCGTGCCCTGC CTGATCCTATCAAAATCTTCGAGATCGGCCCGTGTTATCG TAAAGAGTCCGACGGTAAAGAACATCTGGAGGAGTTTACC ATGCTGAACTTTTGCCAAATGGGTTCAGGTTGTACTCGTG AGAACCTGGAAAGCATCATCACCGATTTTCTGAACCACCT GGGCATTGACTTCAAAATTGTGGGCGACAGCTGTATGGTG TTTGGCGACACCCTGGATGTCATGCACGGCGACCTGGAAC TGTCTAGTGCCGTTGTGGGCCCAATCCCGCTGGATCGTGA GTGGGGTATCGACAAACCTTGGATCGGTGCGGGTTTTGGT CTGGAGCGTCTGCTGAAAGTAAAACACGACTTCAAGAACA TCAAACGTGCTGCACGTTCCGAGTCCTATTACAATGGTAT TTCTACTAACCTGTAA (SEQ ID NO:59)

[0420] protein: MASVKVAVRVRPMNRREKDLEAKFIIQMEKSKTTITNLKI PEGGTGDSGRERTKTFTYDFSFYSADTKSPDYVSQEMVFK TLGTDVVKSAFEGYNACVFAYGQTGSGKSYTMMGNSGDSG LIPRICEGLFSRINETTRWDEASFRTEVSYLEIYNERVRD LLRRKSSKTFNLRVREHPKEGPYVEDLSKHLVQNYGDVEE LMDAGNINRTTAATGMNDVSSRSHAIFTIKFTQAKFDSEM PCETVSKIHLVDLAGSERADATGATGVRLKEGGNINKSLV TLGNVISALADLSQDAANTLAKKKQVFVPYRDSVLTWLLK DSLGGNSKTIMIATISPADVNYGETLSTLRYANRAKNIIN KPTINEDANVKLIRELRAEIARLKTLLAQGNQIALLDSPT YTDIEMNRLGKGAPGSAGSAAGSGMASNDYTQQATQSYGA YPTQPGQGYSQQSSQPYGQQSYSGYSQSTDTSGYGQSSYS SYGQSQNTGYGTQSTPQGYGSTGGYGSSQSSQSSYGQQSS YPGYGQQPAPSSTSGSYGSSSQSSSYGQPQSGSYSQQPSY GGQQQSYGQQQSYNPPQGYGQQNQYNSSSGGGGGGGGGGN YGQDQSSMSSGGGSGGGYGNQDQSGGGGSGGYGQQDRGGR GRGGSGGGGGGGGGYNRSSGGYEPRGRGGGRGGRGGMGG SDRGGFNKFGGPRDQGSRHDSEQDNSDNTIFVQGLGENV TIESWADYFKQIGIIKTNKKTGQPMINLYTDRETGKLKGE ATVSFDDPPSAKAAIDWFDGKEFSGNPIKVSFATRRADFN RGGGNGRGGRGRGGPMGRGGYGGGGSGGGGRGGFPSGGGG GGGQQRAGDWKCPNPTCENMNFSWRNECNQCKAPKPDGPG GGPGGSHMGGNYGDDRRGGRGGATMDAQTRRRERRAEKQA QWKAANPPLDGAGAGAGAGAGLATMDAQTRRRERRAE KQAQWKAANPPLDGAGAGAGAGAGLATMDAQTRRRER RAEKQAQWKAANPPLDGAGAGAGAGAGGLATMDAQTRR RERRAEKQAQWKAANPPLDYKDDDDKGTGAPGSAGSAAGS GACPVPLQLPPLERLTLDKPLNTLISATGLWMSRTGTI HKIKHHEVSRSKIYIEMACGDHLVVNNSRSSRTARALRHH KYRKTCKRCRVSDEDLNKFLTKANEDQTSVKVKVVSAPTR TKKAMPKSWARAPKPLENTEAQPSGSKFSPAIPVSTQ ESVSVPASVSTSISSISTGATASALVKGNTNPITSMSAPV QASAPALTKSQTDRLEVLLNPKDEISLNSGKPFRELESEL LSRRKKDLQQIYAEERENYLGKLEREITRFFVDRGFLEIK SPILIPLEYIERMGIDNDTELSKQIFRVDKNFCLRPMLAP NLANYLRKLDRALPDPIKIFEIGPCYRKESDGKEHLEEFT MLNFCQMGSGCTRENLESIITDFLNHLGIDFKIVGDSCMV FGDTLDVMHGDLELSSAVVGPIPLDREWGIDKPWIGAGFG LERLLKVKHDFKNIKRAARSESYYNGISTNL(SEQ ID NO: 60)

[0421] KIF16B - FUS - MCP - PylRS AF

[0422] DNA: ATGGCATCGGTCAAGGTGGCCGTGAGGGTCCGGCCCATGA ATCGCAGGGAAAAGGACTTGGAGGCCAAGTTCATTATTCA GATGGAGAAAAGCAAAACGACAATCACAAACTTAAAGATA CCAGAAGGAGGCACTGGGGACTCAGGAAGAGAACGGACCA AGACCTTCACCTATGACTTTTCTTTTTATTCTGCTGATAC AAAAAGCCCAGATTACGTTTCACAAGAAATGGTTTTCAAA ACCCTCGGCACAGATGTCGTGAAGTCTGCATTTGAAGGTT ATAATGCTTGTGTCTTTGCATATGGGCAAACTGGATCTGG AAAGTCATACACTATGATGGGAAATTCTGGAGATTCTGGC TTAATACCTCGGATCTGTGAAGGACTCTTCAGTCGGATAA ATGAAACCACCAGATGGGATGAAGCTTCTTTTCGAACTGA AGTCAGCTACTTAGAAATTTATAACGAACGTGTGAGAGAT CTACTTCGGCGGAAGTCATCTAAAACCTTCAATTTGAGAG TCCGTGAGCATCCCAAAGAAGGCCCTTATGTTGAGGATTT ATCCAAACATTTAGTACAGAATTATGGTGACGTAGAAGAA CTTATGGATGCGGGCAATATCAACCGGACCACCGCAGCGA CTGGGATGAACGACGTCAGTAGCAGGTCTCATGCCATCTT CACCATCAAGTTCACTCAGGCTAAATTTGATTCTGAAATG CCATGTGAAACCGTCAGTAAGATCCACTTGGTTGATCTTG CCGGAAGTGAGCGTGCAGATGCCACCGGAGCCACCGGGGT TAGGCTAAAGGAAGGGGGAAATATTAACAAGTCCCTCGTG ACTCTGGGGAACGTCATTTCTGCCTTAGCTGATTTATCTC AGGATGCTGCAAATACTCTTGCAAAGAAGAAGCAAGTTTT CGTGCCTTACAGGGATTCTGTGTTGACTTGGTTGTTAAAA GATAGCCTTGGAGGAAACTCTAAAACTATCATGATTGCCA CCATTTCACCTGCTGATGTCAATTATGGAGAAACCCTAAG TACTCTTCGCTATGCAAATAGAGCCAAAAACATCATCAAC AAGCCTACCATTAATGAGGATGCCAACGTCAAACTTATCC GTGAGCTGCGAGCTGAAATAGCCAGACTGAAAACGCTGCT TGCTCAAGGGAATCAGATTGCCCTCTTAGACTCCCCCACA TATACAGATATTGAAATGAACAGATTGGGAAAGGGCGCCC CCGGCTCCCGCCGGCTCCGCCGCCGGCTCCGGCATGGCCTC AAACGATTATACCCAACAAGCAACCCAAAGCTATGGGGCC TACCCCACCCAGCCCGGGCAGGGCTATTCCCAGCAGAGCA GTCAGCCCTACGGACAGCAGAGTTACAGTGGTTATAGCCA GTCCACGGACACTTCAGGATATGGGCCAGAGCAGCTATTCT TCTTATGGCCAGAGCCAGAACACAGGCTATGGAACTCAGT CAACTCCCCAGGGATATGGCTCGACTGGCGGCTATGGCAG TAGCCAGAGCTCCCAATCGTCTTACGGGCAGCAGTCCTCC TACCCTGGCTATGGCCAGCAGCCAGCTCCCAGCAGCACCT CGGGAAGTTACGGTAGCAGTTCTCAGAGCAGCAGCTATGG GCAGCCCCAGAGTGGGAGCTACAGCCAGCAGCCTAGCTAT GGTGGACAGCAGCAAAGCTATGGACAGCAGCAAAGCTATA ATCCCCCTCAGGGCTATGGACAGCAGAACCAGTACAACAG CAGCAGTGGTGGTGGAGGTGGAGGTGGAGGTGGAGGTAAC TATGGCCAAGATCAATCCTCCATGAGTAGTGGTGGTGGCA GTGGTGGCGGTTATGGCAATCAAGACCAGAGTGGTGGAGG TGGCAGCGGTGGCTATGGACAGCAGGACCGTGGAGGCCGC GGCAGGGGTGGCAGTGGTGGCGGCGGCGGCGGCGGCGGTG GTGGTTACAACCGCAGCAGTGGTGGCTATGAACCCAGAGG TCGTGGAGGTGGCCGTGGAGGCAGAGGTGGCATGGGCGGA AGTGACCGTGGTGGCTTCAATAAATTTGGTGGCCCTCGGG ACCAAGGATCACGTCATGACTCCGAACAGGATAATTCAGA CAACAACACCATCTTTGTGCAAGGCCTGGGTGAGAATGTT ACAATTGAGTCTGTGGCTGATTACTTCAAGCAGATTGGTA TTATTAAGACAAACAAGAAAACGGGACAGCCCATGATTAA TTTGTACACAGACAGGGAAACTGGCAAGCTGAAGGGAGAG GCAACGGTCTCTTTTGATGACCCACCTTCAGCTAAAGCAG CTATTGACTGGTTTGATGGTAAAGAATTCTCCGGAAATCC TATCAAGGTCTCATTTGCTACTCGCCGGGCAGACTTTAAT CGGGGTGGTGGCAATGGTCGTGGAGGCCGAGGGCGAGGAG GACCCATGGGCCGTGGAGGCTATGGAGGTGGTGGCAGTGG TGGTGGTGGCCGAGGAGGATTTCCCAGTGGAGGTGGTGGC GGTGGAGGACAGCAGCGAGCTGGTGACTGGAAGTGTCCTA ATCCCACCTGTGAGAATATGAACTTCTCTTGGAGGAATGA ATGCAACCAGTGTAAGGCCCCTAAACCAGATGGCCCAGGA GGGGGACCAGGTGGCTCTCACATGGGGGGTAACTACGGGG ATGATCGTCGTGGTGGCAGAGGAGGCGATTACAAGGATGA CGACGATAAGGGTACCGGCGCCCCCGGCTCCGCCGGCTCC GCCGCCGGCTCCGGCGCTTCTAACTTTACTCAGTTCGTTC TCGTCGACAATGGCGGAACTGGCGACGTGACTGTCGCCCC AAGCAACTTCGCTAACGGGATCGCTGAATGGATCAGCTCT AACTCGCGTTCACAGGCTTACAAAGTAACCTGTAGCGTTC GTCAGAGCTCTGCGCAGAATCGCAAATACACCATCAAAGT CGAGGTGCCTAAAGGCGCCTGGCGTTCGTACTTAAATATG GAACTAACCATTCCAATTTTCGCCACGAATTCCGACTGCG AGCTTATTGTTAAGGCAATGCAAGGTCTCCTAAAAGATGG AAACCCGATTCCCTCAGCAATCGCAGCAAACTCCGGCATC TACGGTACCGGCGCCCCCGGCTCCGCCGGCTCCGCCGCCG GCTCCGGCGCGTGCCCGGTGCCGCTGCAGCTGCCGCCGCT GGAACGCCTGACCCTGGATGATAAAAAACCGCTGAATACC CTGATCTCTGCTACTGGTCTGTGGATGAGTCGTACCGGAA CCATTCATAAAATCAAACACCACGAGGTTAGCCGTTCGAA AATCTATATTGAGATGGCGTGTGGCGATCATCTGGTTGTG AACAATAGCCGCTCTTCTCGTACAGCACGTGCACTGCGTC ACCACAAATATCGTAAAACCTGTAAACGTTGCCGTGTGTC CGATGAGGATCTGAACAAATTCCTGACAAAAGCCAATGAG GACCAAACAAGCGTGAAAGTGAAAGTCGTTAGCGCTCCTA CCCGTACTAAAAAAGCAATGCCGAAATCCGTTGCTCGTGC CCCTAAACCACTGGAAAACACTGAAGCAGCACAGGCACAG CCGTCTGGAAGCAAATTCTCTCCGGCCATTCCTGTTTCTA CCCAGGAGTCCGTTTCTGTTCCAGCAAGTGTGAGCACCAG CATTAGCAGTATTAGCACCGGTGCCACCGCTAGCGCCCTG GTTAAAGGCAATACCAATCCGATTACAAGCATGTCTGCCC CGGTTCAAGCATCAGCTCCAGCACTGACAAAATCCCAAAC CGATCGTCTGGAGGTTCTGCTGAATCCGAAAGACGAAATC AGCCTGAATTCCGGCAAACCGTTTCGTGAACTGGAGAGCG AACTGCTGTCACGTCGTAAAAAAGACCTGCAACAAATCTA TGCCGAAGAACGTGAGAACTATCTGGGGAAACTGGAACGT GAAATCACCCGCTTTTTCGTGGATCGTGGCTTTCTGGAGA TCAAATCCCCGATTCTGATTCCTCTGGAGTATATCGAGCG TATGGGCATCGACAATGATACCGAACTGAGCAAACAAATT TTCCGTGTGGATAAAAAC...

Claims

1. An assembler fusion protein (AFP), comprising: (a) (a1) is derived from an intracellular targeting polypeptide, the intracellular targeting polypeptide being a Targeting and therefore locally concentrating intracellular structural elements within or directly adjacent to the cytoplasm a polypeptide segment (IC-TP segment), (a2) A method for producing a phase-separated polypeptide, the phase-separated polypeptide being a high-protein polypeptide in a cytoplasm. The ability to self-associate in the cytoplasm of cells to create sites of high localized concentration. a polypeptide segment (PSP segment), At least one first polypeptide acting as an assembler (AP) selected from A segment, (b) b1) an RNA-targeting polypeptide (RNA-TP) segment, and b2) an orthogonal aminoacyl-tRNA synthetase (O-RS) segment; At least one second polypeptide segment acting as an effector selected from ent (EP) and Including, The polypeptide segments are operably linked on the AFP. Assembler Fusion Proteins (AFPs).

2. An assembler fusion protein (AFP) comprising at least two AFPs according to claim 1. A combination of.

3. A fusion protein (RNA-TP / O-RS fusion protein), (i) at least one RNA targeting polypeptide (RNA-TP) segment; (ii) at least one orthogonal aminoacyl-tRNA synthetase (O-RS) segment; Mention and; Including, The polypeptide segments are connected to each other on the RNA-TP / O-RS fusion protein. are functionally linked to each other, Fusion proteins (RNA-TP / O-RS fusion proteins).

4. (i) at least one AFP according to claim 1 or at least one AFP according to claim 2 a nucleotide sequence encoding a combination of one AFP; or (ii) a nucleic acid sequence complementary to the nucleotide sequence of (i); (iii) both (i) and (ii); A nucleic acid molecule or a combination of two or more nucleic acid molecules comprising:

5. (i) a nucleic acid encoding at least one RNA-TP / O-RS fusion protein according to claim 3; a nucleotide sequence to be read, or (ii) a nucleic acid sequence complementary to (i); or (iii) both (i) and (ii); A nucleic acid molecule or a combination of two or more nucleic acid molecules comprising:

6. The nucleotide sequence of a nucleic acid molecule or combination of nucleic acid molecules according to claim 4 or claim 5. An expression cassette comprising:

7. An expression vector comprising at least one expression cassette according to claim 6.

8. At least one nucleic acid molecule or a set of nucleic acid molecules according to claim 4 or claim 5. In combination with at least one expression cassette according to claim 6 or at least one expression cassette according to claim 7 Also, a cell containing an expression vector.

9. (i) at least one EP selected from the RNA-TP segment; and (ii) an O-R and at least one EP selected from the S segment. A nucleotide sequence encoding or complementary to a nucleotide sequence encoding an AFP. The cell of claim 8, comprising a nucleic acid sequence.

10. At least one of the at least two AFPs contains at least one RNA-TP segment, At least one other of the at least two AFPs comprises at least one O-RS segment. A nucleic acid encoding or coding for a combination of at least two AFPs according to claim 1. The cell of claim 8, comprising a nucleotide sequence that is complementary to the nucleotide sequence.

11. A nucleic acid sequence encoding at least one RNA-TP / O-RS fusion protein according to claim 3. or a nucleotide sequence that is complementary to a nucleotide sequence that encodes the same. The cell according to claim 8.

12. A polypeptide of interest that contains one or more non-standard amino acid (ncAA) residues along its amino acid sequence. A method for preparing a peptide (POI), comprising: The method comprises the steps of: expressing the POI by the cell, The cells: (i) the one or more ncAA residues of the POI are selected by a selector codon(s). the encoded nucleotide sequence encoding the POI (CSPOI); (ii) at least one RNA of an AFP in a cell that is operably linked to a CSPOI; - a targeting nucleotide sequence (TN) capable of interacting with the TP segment; (iii) an anticodon (single or multiple) complementary to the selector codon (single or multiple) of the CSPOI; or more) orthogonal tRNAs ncAA (O-tRNA ncAA )molecule ; Including, The O-tRNA ncAA The molecule binds one or more O-R of at least one of the AFPs in the cell. Together with the S segment, one or more orthogonal O-RS / O-tRNA ncAA pair which are capable of introducing said one or more ncAA residues onto the amino acid sequence of the POI. forgiveness, The method optionally further comprises recovering the expressed POI. method.

13. A polypeptide of interest that contains one or more non-standard amino acid (ncAA) residues along its amino acid sequence. A method for preparing a peptide (POI), comprising: The method comprises: producing PO by the cell of claim 11 in the presence of one or more ncAAs. expressing I, The cells: (iv) the one or more ncAA residues of the POI are selected by a selector codon(s). a nucleotide sequence encoding a POI (CSPOI); (v) an RNA-TP / O-RS fusion protein operably linked to a CSPOI and in a cell; A targeting nucleic acid capable of interacting with at least one RNA-TP segment of a protein. Nucleotide sequence (TN); (vi) an anticodon(s) complementary to the selector codon(s) of the CSPOI; one or more orthogonal tRNAs ncAA (O-tRNA ncAA )molecule; Including, The O-tRNA ncAA The molecule is a fusion protein of the RNA-TP / O-RS in a cell. or more O-RS segments together to form one or more orthogonal O-RS / O-tRNs; A ncAA pair, which are the one or more ncA pairs on the amino acid sequence of the POI. Allows the introduction of A residues, The method optionally further comprises recovering the expressed POI. method.

14. (i) encoding a polypeptide of interest (POI), said POI being regulated by a selector codon; The nucleic acid sequence includes one or more non-standard amino acid (ncAA) residues encoded on the CSPOI. A nucleotide sequence (CSPOI); (ii) a targeting nucleotide sequence (TN); Including, The RNA molecule containing the TN is coupled to an RNA targeting polypeptide (RNA-T P) can interact with Nucleic acid molecule.

15. The kit includes: at least one ncAA corresponding to at least one ncAA residue of the POI or Salt and - at least one expression vector according to claim 7, Including, A polypeptide of interest (PO) having at least one non-standard amino acid (ncAA) residue A kit for preparing I).

Citation Information

Patent Citations

  • Compositions and methods comprising aspartyl-tRNA synthetase having non-standard biological activity

    JP2012522510A

  • Expression of soluble virus-fusion glycoproteins in mammalian cells

    JP2014505477A

  • Methods and compositions for producing orthogonal trna-aminoacyl trna synthetase pairs

    JP2016112021A

  • Orthogonal cas9 proteins for RNA-induced gene regulation and editing

    JP2016523560A

  • Viral particle for RNA transfer, especially into cells involved in immmune response

    WO2017194903A2