Compositions and methods for modifying target molecules
By using a tyrosinase-catalyzed reaction to covalently bind the target molecule to the phenolic biomolecule, the problem of site specificity and functional preservation of target molecule modification in existing technologies is solved, and a stable conjugation between the target molecule and the biomolecule is achieved.
Patent Information
- Application Number
- CN202511625680.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-04
- Filing Date
- 2020-03-19
- Publication Date
- 2026-02-13
AI Technical Summary
Existing technologies struggle to modify target molecules in a simple and site-specific manner, especially when conjugating biomolecules to target molecules, they cannot effectively maintain the function of the biomolecules.
By contacting the thiol moiety in the target molecule with a biomolecule containing a phenolic or catechol moiety, a quinone intermediate is generated through a tyrosinase-catalyzed reaction and reacted with a nucleophile, thus achieving covalent binding between the target molecule and the biomolecule.
Selective modification of target molecules was achieved, the function of biomolecules was maintained, and the stability and activity of the conjugates were improved.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
[0001] This application is a divisional application of Chinese invention patent application number CN202080037231.3. Technical Field
[0002] This application relates to compositions and methods for modifying target molecules.
[0003] Cross-referencing
[0004] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 822,616, filed March 22, 2019, and U.S. Provisional Patent Application No. 62 / 910,836, filed October 4, 2019, the entire contents of which are incorporated herein by reference.
[0005] Incorporate by reference into the sequence list provided as a text file.
[0006] The sequence list is also provided as a text file, “BERK-405WO_SEQ_LISTING_ST25.txt”, created on March 17, 2020, and measuring 8,056 KB in size. The contents of this text file are incorporated herein by reference in their entirety.
[0007] Statement on Federally Funded Research
[0008] This invention was developed with government grants under National Science Foundation licenses 1059083 and 1808189. The government holds certain rights to this invention. Background Technology
[0009] The conjugation of biomolecules to target molecules to create conjugates while preserving the functions of both the biomolecule and the target molecule has long been a goal of chemical biology and biopharmaceutical research. Examples of conjugates include protein-peptide conjugates for vaccine development, antibody-drug conjugates, and antibody-protein conjugates for immunotherapy.
[0010] Although many techniques have been developed to allow the attachment of medium-sized molecules to proteins, developing simple biomolecular modification procedures that can attach proteins or biomolecules to any location on the protein surface in a site-specific manner is challenging.
[0011] There is a need for improved target molecule modification procedures that can modify target molecules in a simple yet site-specific manner. Summary of the Invention
[0012] This disclosure provides a method for chemically selectively modifying a target molecule. The method involves contacting a target molecule comprising a thiol moiety with a biomolecule comprising a reactive moiety, wherein the reactive moiety is generated by reacting the biomolecule comprising a phenolic or catechol moiety with an enzyme capable of oxidizing the phenolic or catechol moiety. The contact is performed under conditions sufficient to conjugate the target molecule to the biomolecule, thereby producing a modified target molecule. This disclosure provides compositions comprising a target molecule comprising a thiol moiety and a biomolecule comprising a phenolic or catechol moiety. This disclosure provides kits for performing the method. This disclosure also provides modified target molecules and methods of using them. Attached Figure Description
[0013] The invention is best understood when read in conjunction with the following detailed description and the accompanying drawings. It should be emphasized that, by convention, the various features in the drawings are not to scale. Rather, for clarity, the dimensions of the various features have been arbitrarily enlarged or reduced. The drawings include the following figures. It should be understood that the drawings described below are for illustrative purposes only. The drawings are not intended to limit the scope of the teachings of the invention in any way.
[0014] Figure 1A The activation of phenolic and catechol moieties with tyrosinase to provide quinone intermediates is shown, followed by reaction of the quinone intermediates with a potential nucleophile.
[0015] Figure 1B An exemplary topic chemoselective modification reaction of the target protein with solvent-exposed thiols (A) and a tyrosine / phenol-containing conjugate (B) is shown to provide a covalently bound conjugate product (C).
[0016] Figure 2 Figure A depicts ESI-TOF data from an experiment showing MS2 N87C modified with α-endorphin and maleimide blocking, demonstrating that the addition reaction is blocked by tyrosinase-catalyzed reactions via maleimide-terminated thiols on the protein, and that the maleimide reaction is also blocked when tyrosinase is performed first. The figure indicates that surface cysteine is the modified residue. Figure 2 Figure B depicts the stability studies of the protein-peptide conjugates under various conditions. All samples were stored in 50 mM phosphate buffer under these conditions.
[0017] Figure 3 An exemplary example of a biomolecule containing a phenolic moiety that is compatible with the subject approach is shown.
[0018] Figure 4ESI-TOF data showing the conjugation of various peptides with cysteine-containing mutants of the MS2 viral capsid are presented. The peptides consist of the following sequences with acylated N-termini: 2NLS: Ac-YGPKKKRKVGGSPKKKRKV (SEQ ID NO: 943); IL13: Ac-GYACGEMGWVRCGGSK (SEQ ID NO: 944); R8: Ac-YGRRRRRRRR (SEQ ID NO: 945); and HIV-Tat: Ac-YGRKKRRQRRRPPQ (SEQ ID NO: 946).
[0019] Figure 5 Figure A shows ESI-TOF data demonstrating that Cas9 (C80, C574) was modified twice with endorphins. Figure 5 Figure B depicts the in vitro DNA cleavage analysis, demonstrating that the peptide (terminus)-modified Cas9 (RNP) maintained cleavage activity even before the addition of guide RNA (apo). For each treatment, RNPs were added along a concentration gradient to determine activity on the target DNA strand. Figure 5 Figure C shows ESI-TOF data demonstrating successful Cas9-GFP conjugation. The sequences are shown below: GYGGS (SEQ ID NO: 1021), MYGGS (SEQ ID NO: 1022). Figure 5 The D-plot depicts the in vitro cleavage analysis, showing that GFP-modified Cas9 retained its activity compared to the control. The sequences are shown below: MYGGS (SEQ ID NO: 1022), SGGGGY (SEQ ID NO: 1040).
[0020] Figure 6 It was demonstrated that Cas9 modified with a peptide containing two copies of the SV40 nuclear localization sequence can enter and edit neural progenitor cells, thereby allowing for a 20-fold increase in editing efficiency.
[0021] Figure 7 The ESI-TOF data shown are for proteins containing phenols modified with small molecule thiols.
[0022] Figure 8 and Figure 9 The amino acid sequence of mushroom tyrosinase was provided. Figure 8 The sequence is shown in SEQ ID NO: 971. Figure 9 The sequence is shown in SEQ ID NO: 972.
[0023] Figures 10A to 10Z and Figures 10AA to 10VVThe amino acid sequence of Bacillus megaterium tyrosinase was provided. Figures 10A to 10Z The sequence is shown in SEQ ID NO: 973-998. Figures 10AA to 10VV The sequence is shown in SEQ ID NO: 999-1020.
[0024] Figure 11 The abTYR-peptide charge screening method is illustrated: peptides containing 5-meric tyrosine residues were coupled to Y182C GFP and pAF MS2 using abTYR. The resulting reaction mixture was analyzed using Q-TOF mass spectrometry. Reaction conditions: 50 M μM GFP, 250 μM peptide, 0.167 μM tyrosinase, 10 mM buffer, pH 6.5, 30 min at room temperature, all reactions quenched with 10 mM tyrosine. The sequences are as follows: GGGGY (SEQ ID NO: 1024), RGGGY (SEQ ID NO: 1025), RGRGY (SEQ ID NO: 1026), RRRGY (SEQ ID NO: 1027), RRRRY (SEQ ID NO: 1028), EGGGY (SEQ ID NO: 1029), EGEGY (SEQ ID NO: 1030), EEEGY (SEQ ID NO: 1031), EEEEY (SEQ ID NO: 1032), GGGWY (SEQ ID NO: 1033), GGWGY (SEQ ID NO: 1034), RRRWY (SEQ ID NO: 1035), RRWRY (SEQ ID NO: 1036), EEEWY (SEQ ID NO: 1037), EEWEY (SEQ ID NO: 1038).
[0025] Figure 12 A to Figure 12 B illustrates the abTYR and bmTYR models: due to the abundance of glutamic acid and aspartic acid residues, abTYR (a) has a total negative charge around its active site (red residues). Conversely, bmTYR (b) has a slight positive charge around its active site (blue residues).
[0026] Figure 13The bmTYR charge screening method was demonstrated: a peptide containing a 5-meric tyrosine residue was conjugated to Y182C GFP using bmTYR. Analysis of the resulting reaction mixture using Q-TOF mass spectrometry showed that bmTYR preferred negatively charged substrates. Reaction conditions: 50 M μM GFP, 250 μM peptide, 0.2 μM tyrosinase, 10 mM buffer, pH 6.5, at 37°C for 30 min. All reactions were quenched with 10 mM tyrosine. The sequences are shown below: GGGGY (SEQ ID NO: 1024), GGGWY (SEQ ID NO: 1033), EEEGY (SEQ ID NO: 1031), RRRGY (SEQ ID NO: 1027).
[0027] Figure 14 The comparison of abTYR and bmTYR with respect to EGGGY (SEQ ID NO: 1029) and EEEEY (SEQ ID NO: 1032) peptides is shown. Reaction conditions (abTYR): 50 M μM GFP, 250 μM peptide, 0.167 μM tyrosinase, 10 mM buffer pH 6.5, 30 min at room temperature. Reaction conditions (bmTYR): 10 M μM GFP, 50 μM peptide, 0.8 μM tyrosinase, 10 mM buffer pH 6.5, 1 h at room temperature, all reactions quenched with 10 mM tyrosine.
[0028] Figure 15 A to Figure 15 C illustrates oxidative coupling strategies for protein modification. a) Chemical and physical methods utilizing ortho-quinones and ortho-iminoquinones for coupling with N-terminal proline residues and aminophenyl groups. b) Tyrosinase-mediated phenol oxidation for coupling with N-terminal proline residues. c) Tyrosine-labeled proteins for selective tyrosinase-mediated generation of ortho-quinones at the N- or C-terminus of the protein, followed by coupling with an exogenous amine nucleophile.
[0029] Figure 16 A to Figure 16B shows the linkage of the Tyr-containing peptide to MS2 (pAF-MS2) containing p-aminophenylalanine. The N-Ac-α-endorphin has an accessible tyrosine residue at its N-terminus. This site can be oxidized by tyrosinase and coupled to the pAF-MS2 capsid containing an aniline group introduced using the Schultz amber codon suppression method. The sequence is shown below: GGFMTSEKSQTPLVT (SEQ ID NO: 1039). (b) The positions of 180 aniline groups are shown in pink on the whole viral capsid (PDB ID: 2MS2). ESI-TOF MS analysis showed almost complete conversion to the expected product (expected: 15589 Da). No over-modification was observed.
[0030] Figure 17 A to Figure 17 E shows the efficiency of tyrosinase-mediated coupling with amine nucleophiles in *Agaricus bisporus*. a) C-terminal-GGY-labeled trastuzumab scFv was used as a model coupling partner. b) Crystal structures of the variable domains of the heavy and light chains of trastuzumab constituting scFv. c) Representative mass spectra of the initiating scFv-GGY before and after coupling with 150 μM aniline. d) Screening of 4-aminophenyl-derived nucleophiles at concentrations from 25 μM to 750 μM. e) Screening of pyrrolidine and piperazine-derived nucleophiles at concentrations from 100 μM to 5000 μM. Conversion rates were estimated by integral analysis using TOF-LCMS. Representative spectra are shown in supporting figure X.
[0031] Figure 18 The images show tyrosine-labeled protein substrates successfully conjugated using *Agaricus bisporus* tyrosinase. The C-terminus is highlighted in red, and internal tyrosine residues are highlighted in orange. The reaction was carried out with tyrosinase and aniline in phosphate buffer at pH 6.5. a) N-terminal labeled ubiquitin. b) C-terminal-(GGGGS)2GGY labeled sfGFP. (SEQ ID NO: 947) c) C-terminal-GGY labeled trastuzumab scFv. The sequence is shown below: SGGGGY (SEQ ID NO: 1040).
[0032] Figure 19 A to Figure 19 C shows a flow cytometry study of fluorophore-conjugated trastuzumab scFv binding to SKBR3 (HER2+) cells. a) Oxidative coupling of GGY-labeled scFv with 12 U / L Agaricus bisporus tyrosinase and 50 μM aniline-Oregon Green 488. b) ESI TOF-MS showed that scFv-GGY was coupled at 85% conversion. The unlabeled form of scFv was unmodified.
[0033] Figure 20 A to Figure 20 B illustrates the exploration of C-terminal linkers and the utility of *Bacillus megaterium* tyrosinase. a) Linkers of various types and lengths were attached to the C-terminus of protein L, including two linkers utilizing the natural domain indirect linker sequences of domains 4 and 5. Standard coupling reactions with *Bacillus megaterium* tyrosinase were performed on protein L variants. The sequence is shown in SEQ ID NO: 1041. b) Transformations were observed by TOF-LCMS following treatment with *Bacillus megaterium* tyrosinase. None of these variants could be modified by *Agaricus bisporus* tyrosinase. The sequences are shown below: (G4S)2GGY (SEQ ID NO: 947), (G4S)3GGY (SEQ ID NO: 1042), A(EAAAK)2AGGY (SEQ ID NO: 1043), (AP)3GGY (SEQ ID NO: 1044), AN 20 GGY (SEQ ID NO: 1045), EIKRTGGY (SEQ ID NO: 1046), G4SGGY (SEQ ID NO: 968).
[0034] Figure 21 A to Figure 21 C shows the C-terminal tyrosine-tagged MBP mediated by Bacillus megaterium. a) Crystal structure of MBP with the C-terminus highlighted in red. Tyrosine residues are shown in orange. Conjugated maltose is shown in yellow. b) Data with MBP-SSGGGGY (SEQ ID NO: 948); c) Data with MBP-GGY.
[0035] Figure 22 A to Figure 22 D shows the detection of HER2+ cells using a “secondary” affinity reagent of the protein-L-OG 488 conjugate. a) Detection protocol: Untyrosine-labeled trastuzumab scFv binds to HER2+ SK-BR-3 cells and is recognized by OG488-modified protein-L. b) Secondary affinity reagent prepared from the -AN20GGY-terminated protein-L variant using 25 μM OG 488-aniline and Bacillus megaterium tyrosinase. c) Mass spectra of protein-L-AN20GGY before and after modification. d) Flow cytometry fluorescence data of SK-BR-3 cells treated according to the above protocol and negative controls. MDA-MB-468 cells were used as HER2- controls. The sequence is shown below: AN 20 GGY (SEQ ID NO: 1045).
[0036] Figure 23The images show C-terminal -GGY tyrosine-labeled and unlabeled trastuzumab scFv subjected to oxidative coupling conditions. 12 U / mL abTYR, 150 μM aniline, 20 mM sodium phosphate buffer, pH 6.5, 1 hour.
[0037] Figure 24 A to Figure 24 D shows the changes in abTyr and aniline concentrations during the conversion of trastuzumab scFv labeled with -GGY. a) Reaction protocol b) Representative mass spectra: 1000 M aniline with variable abTYR concentration c) Conversion % of trastuzumab (“Tras.”) scFv-GGY to aniline conjugate listed in the table d) Graphical representation of the conversion of Tras. scFv-GGY to aniline conjugate.
[0038] Figure 25 The experiment showed the addition of nucleophilic reagents at later stages. Aniline was added to the abTYR-mediated oxidative coupling at 5, 10, 20, 40, or 60 minutes after tyrosinase.
[0039] Figure 26 Representative spectra of oxidative coupling reactions with 4-aminophenyl-derived nucleophiles are shown. a) o-Toluidine, b) 2,6-Dimethylaniline, c) 4-aminophenyl-N-methylamide.
[0040] Figure 27 The representative spectrum of the oxidative coupling reaction is shown.
[0041] Figure 28 The oxidative coupling reaction of protein L variant is shown. The sequences are as follows: (G4S)2GGY (SEQ ID NO: 947); (G4S)3GGY (SEQ ID NO: 1042), A(EAAAK)2AGGY (SEQ ID NO: 1043), (AP)3GGY (SEQ ID NO: 1044).
[0042] Figure 29 A to Figure 29B shows the stability study of trastuzumab scFv-GGY in protein storage buffer (20 mM Na2HPO4, 150 mM NaCl, containing 15% glycerol, pH 7.4) at 4 °C with 10 mM dithiothreitol (DTT). TOF-LCMS spectra in each column are from identical aliquots sampled at specified time points. The calculated masses of scFv-GGY for disulfide reduction of uncoupled and coupled proteins are 26,337.2 Da and 26,442.2 Da, respectively. Aniline coupling + reduction + DTT = 26,594.45 Da a) abTYR-mediated oxidative coupling with aniline. b) No oxidative coupling reaction.
[0043] Figure 30 A to Figure 30 B shows the stability study of trastuzumab scFv-GGY in protein storage buffer (20 mM Na2HPO4, 150 mM NaCl, containing 15% glycerol, pH 7.4) stored at 4 °C. TOF-LCMS spectra in each column are from identical aliquots sampled at specified time points. The calculated masses of scFv-GGY for disulfide reduction of uncoupled and coupled proteins are 26,337.2 Da and 26,442.2 Da, respectively. a) abTYR-mediated oxidative coupling with aniline and exchange into protein storage buffer. b) No oxidative coupling reaction.
[0044] Figure 31 A to Figure 31 B shows the stability study of trastuzumab scFv-GGY in protein storage buffer (20 mM Na2HPO4, 150 mM NaCl, containing 15% glycerol, pH 7.4) with 10 mM glutathione at 4 °C. TOF-LCMS spectra in each column were obtained from identical aliquots sampled at specified time points. The calculated masses of scFv-GGY for disulfide reduction of uncoupled and coupled proteins were 26,337.2 Da and 26,442.2 Da, respectively. Aniline coupling + reduction + 1x glutathione = 26,747.58 Da; Aniline coupling + reduction + 2x glutathione = 27,052.89 Da. a) abTYR-mediated oxidative coupling with aniline. b) No oxidative coupling reaction.
[0045] Figure 32This study demonstrates the exchange of thiols in the oxidative coupling reaction products. Trastuzumab scFv-GGY was exchanged into protein storage buffer (20 mM Na2HPO4, 150 mM NaCl, containing 15% glycerol, pH 7.4) with 10 mM glutathione and stored at 4°C. After 24 hours, a portion of the sample was analyzed by TOF-LCMS, and the remaining portion was exchanged into protein storage buffer with 10 mM DTT and stored at 4°C for another 24 hours. A second sample was then analyzed by TOF-LCMS.
[0046] Figure 33 The average mass of the protein constructs is shown. The sequences are as follows: GGGGSGGY (SEQ ID NO: 968); (GGGGS)2GGY (SEQ ID NO: 947); (AP)4GGY (SEQ ID NO: 1061); AN 20 GGY (SEQ ID NO: 1045), SSGGGGY (SEQ ID NO: 948), (GGGGS)3GGY (SEQ ID NO: 1042), AEAAAKEAAAKAGGY (SEQ ID NO: 1043), (AP)3GGY (SEQ ID NO: 1044), EIKRTGGY (SEQ ID NO: 1046), GGGGSGGY (SEQ ID NO: 968).
[0047] Figures 34A to 34E The amino acid sequence of the protein construct is provided. Figures 34A to 34E The sequence is shown in SEQ ID NO: 1049-1053.
[0048] Figure 35 The use of the D55K mutant of Bacillus megaterium tyrosinase (bmTYR) to couple phenol-labeled nucleic acids to cysteine-containing proteins was described.
[0049] Figures 36A to 36C The method of this disclosure is described to couple nucleic acids to peptides.
[0050] Figures 37A to 37C The effects of various mutations in bmTYR on its preference for charged substrates were described.
[0051] Figure 38 The study describes the lack of activity of abTYR on activated negatively charged substrates.
[0052] Figures 39A to 39G Protein linkage using the methods of this disclosure is schematically depicted.
[0053] Figures 40A to 40C The stability of the target molecule-biomolecule conjugate in human serum was described.
[0054] Figure 41 The conjugation of Cas9 with the following: i) Ig Fc peptides; ii) and conjugation with nanobodies using the methods of this disclosure.
[0055] Figure 42 Time-of-flight mass spectrometry data of Cas9-nanobody conjugates were depicted.
[0056] Figures 43A to 43B The method described is: i) a method for directly labeling the surface of living mammalian cells. Figure 43A (i) and (ii) using the methods of this disclosure to couple the polypeptide to the cell surface.
[0057] Figures 44A to 44B The reaction in which the target molecule includes two thiol moieties is described.
[0058] definition
[0059] Before further describing the invention, it should be understood that the invention is not limited to the specific embodiments described, and therefore, changes are naturally possible. It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting, as the scope of the invention will be limited only by the appended claims.
[0060] When providing numerical ranges, it should be understood that every intermediate value between the upper and lower limits of the range (unless the context clearly indicates otherwise, the intermediate value is one-tenth of the lower limit unit) and any other specified or intermediate values within the specified range are covered within the invention. The upper and lower limits of these smaller ranges may be independently included within the smaller range and also covered within the invention, conditional on any explicitly excluded limit value within the specified range. When a specified range includes one or two limits, the range excluding any one or both of those included limits is also included in the invention.
[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. While any methods and materials similar or equivalent to those described herein may also be used in the practice or testing of this invention, preferred methods and materials are described hereafter. All publications mentioned herein are incorporated by reference to disclose and describe methods and / or materials in connection with those publications.
[0062] It must be noted that, as used herein and in the appended claims, the singular forms “a / an” and “the” include plural references unless the context clearly indicates otherwise. Thus, for example, reference to “thiol group” includes a plurality of such thiol groups and reference to “the thiol group” includes reference to one or more thiol groups and their equivalents known to those skilled in the art, etc. It should also be noted that the claims may be drafted to exclude any optional elements. Therefore, such statements are intended to serve as a precondition for using exclusive terms such as “merely” or “only” or for using negative limitations in conjunction with the elements of the claims.
[0063] It should be understood that certain features of the invention described in the context of individual embodiments for clarity may also be provided in combination in a single embodiment. Conversely, for brevity, various features of the invention described in the context of individual embodiments may also be provided individually or in any suitable sub-combination. The invention particularly covers all combinations of embodiments relating to the invention, and is disclosed herein as if each combination were disclosed individually and explicitly. Furthermore, the invention also particularly covers all sub-combinations of various embodiments and their elements, and is disclosed herein as if each such sub-combination were disclosed individually and explicitly herein.
[0064] The publications discussed herein provide only their disclosure prior to the filing date of this application. Nothing herein should be construed as an admission that the invention is not entitled to precede such publications by virtue of a prior invention. Furthermore, the publication dates provided may differ from the actual publication dates, which may require independent verification.
[0065] As used herein, the term "affinity tag" refers to a member of a specific binding pair, i.e., two molecules, one of which binds specifically to the other molecule by a chemical or physical means. The complementary member of the affinity tag can be immobilized (e.g., onto a chromatographic carrier, beads, or a flat surface) to produce an affinity chromatographic carrier that specifically binds to the affinity tag. Tagging the compound of interest with an affinity tag allows the separation of the compound from a mixture of unlabeled compounds by affinity, for example, using affinity chromatography. Examples of specific binding pairs include biotin and streptavidin (or avidin), as well as antigens and antibodies, although binding pairs such as nucleic acid hybrids, multihistidines and nickel, and azide and alkynyl groups (e.g., cyclooctynyl) or phosphine groups are also contemplated. Specific binding pairs may include analogues, derivatives, and fragments of the original specific binding member.
[0066] As used herein, the term "biotin moiety" refers to an affinity label that includes biotin or biotin analogues such as desulfobiotin, oxybiotin, 2'-iminobiotin, diaminobiotin, biotin sulfoxide, biocytotin, etc. The biotin moiety is defined with at least 10 -8M binds to streptavidin with its affinity. The biotin moiety may also include a linker, such as -LC-biotin, -LC-LC-biotin, -SLC-biotin, or -PEG. n 1 -Biotin, where n 1 It is 3-12.
[0067] The term "link" or "connector" in terms such as "linking group" or "joint part" refers to a linking part that connects two groups via covalent bonds. A linker can be straight-chain, branched, cyclic, or a single atom. Examples of such linking groups include alkyl, alkenyl, ynyl, aryl, alkylaryl, arylalkylene, and linking parts containing functional groups, including but not limited to: amide (-NH-CO-), ureidyl (-NH-CO-NH-), imide (-CO-NH-CO-), epoxy (-O-), cyclic sulfur (-S-), cyclic dioxy (-OO-), cyclic disulfide (-SS-), carbonyl dioxy (-O-CO-O-), alkyl dioxy (-O-(CH2)nO-), epoxy imino (-O-NH-), cyclic imino (-NH-), carbonyl (-CO-), etc. In some cases, one, two, three, four, or five or more carbon atoms in the linker backbone may optionally be substituted with sulfur, nitrogen, or oxygen heteroatoms. The bonds between the main chain atoms can be saturated or unsaturated, and typically no more than one, two, or three unsaturated bonds will exist in the linker main chain. The linker may contain one or more substituents, such as alkyl, aryl, or alkenyl groups. The linker may include, but is not limited to, poly(ethylene glycol) units (e.g., -(CH2-CH2-O)-); ethers, thioethers, amines, alkyl groups (e.g., (C1-C2-O)-); and other substituents. 12 The linker (alkyl group) can be straight-chain or branched, such as methyl, ethyl, n-propyl, 1-methylethyl (isopropyl), n-butyl, n-pentyl, 1,1-dimethylethyl (tert-butyl), etc. The linker backbone may include cyclic groups, such as aryl, heterocyclic, or cycloalkyl groups, wherein the backbone contains two or more atoms of the cyclic group, for example, two, three, or four atoms. The linker can be cleavable or non-cleavable. The linker can be used in any convenient orientation and / or connection with the linking group.
[0068] "Alkyl" refers to a monovalent saturated aliphatic hydrocarbon group having 1 to 10 carbon atoms, for example 1 to 6 carbon atoms. The term includes, for example, straight-chain and branched hydrocarbon groups, such as methyl (CH3-), ethyl (CH3CH2-), n-propyl (CH3CH2CH2-), isopropyl ((CH3)2CH-), n-butyl (CH3CH2CH2CH2-), isobutyl ((CH3)2CHCH2-), sec-butyl ((CH3)(CH3CH2)CH-), tert-butyl ((CH3)3C-), n-pentyl (CH3CH2CH2CH2CH2-), and neopentyl ((CH3)3CCH2-).
[0069] The term "substituted alkyl" refers to an alkyl group as defined above, wherein one or more carbon atoms (other than the C1 carbon atom) in the alkyl chain are optionally replaced by heteroatoms such as -O-, -N-, -S-, -S(O). n 2 - (where n) 2 It is 0 to 2) substituted with -NR- (where R is hydrogen or alkyl) and has 1 to 5 substituents selected from the group consisting of: alkoxy, substituted alkoxy, cycloalkyl, substituted cycloalkyl, cycloalkenyl, substituted cycloalkenyl, acyl, acylamino, acyloxy, amino, aminoacyl, aminoacyloxy, oxyaminoacyl, azide, cyano, halogen, hydroxy, oxo, thioketone, carboxyl, carboxylalkyl, thioaryloxy, thioheteroaryloxy, thioheterocyclicoxy, mercapto, thioalkoxy, substituted thioalkoxy, aryl, aryloxy, heteroaryl, heteroaryloxy, heterocyclic, heterocyclic, hydroxyamino, alkoxyamino, nitro, -SO-alkyl, -SO-aryl, -SO-heteroaryl, -SO2-alkyl, -SO2-aryl, -SO2-heteroaryl and -NR a R b , where R ’ and R ” They may be the same or different, and are selected from hydrogen, optionally substituted alkyl, cycloalkyl, alkenyl, cycloalkenyl, alkynyl, aryl, heteroaryl and heterocyclic groups.
[0070] "Aryl" or "Ar" refers to a monovalent aromatic carbocyclic group having 6 to 18 carbon atoms in a ring system having a single ring (as present in phenyl) or multiple fused rings (examples of such aromatic ring systems include naphthyl, anthracene, and indene). The fused rings may or may not be aromatic, provided that the bonding point passes through an atom of the aromatic ring. This term includes, for example, phenyl and naphthyl. Unless otherwise limited by the definition of aryl substituents, such aryl groups may optionally be substituted with 1 to 5 or 1 to 3 substituents selected from acyloxy, hydroxyl, mercapto, acyl, alkyl, alkoxy, alkenyl, alkynyl, cycloalkyl, cycloalkenyl, substituted alkyl, substituted alkoxy, substituted alkenyl, substituted alkynyl, substituted cycloalkyl, substituted cycloalkenyl, amino, substituted amino, aminoacyl, acylamino, alkylaryl, aryl, aryloxy, Azide, carboxyl, carboxylalkyl, cyano, halogen, nitro, heteroaryl, heteroaryloxy, heterocyclic, heterocyclic, aminoacyloxy, oxyacylamino, thioalkoxy, substituted thioalkoxy, thioaryloxy, thioheteroaryloxy, -SO-alkyl, -SO-substituted alkyl, -SO-aryl, -SO-heteroaryl, -SO2-alkyl, -SO2-substituted alkyl, -SO2-aryl, -SO2-heteroaryl and trihalomethyl.
[0071] "Amino" refers to the group –NH2.
[0072] The term “substituted amino” refers to a group -NRR, wherein each R is independently selected from the group consisting of: hydrogen, alkyl, substituted alkyl, cycloalkyl, substituted cycloalkyl, alkenyl, substituted alkenyl, cycloalkenyl, substituted cycloalkenyl, alkynyl, substituted alkynyl, aryl, heteroaryl and heterocyclic, provided that at least one R is not hydrogen.
[0073] In addition to the contents disclosed herein, the term “substituted” when used to modify a specified group (radical) may also mean that one or more hydrogen atoms of the specified group are each independently substituted by the same or different substituents as defined below.
[0074] Except for the groups disclosed herein with respect to various terms, unless otherwise stated, they are used to replace one or more hydrogen atoms on a saturated carbon atom in a specified group (any two hydrogen atoms on a single carbon atom can be =O, =NR). 70 =N-OR 70 The substituent for (=N2 or =S substitution) is -R. 60 , halogenated group, =O, -OR 70 -SR 70 -NR 80 R 80 Trihalomethyl, -CN, -OCN, -SCN, -NO, -NO2, =N2, -N3, -SO2R70 、 -SO2O – M + 、 -SO2OR 70 、 -OSO2R 70 、 -OSO2O – M + 、 -OSO2OR 70 、 -P(O)(O – )2(M + )2、 -P(O)(OR 70 )O – M + 、 -P(O)(OR 70 )2、 -C(O)R 70 、 -C(S)R 70 、 -C(NR 70 )R 70 、 -C(O)O – M + 、 -C(O)OR 70 、 -C(S)OR 70 、 -C(O)NR 80 R 80 、 -C(NR 70 )NR 80 R<000Choose from the group consisting of: optionally substituted alkyl, cycloalkyl, heteroalkyl, heterocycloalkylalkyl, cycloalkylalkyl, aryl, arylalkyl, heteroaryl, and heteroarylalkyl, each R 70 Independently hydrogen or R 60 ; Each R 80 R is independent 70 Or choose two other locations R 80 Together with the nitrogen atom it is bonded to, it forms a 5-, 6-, or 7-membered heterocyclic alkyl group, which may optionally comprise 1 to 4 identical or different heteroatoms selected from the group consisting of O, N, and S, wherein N may have -H or C1-C3 alkyl substitution; and each M + It is a counterion with a net single positive charge. Each M + It can be an independent ion, such as a base ion, like K+. + Na + Li + Ammonium ions, such as + N(R 60 )4; or alkaline earth metal ions, such as [Ca 2 + ] 0.5 、[Mg 2+ ] 0.5 or[Ba 2+ ] 0.5 ("Subscript 0.5" means that one of the counter ions of such divalent alkaline earth metal ions can be the ionized form of the compound of the present invention and the other is a typical counter ion, such as chloride ion, or that the two ionized compounds disclosed herein can act as counter ions of such divalent alkaline earth metal ions, or that the dual ionized compound of the present invention can act as a counter ion of such divalent alkaline earth metal ions.) As a specific example, -NR 80 R 80 This refers to compounds including -NH2, -NH-alkyl, N-pyrrolidinyl, N-piperazinyl, 4N-methyl-piperazin-1-yl, and N-morpholinyl.
[0075] Except as otherwise stated herein, the substituent for hydrogen on the unsaturated carbon atom in “substituted” alkenes, alkynes, aryls, and heteroaryls is -R. 60 , halogenated group, -O - M + -OR 70 -SR 70 -S – M + -NR 80 R 80 Trihalomethyl, -CF3, -CN, -OCN, -SCN, -NO, -NO2, -N3, -SO2R 70 -SO3 –M + 、 -SO3R 70 、 -OSO2R 70 、 -OSO3 – M + 、 -OSO3R 70 、 -PO3 -2 (M + )2、 -P(O)(OR 70 )O – M + 、 -P(O)(OR 70 )2、 -C(O)R 70 、 -C(S)R 70 、 -C(NR 70 )R 70 、 -CO2 – M + 、 -CO2R 70 、 -C(S)OR 70 、 -C(O)NR 80 R 80 、 -C(NR 70 )NR 80 R 80 、 -OC(O)R 70 、 -OC(S)R 70 、 -OCO2 – M + 、 -OCO2R 70 、 -OC(S)OR 70 、 -NR 70 C(O)R 70 、 -NR 70 C(S)R 70 、 -NR 70 CO2 – M + 、 -NR 70 CO2R 70 、 -NR 70 C(S)OR 70 、 -NR 70 C(O)NR 80 R 80 、 -NR 70 C(NR 70 )R 70 和 -NR 70 C(NR 70 )NR 80 R 80 ,其中R 60 、R 70 、R 80 和M +As defined above, the condition is that in the case of substituted alkenes or alkynes, the substituent is not -O. - M + -OR 70 -SR 70 or -S – M + .
[0076] Except for the groups disclosed with respect to various terms herein, unless otherwise stated, the substituent for hydrogen on the nitrogen atom in “substituted” heteroalkyl and cycloalkyl groups is -R. 60 -O - M + -OR 70 -SR 70 -S - M + -NR 80 R 80 Trihalomethyl, -CF3, -CN, -NO, -NO2, -S(O)2R 70 -S(O)2O - M + -S(O)2OR 70 -OS(O)2R 70 -OS(O)2O - M + -OS(O)2OR 70 -P(O)(O) - )2(M + )2、-P(O)(OR 70 )O - M + -P(O)(OR) 70 (OR) 70 -C(O)R 70 -C(S)R 70 -C(NR) 70 )R 70 -C(O)OR 70 -C(S)OR 70 -C(O)NR 80 R 80 -C(NR) 70 )NR 80 R 80 -OC(O)R 70 -OC(S)R 70 -OC(O)OR 70 -OC(S)OR 70 -NR 70 C(O)R 70 -NR 70 C(S)R70 -NR 70 C(O)OR 70 -NR 70 C(S)OR 70 -NR 70 C(O)NR 80 R 80 -NR 70 C(NR 70 )R 70 and -NR 70 C(NR 70 )NR 80 R 80 , where R 60 R 70 R 80 and M + As defined above.
[0077] In addition to the disclosure herein, in some embodiments, the substituted group has 1, 2, 3 or 4 substituents, 1, 2 or 3 substituents, 1 or 2 substituents, or 1 substituent.
[0078] It is understood that polymers obtained by defining substituents with other substituents against themselves (e.g., a substituted aryl group having a substituted aryl group that is itself substituted as a substituent, which is further substituted by another substituted aryl group, etc.) are not intended to be included herein. In this case, the maximum number of such substitutions is three. For example, the successive substitutions of substituted aryl groups specifically considered herein are limited to substituted aryl-(substituted aryl)-substituted aryl.
[0079] With respect to any group disclosed herein containing one or more substituents, it should be understood that such groups do not include any substitution or substitution pattern that is sterically impractical and / or synthetically infeasible. Furthermore, the subject compounds include all stereochemical isomers resulting from the substitution of these compounds.
[0080] In some embodiments, substituents may contribute to the optical and / or stereoisomerization of the compound. Salt, solvate, hydrate, and prodrug forms of the compound are also of interest. This disclosure includes all such forms. Therefore, the compounds described herein include their salt, solvate, hydrate, prodrug, and isomer forms, including pharmaceutically acceptable salts, solvates, hydrates, prodrugs, and isomers. In some embodiments, the compound can be metabolized into a pharmaceutically active derivative.
[0081] Unless otherwise stated, references to an atom refer to isotopes that include that atom. For example, a reference to H refers to isotopes that include H. 1 H, 2 H (i.e., D) and3 H (i.e., T), and mentioning C refers to including 12 All isotopes of C and carbon (e.g.) 13 C).
[0082] As used herein, the terms “cleavable linker” or “cleavable joint” refer to a linker or bond that can be selectively broken by a stimulus (e.g., physical, chemical, or enzymatic stimulation) that leaves the bonded portion intact. Several cleavable bonds have been described in the literature (e.g., Brown (1997) Contemporary Organic Synthesis 4(3); 216-237) and Guillier et al. (Chem. Rev. 2000 1000:2091-2157). Disulfide bonds (which can be broken by DDT) and photocleavable linkers are examples of cleavable bonds.
[0083] The term "fluorophore" refers to any molecular entity capable of absorbing energy at a first wavelength and re-emitting energy at a different second wavelength. In some embodiments, the subject biomolecule includes a fluorophore attached to one end or the central location of the biomolecule. In some embodiments, the fluorophore may be attached to one end of the biomolecule. The fluorophore attached to the biomolecule need not be a single molecule, but may include multiple molecules.
[0084] As those skilled in the art will know, fluorophores can be synthetic or biological in nature. More generally, any fluorophore that is stable under coupling conditions and can be adequately suppressed when in close proximity to a quencher can be used, such that a significant change in the fluorescence intensity of the fluorophore in response to a target specifically bound to the probe is detectable. Examples of suitable fluorophores include, but are not limited to, Oregon Green 488 dye, rhodamine and rhodamine derivatives, fluorescein isothiocyanate, fluorescein, 6-carboxyfluorescein (6-FAM), coumarin and coumarin derivatives, anthocyanins and anthocyanin derivatives, Alexa Fluors, DyLightFluors, etc.
[0085] In some embodiments, the biomolecule includes a metal chelating agent. As used herein, a "chelate" relating to a complex between a metal and a chelating ligand refers to a combination of metal ions bonded to one or more ligands to form a heterocyclic structure. The formation of a chelate by neutralizing the positive charge of a metal ion can be achieved through ionic bonds, covalent bonds, or coordinate covalent bonds. In some embodiments, the metal chelating agent includes, but is not limited to, 1,4,7,10-tetraazacyclododecane-1,4,7,10-tetraacetic acid (also known as DOTA or tetraxetan).
[0086] The terms “polynucleotide” and “nucleic acid”, used interchangeably herein, refer to polymeric forms of nucleotides (ribonucleotides or deoxyribonucleotides) of any length. Therefore, the term includes, but is not limited to, single-stranded, double-stranded, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers containing purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derived nucleotide bases.
[0087] The terms “peptide” and “protein”, used interchangeably herein, refer to a polymer of amino acids of any length, which may include coding and non-coding amino acids, chemically or biochemically modified or derived amino acids, and peptides having a modified peptide backbone. The term “fusion protein” or its grammatical equivalent refers to a protein comprising multiple polypeptide components that are typically not linked in their native state, but are usually linked by peptide bonds from their respective amino and carboxyl ends to form a single, continuous polypeptide. Fusion proteins can be combinations of two, three, or even four or more different proteins.
[0088] Generally, polypeptides can have any length, such as 2 or more amino acids, more than 4 amino acids, more than about 10 amino acids, more than about 20 amino acids, more than about 50 amino acids, more than about 100 amino acids, more than about 300 amino acids, and typically up to about 500 or 1000 or more amino acids. "Peptides" are typically 2 or more amino acids long, such as more than 4 amino acids, more than about 10 amino acids, more than about 20 amino acids, and typically up to about 50 amino acids. In some embodiments, peptides are 2 to 30 amino acids long.
[0089] As used herein, the term "target protein" refers to all members of the target family, their fragments and enantiomers, and their protein mimics. Unless otherwise explicitly stated, the target protein of interest described herein is intended to include all members of the target family, their fragments and enantiomers, and their protein mimics. Target proteins can be any protein of interest, such as therapeutic or diagnostic targets, including but not limited to: hormones, growth factors, receptors, enzymes, cytokines, osteoinducible factors, colony-stimulating factors, and immunoglobulins. The term "target protein" is intended to include recombinant and synthetic molecules that can be prepared using any convenient recombinant expression method or any convenient synthetic method, or commercially available, as well as fusion proteins containing target molecules.
[0090] The term "physiological conditions" refers to those conditions that are compatible with living cells, such as the main aqueous conditions that are compatible with living cells, such as temperature, pH, salinity, etc.
[0091] The terms "solid support," "support," and "solid phase carrier" are used interchangeably and refer to a material or group of materials having one or more rigid or semi-rigid surfaces. In many embodiments, at least one surface of the solid support will be substantially flat, although in some embodiments it may be desirable to physically separate the synthesis regions of different compounds using, for example, pores, raised regions, needles, etched trenches, etc. According to other embodiments, the solid support will take the form of beads, resin, gel, microspheres, or other geometries.
[0092] The terms “antibody” and “immunoglobulin” include any isotype of antibody or immunoglobulin, antibody fragments that retain specific binding to antigens, including but not limited to Fab, Fv, scFv, and Fd fragments, chimeric antibodies, humanized antibodies, single-chain antibodies (scAbs), single-domain antibodies (dAbs), single-domain heavy-chain antibodies, single-domain light-chain antibodies, nanobodies, bispecific antibodies, multispecific antibodies, and fusion proteins comprising the antigen-binding (also referred to herein as antigen-binding) portion of an antibody and a non-antibody protein. Antibodies can be detectably labeled, for example with radioactive isotopes, enzymes that produce detectable products, fluorescent proteins, etc. Antibodies can be further conjugated to other parts, such as members of specific binding pairs, such as biotin (a member of the biotin-antibiotin protein specific binding pair), etc. Antibodies can also be bound to solid carriers, including but not limited to polystyrene plates or beads. The term also covers Fab', Fv, F(ab')2, and / or other antibody fragments that retain specific binding to antigens, as well as monoclonal antibodies. As used herein, a monoclonal antibody is an antibody produced by a group of identical cells, all of which are generated from a single cell through repeated cell replication. That is, the cell clone produces only a single antibody species. While monoclonal antibodies can be produced using hybridoma production techniques, other production methods known to those skilled in the art can also be used (e.g., antibodies derived from an antibody phage display library). Antibodies can be monovalent or bivalent. Antibodies can be Ig monomers, which are “Y-shaped” molecules composed of four polypeptide chains: two heavy chains and two light chains linked by disulfide bonds.
[0093] As used herein, the term "humanized immunoglobulin" refers to an immunoglobulin comprising immunoglobulin motifs from different sources, wherein at least a portion contains an amino acid sequence of human origin. For example, a humanized antibody may comprise motifs derived from a non-human immunoglobulin with desired specificity (e.g., mouse) and a human immunoglobulin sequence (e.g., chimeric immunoglobulin), chemically linked together by conventional techniques (e.g., synthesis) or prepared as a continuous polypeptide using genetic engineering techniques (e.g., DNA encoding the protein motif of a chimeric antibody may be expressed to produce a continuous polypeptide chain). Another example of a humanized immunoglobulin is an immunoglobulin containing one or more immunoglobulin chains comprising a complementarity-determining region (CDR) of an antibody derived from a non-human source and a skeletal region of a light chain and / or heavy chain derived from a human source (e.g., CDR-grafted antibodies with or without skeletal alterations). The term humanized immunoglobulin also encompasses chimeric or CDR-grafted single-chain antibodies. See, for example, Cabilly et al., U.S. Patent No. 4,816,567; Cabilly et al., European Patent No. 0,125,023 B1; Boss et al., U.S. Patent No. 4,816,397; Boss et al., European Patent No. 0,120,694 B1; Neuberger, MS et al., WO 86 / 01533; Neuberger, MS et al., European Patent No. 0,194,276 B1; Winter, U.S. Patent No. 5,225,539; Winter, European Patent No. 0,239,400 B1; Padlan, EA et al., European Patent Application No. 0,519,596 A1. For more information on single-chain antibodies, see Ladner et al., U.S. Patent No. 4,946,778; Huston, U.S. Patent No. 5,476,786; and Bird, RE et al., Science, 242: 423-426 (1988)).
[0094] As used herein, the term "nanobody" (Nb) refers to the smallest antigen-binding fragment or single variable domain (V domain) derived from naturally occurring heavy chain antibodies. HHThese are known to those skilled in the art. They are derived from heavy-chain-only antibodies found in camelids (Hamers-Casterman et al., (1993) Nature 363:446; Desmyter et al., (1996) Nature Struct. Biol. 3:803). Immunoglobulins without light polypeptide chains have been found in the "camelidae" family. "Camelidae" includes Old World camels (Bactrian and Dromedary camels) and New World camels (e.g., alpaca (Llama paccos), llama (Llama glama), guanoa (Llama guanicoe), and llama (Llamavicugna)). Single variable domain heavy-chain antibodies are referred to herein as nanobodies or V. HH Antibody.
[0095] "Antibody fragments" include portions of a complete antibody, such as the antigen-binding region or variable region of a complete antibody. Examples of antibody fragments include Fab, Fab', F(ab')2, and Fv fragments; biantibodies; linear antibodies (Zapata et al., ProteinEng. 8(10): 1057-1062 (1995)); domain antibodies (dAb; Holt et al. (2003) TrendsBiotechnol. 21:484); single-chain antibody molecules; and multispecific antibodies formed from antibody fragments. Papain digestion of antibodies produces two identical antigen-binding fragments, called "Fab" fragments, each with a single antigen-binding site; and a residual "Fc" fragment, a name reflecting its tendency to crystallize. Pepsin treatment produces F(ab')2 fragments with two antigen-binding sites and still capable of cross-linking antigens.
[0096] "Fv" is the smallest antibody fragment containing both a complete antigen recognition site and a binding site. This region consists of a tightly bound, non-covalently associated dimer of a heavy-chain variable domain and a light-chain variable domain. It is in this configuration that the three CDRS of each variable domain interact to achieve the V... H -V L The surface of the dimer defines the antigen-binding site. The six CDRs together confer antibody antigen-binding specificity. However, even a single variable domain (or half of the Fv containing only three antigen-specific CDRs) has the ability to recognize and bind to the antigen, although with lower affinity than the entire binding site.
[0097] The “Fab” fragment also contains a constant domain of the light chain and a first constant domain (CH1) of the heavy chain. The Fab fragment differs from the Fab' fragment in that it has several residues added to the carboxyl terminus of the CH1 domain of the heavy chain, including one or more cysteine residues from the antibody hinge region. Fab'-SH is referred to herein as Fab', where the cysteine residues of the constant domain are accompanied by free thiol groups. The F(ab')2 antibody fragment is initially generated as a Fab' fragment pair, which has a hinge cysteine residue between the Fab' fragments. Other chemical conjugates of the antibody fragment are also known.
[0098] The "light chain" of antibodies (immunoglobulins) from any vertebrate species can be classified into one of two distinct types, called κ and λ, based on the amino acid sequence of their constant domain. Immunoglobulins can be classified into different categories based on the amino acid sequence of their heavy chain constant domain. There are five main classes of immunoglobulins: IgA, IgD, IgE, IgG, and IgM, and several of these classes can be further subdivided into subclasses (isotypes), such as IgG1, IgG2, IgG3, IgG4, IgA, and IgA2. These subclasses can be further subdivided into types, such as IgG2a and IgG2b.
[0099] A single-chain Fv, sFv, or scFv antibody fragment contains the antibody's V. H and V L Domains, wherein these domains are present within a single polypeptide chain. In some embodiments, the Fv polypeptide also contains V H With V L The domains contain peptide linkers that enable sFv to form the structures required for antigen binding. For a review of sFv, see Pluckthun in The Pharmacology of Monoclonal Antibodies, Vol. 113, Rosenburg and Moore eds., Springer-Verlag, New York, pp. 269-315 (1994).
[0100] The term "dual antibody" refers to a small antibody fragment having two antigen-binding sites, said fragment containing antibodies from the same polypeptide chain (V). H -V L The light chain variable structural domain (V) in ) L ) connected heavy chain variable structural domain (V HBy using a linker that is too short to allow pairing between two domains on the same chain, the domain is forced to pair with a complementary domain of another chain, creating two antigen-binding sites. Biantibodies are well described, for example, in EP 404,097; WO 93 / 11161; and Hollinger et al. (1993) Proc. Natl. Acad. Sci. USA 90:6444-6448.
[0101] As used herein, the term "affinity" refers to the equilibrium constant of the reversible binding of two agents (e.g., antibody and antigen), and is expressed as the dissociation constant (K). D The affinity can be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or 1,000 times greater than the antibody's affinity for unrelated amino acid sequences. The antibody's affinity for the target protein can be, for example, from about 100 nanomolars (nM) to about 0.1 nM, from about 100 nM to about 1 picomolar (pM), or from about 100 nM to about 1 femtomolar (fM) or higher. As used herein, the term "affinity" refers to the resistance of a complex of two or more agents to dissociation upon dilution. The terms “immunoreactive” and “preferential binding” are used interchangeably in this document in relation to antibody and / or antigen-binding fragments.
[0102] The term "binding" refers to the direct association between two molecules due to interactions such as covalent, electrostatic, hydrophobic, and ionic and / or hydrogen bonding (including interactions such as salt bridges and water bridges). "Specific binding" refers to a binding at least about 10... -7 M or larger, for example, 5x10 -7 M, 10 -8 M, 5 x 10 -8 M binds with greater affinity. "Non-specific binding" refers to binding with an affinity less than approximately 10. -7 The affinity of M, for example, with 10 -6 M, 10 -5 M, 10 -4 Affinity binding of M, etc.
[0103] "Isolated" polypeptides are polypeptides that have been identified, isolated, and / or recovered from components of their native environment. Contaminating components of their native environment are substances that could interfere with the diagnostic or therapeutic use of the polypeptide and may include enzymes, hormones, and other protein or non-protein solutes. In some embodiments, the polypeptide is purified to (1) greater than 90% by weight, greater than 95% by weight, or greater than 98% by weight, as determined by the Lowry method, for example, greater than 99% by weight; (2) purified to a degree sufficient to obtain at least 15 residues of the N-terminal or internal amino acid sequence using a rotary cup sequencer; or (3) purified to homogeneity by sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) under reducing or non-reducing conditions using Coomassie blue or silver staining. Isolated polypeptides include polypeptides in situ in recombinant cells, since at least one component of the polypeptide's native environment will be absent. In some cases, isolated polypeptides are prepared by at least one purification step.
[0104] As will be apparent to those skilled in the art upon reading this disclosure, each individual embodiment described and illustrated herein has discrete components and functions that can be readily separated from or combined with the functions of any other several embodiments without departing from the scope or spirit of the invention. Any stated method may be performed in the order of the stated events or in any other logically possible order.
[0105] When providing numerical ranges, it should be understood that every intermediate value between the upper and lower limits of the range (unless the context clearly indicates otherwise, the intermediate value is one-tenth of the lower limit unit) and any other specified or intermediate values within the specified range are covered within the invention. The upper and lower limits of these smaller ranges may be independently included within the smaller range and also covered within the invention, conditional on any explicitly excluded limit value within the specified range. When a specified range includes one or two limits, the range excluding any one or both of those included limits is also included in the invention.
[0106] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. While any methods and materials similar or equivalent to those described herein may also be used in the practice or testing of this invention, preferred methods and materials are described hereafter. All publications mentioned herein are incorporated by reference to disclose and describe methods and / or materials in connection with those publications.
[0107] It must be noted that, as used herein and in the appended claims, the singular forms “a / an” and “the” include plural references unless the context clearly indicates otherwise. Thus, for example, reference to “target molecule” includes a plurality of such target molecules, and reference to “biomolecule” includes reference to one or more biomolecules and their equivalents known to those skilled in the art. It should also be noted that the claims may be drafted to exclude any optional elements. Therefore, such statements are intended to serve as a precondition for using exclusive terms such as “merely” or “only” or for using negative limitations in conjunction with the elements of the claims.
[0108] It should be understood that certain features of the invention described in the context of individual embodiments for clarity may also be provided in combination in a single embodiment. Conversely, for brevity, various features of the invention described in the context of individual embodiments may also be provided individually or in any suitable sub-combination. The invention particularly covers all combinations of embodiments relating to the invention, and is disclosed herein as if each combination were disclosed individually and explicitly. Furthermore, the invention also particularly covers all sub-combinations of various embodiments and their elements, and is disclosed herein as if each such sub-combination were disclosed individually and explicitly herein.
[0109] The publications discussed herein provide only their disclosure prior to the filing date of this application. Nothing herein should be construed as an admission that the invention is not entitled to precede such publications by virtue of a prior invention. Furthermore, the publication dates provided may differ from the actual publication dates, which may require independent verification. Detailed Implementation
[0110] This disclosure provides a method for chemically selectively modifying a target molecule. The disclosure provides compositions comprising a subject target molecule and a biomolecule, the subject target molecule comprising a thiol moiety and the biomolecule comprising a phenolic or catechol moiety. The disclosure provides a kit comprising a first container of the subject composition and a second container comprising an enzyme capable of oxidizing the phenolic or catechol moiety. The disclosure also provides modified target molecules that can be used for the delivery of biomolecules for gene therapy, novel immunotherapies via antibody conjugates, biomaterial construction, and vaccine development.
[0111] method
[0112] As summarized above, aspects of this disclosure include methods for the chemoselective modification of target molecules. The subject method involves contacting a target molecule comprising a thiol moiety with a biomolecule comprising a reactive moiety, wherein the reactive moiety is generated by reacting the biomolecule comprising a phenolic or catechol moiety with an enzyme capable of oxidizing the phenolic or catechol moiety. The contact is carried out under conditions sufficient to conjugate the target molecule to the biomolecule, thereby producing a modified target molecule.
[0113] In some cases, the subject method for chemoselectively modifying a target molecule includes contacting: i) a target molecule containing a thiol moiety; ii) a biomolecule containing a phenolic or catechol moiety; and iii) an enzyme capable of oxidizing the phenolic or catechol moiety; wherein the enzyme oxidizes the phenolic or catechol moiety of the biomolecule to produce a reactive moiety, thereby producing a biomolecule containing the reactive moiety, and wherein the reactive moiety reacts with the thiol moiety, thereby conjugating the target molecule and the biomolecule to each other to produce a modified target molecule. In some cases, the target molecule contains a single thiol moiety. In some cases, the target molecule contains two thiol moieties.
[0114] Target molecules can be any of a variety of molecules (e.g., peptides; nucleic acids; small molecules; etc.). Similarly, biomolecules can be any of a variety of molecules (e.g., peptides; nucleic acids; small molecules; etc.). In some cases, the target molecule is a peptide; and the biomolecule is a nucleic acid. In some cases, the target molecule is a nucleic acid; and the biomolecule is a peptide. In some cases, the target molecule is a peptide; and the biomolecule is a small molecule (e.g., a cancer chemotherapy agent). In some cases, the target molecule is a first peptide; and the biomolecule is a second peptide, wherein the first peptide and the second peptide can be the same or different.
[0115] The subject-matter approach provides a simple conjugation procedure capable of linking a biomolecule of interest to any site on the surface of a target molecule in a site-specific manner, thereby producing a modified target molecule of interest. In some embodiments, the target molecule is a second biomolecule (e.g., as described herein). In some embodiments, the second biomolecule is a peptide.
[0116] Biomolecules of interest include, but are not limited to, peptides, polynucleotides, carbohydrates, lipids, fatty acids, steroids, purines, pyrimidines, their derivatives, structural analogs, and combinations thereof. In some cases, the biomolecule of interest is an antibody. In some cases, the biomolecule of interest is an antibody fragment or a conjugated derivative thereof. In some cases, the antibody fragment or its conjugated derivative is selected from the group consisting of: Fab fragments, F(ab')2 fragments, single-chain Fv (scFv), biantibodies, nanobodies, and triantibodies. Suitable biomolecules include, for example, small molecules (e.g., cancer chemotherapeutic agents), cytokines, hormones, immunomodulatory peptides, etc. In some cases, the biomolecule is a nucleic acid; and the target molecule is an antibody (e.g., scFv; nanobodies; etc.). In some cases, the biomolecule is a small molecule (e.g., cancer chemotherapeutic agents); and the target molecule is an antibody (e.g., scFv; nanobodies; etc.). In some cases, for example, when the target molecule is an antibody, the biomolecule is linked to the Fc portion of the target molecule. In some cases, the target molecule is an immunoglobulin (Ig) Fc peptide.
[0117] In some embodiments, the biomolecule containing a phenolic or catechol moiety also includes one or more moieties selected from: fluorophores, active small molecules, affinity tags, and metal chelators (e.g., as described herein). In some cases, the biomolecule of interest is a fluorescent protein. In some cases, the fluorescent protein is green fluorescent protein (GFP). In some cases, the biomolecule is an enzyme. In some cases, the biomolecule is a ligand for a receptor. In some cases, the biomolecule is a receptor.
[0118] In some embodiments, the enzyme capable of oxidizing the phenolic or catechol moiety is a phenol oxidase or a catechol oxidase. In some cases, the enzyme is a tyrosinase.
[0119] The term "tyrosinase" as used in this article refers to monophenol monooxygenase (EC 1.14.18.1; CAS No.: 9002-10-2), an enzyme that catalyzes the oxidation of phenols (such as tyrosine). It is a copper-containing enzyme found in plant and animal tissues that catalyzes the production of melanin and other pigments from tyrosine through oxidation. All tyrosinases share a common binuclear type 3 copper center within their active site. Here, two copper atoms are each coordinated to three histidine residues. Matoba et al., "Crystallographic evidence that the dinuclear copper center of tyrosinase is flexible during catalysis," J Biol Chem. 2006, March 31; 281(13):8981-90. A three-dimensional model of the tyrosinase catalytic center was published in Epub on January 25, 2006.
[0120] In some embodiments, the subject enzyme is attached to a solid support system, such as beads, resin, gel, microspheres, or other geometries. In some cases, the solid support is glass beads. In some cases, the solid support is resin beads. Using an enzyme attached to a solid support system allows for a variety of uses of the subject enzyme and facilitates the purification of the subject target molecule by allowing the enzyme to be easily removed from the reaction mixture. In some embodiments, the subject enzyme attached to a solid support system allows the subject method to be carried out in a continuous flow system. In some embodiments, the subject enzyme attached to a solid support system facilitates high-volume processing of the subject method.
[0121] In some cases, the phenolic component is present in the tyrosine residue. In some cases, the tyrosine residue is part of the biomolecule of interest. In some cases, the tyrosine residue is synthesized and introduced into the biomolecule of interest. In some other cases, the tyrosine residue is linked to the biomolecule of interest via a linker (e.g., as described herein). Tyrosine residues can be introduced into peptide biomolecules using standard recombination techniques, such as by modifying the nucleotide sequence encoding the peptide biomolecule.
[0122] In some cases, the phenolic or catechol moiety is part of a non-natural (non-genetically encoded) amino acid to be introduced into a biomolecule of interest. For example, amber codon (TAG) inhibition can be used to incorporate non-genetically encoded amino acid residues containing a phenolic or catechol moiety. See, for example, Chin et al. (2002) J. Am. Chem. Soc. 124:9026; Chin and Schultz (2002) Chem. Biol. Chem. 3:1135; Chin et al. (2002) Proc. Natl. Acad. Sci.USA 99:11020; US 2015 / 0240249; and US 2018 / 0171321. As another example, orthogonal RNA synthases and / or orthogonal tRNAs can be used to introduce non-genetically encoded amino acids into biomolecules, wherein the non-genetically encoded amino acid contains a phenolic or catechol moiety.
[0123] In some embodiments of the subject method, the thiol moiety present in the target molecule is part of a cysteine residue. In some cases, the cysteine residue is a naturally occurring cysteine residue. In others, the cysteine residue is a residue introduced into the target molecule synthetically.
[0124] In some embodiments, the reactive moiety is an ortho-quinone or semi-quinone group, or a combination thereof. In some embodiments, the subject method provides a reaction between an ortho-quinone reactive intermediate and a thiol moiety, as depicted in Scheme 1 below:
[0125]
[0126] Where Y 1 L is any convenient biomolecule that optionally contains one or more components selected from: active small molecules, cleavable probes, fluorophores, and metal chelators; L is an optional linker (e.g., as described herein); X 1 Selected from hydrogen and hydroxyl groups; Y 2 It is any convenient biomolecule; and n is an integer from 1 to 3.
[0127] As shown in Scheme 1, in some embodiments, a biomolecule containing a phenolic or catechol moiety (e.g., having formula (I)) is activated with an enzyme capable of oxidizing the phenolic or catechol moiety. In some cases, activation is achieved with a tyrosinase in the presence of oxygen to produce an intermediate containing a reactive moiety (e.g., an orthoquinone of formula (II) and / or a semiquinone group of formula (IIA)), and the reactive moiety reacts with a target molecule containing a thiol-based nucleophile (e.g., having formula (III)) to result in the target molecule conjugating to the biomolecule, thereby producing a modified target molecule (e.g., having formula (genIV)). In some embodiments, the target molecule of formula (III) may contain any convenient biomolecule, such as those described herein. In some cases, Y in formula (III) 2 It is a polypeptide. In some cases, the modified molecule is described by formula (IV). In some cases, the modified target molecule is described by formula (IVA). In some cases, the modified target molecule is described by any of formulas (IV)-(IVL), as described herein.
[0128] In some embodiments, the subject method provides a reaction between an orthoquinone reactive intermediate and a thiol moiety, as depicted in Scheme 2 below:
[0129]
[0130] As illustrated in Scheme 2, in some embodiments, a biomolecule containing a phenolic moiety (e.g., having formula (IB)) is activated with a tyrosinase in the presence of oxygen to produce an intermediate containing a reactive moiety (e.g., an ortho-quinone of formula (II)), and said reactive moiety reacts with a target molecule containing a thiol-based nucleophile (e.g., having formula (III)), resulting in the target molecule conjugating to the biomolecule, thereby producing a modified target molecule (e.g., having formula (IVM)). In some embodiments, the target molecule of formula (III) may contain any convenient biomolecule, for example, as described herein. In some cases, Y in formula (III) 2 It is a polypeptide. In some cases of the modified molecule of formula (IVM), the thiol group is located at the 3 position of the catechol ring. In some cases of the modified molecule of formula (IVM), the thiol group is located at the 5 position of the catechol ring. In some cases of the modified molecule of formula (IVM), the thiol group is located at the 6 position of the catechol ring.
[0131] In some embodiments, the biomolecule of formula (I) can be any of formulas (IA)-(IDb), for example, as described herein and discussed in more detail below. In some embodiments, the modified target molecule can have any of formulas (IV)-(IVL), for example, as described herein and discussed in more detail below. In some embodiments, the modified target molecule is a single-conjugation product, for example, as shown in formulas (IV1)-(IV3), (IVA1)-(IVA3), (IVB1)-(IVB3), (IVC1)-(IVC3), (IVD1)-(IVD3), (IVE1)-(IVE3), (IVF1)-(IVF3), (IVG1)-(IVG3), (IVH1)-(IVH3), and (IVJ1)-(IVJ3). In some cases, the modified target molecule is a product of double conjugation, for example, as shown in formulas (IV4)-(IV5), (IVA4)-(IVA5), (IVB4)-(IVB5), (IVC4)-(IVC5), (IVD4)-(IVD5), (IVE4)-(IVE5), (IVF4)-(IVF5), (IVG4)-(IVG5), (IVH4)-(IVH5), and (IVJ4)-(IVJ5).
[0132] In some embodiments, the main process is performed at a pH of 4 to 9, such as 4.2, 4.5, 4.8, 5.0, 5.2, 5.5, 5.8, 6.0, 6.2, 6.5, 6.8, 7.0, 7.2, 7.5, 7.8, 8.0, 8.2, 8.5, 8.8, or 9. In some embodiments, the main process is performed at a pH of 5 to 8, such as 5.2, 5.5, 5.8, 6.0, 6.2, 6.5, 6.8, 7.0, 7.2, 7.5, 7.8, or 8.0. In some cases, the main process is performed at a pH of 6 to 7.5, such as 6.0, 6.3, 6.4, 6.5, 6.6, 6.8, 7.0, 7.2, 7.4, or 7.5. In some embodiments, the main process is performed at a neutral pH. As used herein, the term "neutral pH" means a pH between approximately 7.0 and approximately 7.4. The term "neutral pH" includes pH values of approximately 7.0, 7.05, 7.1, 7.15, 7.2, 7.25, 7.3, 7.35, and 7.4.
[0133] In some embodiments, the subject method can be performed under physiological conditions. In some embodiments, the method is performed in vitro on living cells. In other embodiments, the method is performed ex vivo on living cells.
[0134] In some embodiments, the subject method can be carried out in an aqueous medium in the presence of one or more buffer solutions. Buffer solutions of interest include, but are not limited to, phosphate buffer, 2-amino-2-(hydroxymethyl)propane-1,3-diol (TRIS), 4-[4-(2-hydroxyethyl)piperazin-1-yl]ethanesulfonic acid (HEPES), etc. In some embodiments, the subject method can be carried out in an organic solvent. In some cases, the organic solvent is a water-miscible solvent. In some cases, the organic solvent is a dipolar aprotic solvent. In some cases, the organic solvent is selected from acetonitrile, dimethylformamide, methanol, and acetone. In some cases, the organic solvent is present in an amount of 1% to 20% relative to water, such as 2%, 5%, 10%, 15%, or 20%. In some cases, the subject method is carried out in 1% to 20%, such as 5%, 10%, 15%, or 20% acetonitrile. In some cases, the subject method is carried out in 1% to 20%, such as 5%, 10%, 15%, or 20% dimethylformamide. In some cases, the main process is carried out in 1% to 20%, such as 5%, 10%, 15%, or 20% methanol. In other cases, the main process is carried out in 1% to 20%, such as 5%, 10%, 15%, or 20% acetone.
[0135] In some embodiments of the subject method, the modified target molecule is a product of dual or triple conjugation (e.g., see Formula (IV), collectively referred to herein as a "multiple conjugation product" when n is 2 or 3). In some embodiments of the subject method, the multiple conjugation product is present in less than 1 part / 10 parts by weight relative to a single conjugation product (e.g., see Formula (IV), when n is 1), such as less than 1 part / 20 parts, less than 1 part / 25 parts, less than 1 part / 50 parts, less than 1 part / 75 parts, less than 1 part / 100 parts, or even less. In some embodiments of the subject method, no multiple conjugation product is observed.
[0136] In some embodiments of this method, the modified target molecule is stable within a range of pH and temperature values, and in the presence of a variety of other molecules. In some cases, the modified target molecule is stable at temperatures ranging from 0°C to 50°C, such as 4°C to 40°C, or 4°C to 37°C. In some cases, the modified target molecule is stable in a pH range of 4 to 9, such as at pH 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, or 9. In some cases, the modified target molecule is stable in the presence of biologically relevant molecules. In some cases, the modified target molecule is stable in the presence of molecules such as guanidino groups of arginine residues, primary amines of lysine residues, and aniline moieties. In some cases, the modified target molecule is stable under physiological conditions; for example, in some cases, the modified target molecule is stable in human serum. In some cases, the modified target molecule (also referred to herein as a “target molecule-biomolecule conjugate”) is stable in human serum at 37°C for a period of at least 2 days, at least 3 days, at least 4 days, at least 5 days, at least 6 days, at least 7 days, at least 10 days, or at least 14 days. In some cases, the modified target molecule (also referred to herein as a “target molecule-biomolecule conjugate”) is stable in human serum at 37°C for a period of about 2 days to about 7 days, about 7 days to about 10 days, or about 10 days to about 14 days.
[0137] As described above, in some cases, the target molecule comprises a single thiol moiety. In other cases, the target molecule comprises two thiol moieties. Besides being capable of this in the initial oxidative coupling reaction, a nucleophile with a second thiol moiety can be added during reoxidation in a second time step near the newly formed catechol. The intramolecular nature of this second addition can prevent or minimize the second addition of glutathione or other molecules with free thiols in the biological environment. An example of using a dithiol nucleophile (dithiol target molecule) is schematically depicted in… Figure 44A middle.
[0138] Another implementation of this strategy can be a protein-coupled pair having two cysteine residues. The cysteine residues may be closely adjacent to each other due to their position in the protein's amino acid sequence, or they may be spatially close due to their position in the protein's three-dimensional structure. The polypeptide may include dithiols, wherein the polypeptide contains, for example, sequences such as CC, CGC, CGGC (SEQ ID NO: 1055), or CGGGC (SEQ ID NO: 1056). For example, the polypeptide may contain an amino acid sequence of the following general formula: X n1 C(X) n2 CX n3(SEQ ID NO: 1057), where X is any natural (encoded) or non-natural (non-encoded) amino acid, n1 and n3 are each independently zero or an integer from 1 to 5000 (or greater than 5000), and n2 is zero or an integer from 1 to approximately 10. Examples of such dithiol target molecules are illustratively described in Figure 44B middle.
[0139] Tyrosinase polypeptide
[0140] Tyrosinase peptides suitable for producing reactive moieties (e.g., ortho-quinones) include those with... Figure 8 , Figure 9 , Figures 10A to 10Z and Figures 10AA to 10VVThe tyrosinase polypeptide shown has at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity. In some cases, the tyrosinase polypeptide is a *Agaricus bisporus* tyrosinase polypeptide. In some cases, the tyrosinase polypeptide is a *Bacillus megaterium* tyrosinase polypeptide. In some cases, the tyrosinase polypeptide is a *Streptomyces castaneoglobisporus* tyrosinase polypeptide. In some cases, the tyrosinase polypeptide is a *Citrobacter freundii* tyrosinase polypeptide. In some cases, the tyrosinase polypeptide is a *Homo sapiens* tyrosinase polypeptide. In some cases, the tyrosinase polypeptide is a *Malus domestica* tyrosinase polypeptide. In some cases, the tyrosinase polypeptide is an *Aspergillus oryzae* tyrosinase polypeptide. In some cases, the tyrosinase polypeptide is from tomato (Solanum lycopersicum). In others, it is from Burkholderia thailandensis. And in still others, it is from walnut (Juglans regia). See, for example, Pretzler et al., Sci. Rep. 2017, 7 (1), 1810; Ren et al., BMC Biotechnol. 2013, 13, 18; Faccio et al., ProcessBiochem. 2012, 47 (12), 1749–1760; Fairhead et al., FEBS J. 2010, 277 (9), 2083–2095; Do et al., Sci. Rep. 2017, 7 (1), 17267; Elsayed and Danial J. Appl. Pharm. Sci. 2018, 8 (09), 93–101; Lopez-Tejedor and Palomo Protein Expr. Purif. 2018, 145, 64–70; and Fairhead et al., Nature Biotechnol. 2012, 29 (2). 183–191.
[0141] In some cases, tyrosinase peptides selectively act (e.g., producing reactive moieties such as ortho-quinones) on substrates (biomolecules) containing a phenolic moiety (e.g., tyrosine) or a catechol moiety, wherein the substrate is neutral or positively charged within 50 Å (e.g., within 50 Å, within 40 Å, within 30 Å, or within 20 Å) of the phenolic or catechol moiety. For example, with Figure 8 or Figure 9 Any of the tyrosinases shown having at least 75%, 80%, 90%, 95%, 98%, 99%, or 100% amino acid sequence identity can selectively modify the phenolic or catechol moiety on the substrate, wherein the substrate is neutral or positively charged within 50 Å (e.g., within 50 Å, within 40 Å, within 30 Å, or within 20 Å) of the phenolic or catechol moiety. For example, when the biomolecule is a polypeptide, in some cases, the biomolecule contains at least two neutral or positively charged amino acids within 10 amino acids of the phenolic moiety (e.g., tyrosine) or catechol moiety. For example, when the biomolecule is a polypeptide, in some cases, the biomolecule contains 2, 3, 4, 5, 6, 7, 8, 9, or 10 neutral or positively charged amino acids within 10 amino acids of the phenolic moiety (e.g., tyrosine) or catechol moiety. For example, when the biomolecule is a polypeptide, in some cases, the biomolecule contains the amino acid sequence RRRY (SEQ ID NO: 949), YRRR (SEQ ID NO: 950), RRRRY (SEQ ID NO: 951), or YRRRR (SEQ ID NO: 952).
[0142] In some cases, tyrosinase peptides selectively act (e.g., producing reactive moieties such as ortho-quinones) on substrates (biomolecules) containing a phenolic moiety (e.g., tyrosine) or a catechol moiety, wherein the substrate is negatively charged within 50 Å (e.g., within 50 Å, within 40 Å, within 30 Å, or within 20 Å) of the phenolic or catechol moiety. For example, with Figures 10A to 10Z and Figures 10AA to 10VVAny of the tyrosinases shown having at least 75%, 80%, 90%, 95%, 98%, 99%, or 100% amino acid sequence identity can selectively modify the phenolic or catechol moiety on the substrate, wherein the substrate is negatively charged within 50 Å (e.g., within 50 Å, within 40 Å, within 30 Å, or within 20 Å) of the phenolic or catechol moiety. For example, when the biomolecule is a polypeptide, in some cases, the biomolecule contains at least two negatively charged amino acids within 10 amino acids of the phenolic moiety (e.g., tyrosine) or catechol moiety. For example, when the biomolecule is a polypeptide, in some cases, the biomolecule contains 2, 3, 4, 5, 6, 7, 8, 9, or 10 negatively charged amino acids within 10 amino acids of the phenolic moiety (e.g., tyrosine) or catechol moiety. For example, when the biomolecule is a polypeptide, in some cases, the biomolecule contains the amino acid sequence EEEY (SEQ ID NO: 953), YEEE (SEQ ID NO: 954), EEEEY (SEQ ID NO: 955), or YEEEE (SEQ ID NO: 956).
[0143] In some cases, tyrosinase peptides contain... Figure 10M The illustrated tyrosinase amino acid sequence has an amino acid sequence with at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity, wherein the tyrosinase polypeptide contains an amino acid substitution for D55, for example, where D55 is substituted with Lys. Such tyrosinase polypeptides are particularly useful when the biomolecule has a net negative charge and / or the region surrounding the phenolic or catechol moiety has a net negative charge (e.g., when the phenolic group is Tyr, Tyr can be present in the EEEEEY (SEQ ID NO: 955) or EEEY (SEQ ID NO: 953) peptide). Such tyrosinase polypeptides are particularly useful when the biomolecule is a nucleic acid.
[0144] In some cases, tyrosinase peptides contain... Figure 10CThe tyrosinase amino acid sequences shown have amino acid sequence identity of at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%, wherein the tyrosinase polypeptide contains an amino acid substitution of R209, for example, wherein R209 is substituted with His. Such tyrosinase polypeptides are particularly useful when the biomolecule has a net positive charge and / or the region surrounding the phenolic or catechol moiety has a net positive charge (e.g., when the phenolic group is Tyr, Tyr can be present in RRRY (SEQ ID NO: 949) or RRRRY (SEQ ID NO: 951) peptides).
[0145] Cell surface modification
[0146] In some embodiments, the subject method is used to modify cell surfaces. Therefore, in one aspect, the present invention provides a method for in vitro modification of cell surfaces. This method generally includes reacting a thiol group in a target molecule with a biomolecule containing a reactive moiety to provide a chemoselective conjugation at the cell surface. In some embodiments, the method includes modifying a target molecule on a cell surface with a thiol moiety; and reacting the thiol moiety in the target molecule with a biomolecule containing a reactive moiety (e.g., an orthoquinone moiety). In other embodiments, the method includes activating a biomolecule containing a phenolic moiety on a cell surface to generate a biomolecule containing a reactive moiety; and reacting the reactive moiety in the biomolecule with a target molecule containing a thiol moiety.
[0147] Modify target molecules with detectable labels, drugs and other molecules.
[0148] In some embodiments, this disclosure provides for the linking of a biomolecule of interest to a target molecule comprising a thiol moiety. The method typically involves reacting the thiol-containing target molecule with a subject biomolecule comprising a reactive moiety (e.g., an orthoquinone moiety). Target molecules and biomolecules of interest include, but are not limited to, peptides, polynucleotides, carbohydrates, fatty acids, steroids, purines, pyrimidines, derivatives; and so on.
[0149] The connection between the biomolecules of interest and the carrier
[0150] Biomolecules containing reactive moieties may also include one or more hydrocarbon linkers (e.g., alkyl groups or derivatives thereof, such as alkyl esters or PEGs) that conjugate to a moieties providing attachment to a solid matrix (e.g., to facilitate analysis) or to a moieties providing easily separable portions (e.g., haptens recognized by antibodies bound to magnetic beads). In one embodiment, the methods of the present invention are used to provide proteins (or other molecules containing or potentially modified to contain thiols) to attach onto a chip in a defined orientation. For example, the methods and compositions of this disclosure can be used to deliver tags or other portions (e.g., as described herein) to the thiol of a target molecule, such as a polypeptide having a thiol moiety at a selected site (e.g., at or near the N-terminus). The tags or other portions can then be used as attachment sites for attaching the molecules to a carrier (e.g., a solid or semi-solid carrier, such as a carrier suitable for use in microchips in high-throughput analysis).
[0151] Attachment of biomolecules for delivery to target sites
[0152] In some embodiments, the biomolecule containing the reactive portion will comprise a small molecule drug, toxin, or other molecule for delivery to cells. In some embodiments, the small molecule drug, toxin, or other molecule will provide pharmacological activity. In some embodiments, the small molecule drug, toxin, or other molecule will serve as a target for the delivery of other molecules.
[0153] Small molecule drugs can be small organic or inorganic compounds with a molecular weight greater than 50 and less than about 2,500 Daltons. Small molecule drugs may contain functional groups necessary for interaction with protein structures, particularly hydrogen bonding, and may include at least one amine, carbonyl, hydroxyl, or carboxyl group, and may contain at least two functional chemical groups. Drugs may contain cyclic carbon or heterocyclic structures and / or aromatic or polyaromatic structures substituted with one or more of the above functional groups. Small molecule drugs are also found in biomolecules, including peptides, sugars, fatty acids, steroids, purines, pyrimidines, derivatives, structural analogs, or combinations thereof.
[0154] In another embodiment, the subject biomolecule containing the reactive portion comprises a pair of binding partners (e.g., a ligand; a ligand-binding portion of a receptor; an antibody; an antigen-binding fragment of an antibody; an antigen; a hapten; a lectin; a lectin-bound carbohydrate; etc.). For example, the biomolecule may contain a polypeptide that acts as a viral receptor and, when bound to a viral envelope protein or viral capsid protein, facilitates viral attachment to the cell surface displaying the biomolecule. Alternatively, the biomolecule may contain an antigen specifically bound by an antibody (e.g., a monoclonal antibody) to facilitate the detection and / or isolation of host cells displaying the antigen on their cell surface. In another instance, the biomolecule comprises a ligand-binding portion of a receptor or a receptor-binding portion of a ligand.
[0155] compound
[0156] Biomolecules containing phenolic or catechol moieties
[0157] In some embodiments of the subject method, biomolecules containing phenolic or catechol moieties are described by formula (I).
[0158]
[0159] Where Y 1 It is a biomolecule, which optionally comprises one or more portions selected from: active small molecules, affinity tags, fluorophores, and metal chelators; X 1 It is selected from hydrogen and hydroxyl; and L is an optional connector.
[0160] In some implementations of formula (I), X 1 It is hydrogen that allows the biomolecule to contain the phenolic moiety. In other embodiments of formula (I), X 1 It is the hydroxyl group that makes biomolecules contain the catechol moiety.
[0161] In some embodiments of formula (I), the phenolic moiety is present in the tyrosine residue. In certain cases, biomolecules containing the phenolic moiety of formula (I) have formula (IB) or (IC):
[0162]
[0163] Where R 2 Selected from alkyl and substituted alkyl groups; and R 3 Selected from hydrogen, alkyl, substituted alkyl, peptide and polypeptide.
[0164] In some embodiments of the subject method, the biomolecule containing a phenolic moiety or catechol (e.g., having formula (I)) includes a linker (e.g., as described herein). Suitable linkers include, but are not limited to, carboxylic acids, alkyl esters, aryl esters, substituted aryl esters, aldehydes, amides, arylamides, alkyl halides, thioesters, sulfonyl esters, alkyl ketones, aryl ketones, substituted aryl ketones, halosulfonyl groups, nitriles, PEGs, and peptide linkers.
[0165] Exemplary linkers for attaching a phenolic moiety to a biomolecule of interest will include PEG linkers in some embodiments. The term “PEG” as used herein refers to polyethylene glycol or modified polyethylene glycol. Modified polyethylene glycol polymers include methoxy polyethylene glycol and unsubstituted or substituted polymers at one end with an alkyl, substituted alkyl, or functional group (e.g., as described herein). Any convenient linking group can be used at the end of the PEG to attach that group to the moiety of interest, including but not limited to alkyl, aryl, hydroxyl, amino, acyl, acyloxy, carboxyl ester, and amide-terminal and / or substituent groups. In some cases, the linker comprises more than one PEG unit, such as 2, 3, 4, 5, 6, 7, 8, 9, or 10 PEG units. In some cases, the linker comprises fewer than 10 PEG units, such as 9, 8, 7, 6, 5, 4, 3, 2, or 1 PEG unit. In some cases, the linker consists of four or fewer PEG units.
[0166] In some cases, biomolecules containing phenolic moieties are described by formula (IA):
[0167]
[0168] in:
[0169] Y 1 It is a biomolecule, which optionally includes one or more groups selected from the following: active small molecules, affinity tags, fluorophores and metal chelators;
[0170] Each R 1 Independently selected from hydrogen, acyl, substituted acyl, alkyl, and substituted alkyl;
[0171] X 1 Selected from hydrogen and hydroxyl; and
[0172] L 1 It is a linker selected from straight-chain or branched alkyl groups, straight-chain or branched substituted alkyl groups, polyethylene glycol (PEG), substituted PEG, and one or more peptides.
[0173] In some implementations, X 1 It is hydrogen that makes compound (IA) have formula (IAa):
[0174]
[0175] In some embodiments of any of formulas (IA)-(IAa), at least one R 1 It's hydrogen. In some cases, two R... 1 The groups are all hydrogen. In some cases, an R group... 1 The group is hydrogen, and another R 1The group is selected from alkyl, substituted alkyl, acyl, and substituted acyl groups. In some cases, an R... 1 The group is hydrogen and another R 1 The group is alkyl. In some cases, an R 1 The group is hydrogen and another R 1 The group is a substituted alkyl group. In some cases, an R 1 The group is hydrogen and another R 1 The group is an acyl group. In some cases, an R... 1 The group is hydrogen, and another R 1 The group is a substituted acyl group. In some cases, the acyl group has the formula -C(O)R 4 , where R 4 It is a lower alkyl group, such as methyl, ethyl, propyl, butyl, pentyl, or hexyl. In some cases, the substituted acyl group has the formula -C(O)R. 4 NH2, where R 4 It is a lower alkyl group, such as methyl, ethyl, propyl, butyl, pentyl, or hexyl. In some cases, the substituted acyl group has the formula -C(O)CH2NH2.
[0176] In some implementations of any of formulas (IA)-(IAa), L 1 It can be a straight-chain or branched alkyl group. In some cases, L 1 It is a lower alkyl group, such as methyl, ethyl, propyl, butyl, pentyl, or hexyl. In some cases, L... 1 It is a substituted alkyl group. In some cases, L 1 It is a substituted lower alkyl group. In some cases, L 1 It is a PEG or a substituted PEG (e.g., as described herein). In some other cases, L 1 It is a peptide. In some other cases, L... 1 It is a polypeptide. In some cases, L 1 These are linear connectors with lengths of 1 to 12 atoms, such as those with lengths of 1-10, 1-8, or 1-6 atoms, for example, linear connectors with lengths of 1, 2, 3, 4, 5, or 6 atoms. Connector L 1 It can be (C) 1-6 )alkyl connectors or substituted (C 1-6 The alkyl linker is optionally substituted with a heteroatom or linking functional group, such as ester (-CO2-), amide (CONH), urethane (OCONH), ether (-O-), thioether (-S-), and / or amino (-NR-, where R is H or alkyl). In some cases, the linker L 1 It may include a ketone group (C=O). In some cases, the ketone group, together with an amino, thiol, or ether group in the connector chain, can provide an amide, ester, or thioester group for attachment.
[0177] In some embodiments, the linking group L or L 1 It is a detachable connector, for example, as described herein.
[0178] In some embodiments, biomolecules containing phenolic or catechol moieties are described by formula (ID):
[0179]
[0180] Where Y 1 It is a biomolecule, which optionally comprises one or more groups selected from: active small molecules, affinity tags, fluorophores, and metal chelators; X 1 The integer n is selected from hydrogen and hydroxyl; and n is an integer from 0 to 20. In some cases, n is 10 or less, such as 9, 8, 7, 6, 5, 4, 3, 2, 1, or 0. In some cases, n is 5. In some cases, n is 4. In some cases, n is 3. In some cases, n is 2. In some cases, n is 1. In some cases, n is 0.
[0181] In some implementations, n is 1, such that compound (ID) has formula (IDa):
[0182]
[0183] In some cases of formula (ID) or (IDa), X 1 It is hydrogen that allows the biomolecule to include the phenolic moiety. In other embodiments of formula (ID) or (IDa), X 1 It is the hydroxyl group that makes biomolecules contain the catechol moiety.
[0184] In some cases, compounds of formula (IDa) have the following characteristics:
[0185]
[0186] Compounds of any of formulas (ID)-(IDb) can be prepared by reacting tyramine or a corresponding phenolic or catechol-containing amine with a biomolecule including an N-hydroxysuccinimide (NHS) ester or maleimide group in a suitable solvent. For example, compounds of formula (IDb) can be prepared by reacting NHS-ester (Y... 1 -NHS) is reacted with tyramine in anhydrous dimethylformamide (DMF) to provide compound (IDb) as shown in Scheme 3 below:
[0187]
[0188] It should be understood that biomolecules containing either a phenolic or catechol moiety (e.g., having any of the formulas (I)-(IDb)) can be prepared by any convenient method. Numerous general references providing commonly known chemical synthetic schemes and conditions suitable for synthesizing moieties containing both the subject phenolic and catechol moieties are available (see, for example, Smith and March, March's Advanced Organic Chemistry: Reactions, Mechanisms, and Structure, 5th ed., Wiley-Interscience, 2001; or Vogel, A Textbook of Practical Organic Chemistry, Including Qualitative Organic Analysis, 4th ed., New York: Longman, 1978). As disclosed herein, in some cases, the subject phenolic moiety is present in a tyrosine residue. A tyrosine residue can be part of the biomolecule of interest. In other cases, the tyrosine moiety can be synthesized and introduced into the biomolecule of interest. For example, when the biomolecule is a peptide or polypeptide, tyrosine residues can be introduced via standard solid-phase Fmoc peptide chemistry (Fields GB, Noble RL. Solid phase peptide synthesis utilizing 9-fluorenylmethoxycarbonylamino acids. Int J Pept Protein Res 35: 161–214, 1990). In some cases, the phenolic or catechol moiety is part of a non-natural (non-genetically encoded) amino acid that is introduced into the biomolecule of interest. For example, amber codon (TAG) inhibition can be used to incorporate non-genetically encoded amino acid residues containing phenolic or catechol moieties. See, for example, Chin et al. (2002) J. Am. Chem. Soc. 124:9026; Chin and Schultz (2002) Chem. Biol. Chem. 3:1135; Chin et al. (2002) Proc. Natl. Acad. Sci. USA 99:11020; US2015 / 0240249; and US 2018 / 0171321. As another example, orthogonal RNA synthases and / or orthogonal tRNAs can be used to introduce non-genetically encoded amino acids into biomolecules, wherein the non-genetically encoded amino acids contain a phenolic or catechol moiety.
[0189] In some embodiments of any of formulas (I)-(IDb), the biomolecule of interest comprises one or more groups selected from: an active small molecule, an affinity tag, a fluorophore, and a metal chelator. In some cases, the fluorophore is a rhodamine dye. In some cases, the fluorophore is a xanthan dye. In some cases, the fluorophore is Oregon Green 488. In some cases, the metal chelator is 1,4,7,10-tetraazacyclododecane-1,4,7,10-tetraacetic acid (also known as DOTA or tetraxetan). In some cases, the affinity tag is a biotin moiety (e.g., as described herein).
[0190] In some cases, biomolecules containing phenolic moieties are composed of... Figure 3 The structure is described in the figure.
[0191] Target molecules containing thiol moieties
[0192] Molecules containing thiol moieties suitable for the subject method, and methods for preparing thiol-containing molecules suitable for the subject method, are well known in the art.
[0193] Target molecules can be naturally occurring or synthetically or recombinantly generated, and can be isolated, substantially purified, or present in the natural environment of unmodified molecules. Thiol-containing target molecules are based on these unmodified molecules (e.g., on the cell surface or within cells, including within host animals, such as mammals, such as rodent hosts (e.g., rats, mice), hamsters, dogs, cats, cattle, pigs, etc.). In some embodiments, the target molecule is present in an in vitro cell-free reaction. In other embodiments, the target molecule is present in cells and / or displayed on the cell surface. In many embodiments of interest, the target molecule is in living cells; on the surface of a living cell; in a living organism, such as in a living multicellular organism. Suitable living cells include cells that are part of a living multicellular organism; cells isolated from a multicellular organism; immortalized cell lines; and so on.
[0194] The target molecule can be composed of D-amino acids, L-amino acids, or both, and can be further modified naturally, synthetically, or recombinantly to include other parts. For example, the target molecule can be a lipoprotein, glycoprotein, or other such modified protein.
[0195] Generally, the target molecule contains at least one thiol moiety for reacting with a biomolecule containing the reactive moiety according to the invention, but may contain two or more, three or more, five or more, ten or more thiol moieties. The number of thiol moieties that may be present in the target molecule will vary depending on the intended application of the modified target molecule, the properties of the target molecule itself, and other considerations that will be readily apparent to those skilled in the art when practicing the methods disclosed herein.
[0196] Target molecules can be modified to include a thiol moiety at the point where they need to connect with a biomolecule containing a reactive moiety. For example, when the target molecule is a peptide or polypeptide, the target molecule substrate can be modified to contain an N-terminal thiol moiety, thereby producing a subject target peptide or polypeptide containing a thiol moiety. It should be understood that any convenient position on the peptide or polypeptide substrate can be modified to contain a thiol moiety, thereby producing a target peptide or polypeptide for subject methods.
[0197] In some implementations, the target molecule containing the thiol moiety is a CRISPR-Cas effector peptide.
[0198] In some cases, the thiol moiety is present in the cysteine residue. In some cases, the cysteine residue is native to the CRISPR-Cas effector peptide. In other cases, the cysteine residue is introduced into the CRISPR-Cas effector peptide. For example, the cysteine residue can be introduced by standard solid-phase Fmoc peptide chemistry (Fields GB, Noble RL. Solid phase peptide synthesis utilizing 9-fluorenylmethoxycarbonyl aminoacids. Int J Pept Protein Res 35: 161–214, 1990).
[0199] Modified target molecules
[0200] In some embodiments of the subject method, the resulting modified target molecule has formula (IV) or (IVA), or a combination thereof. Therefore, aspects of this disclosure include compounds of formula (IV) or (IVA):
[0201]
[0202] Wherein Y1 is a biomolecule, which optionally includes one or more portions selected from: active small molecules, affinity tags, fluorophores, and metal chelators; L is an optional linker; Y2 is a second biomolecule; and n is an integer from 1 to 3.
[0203] In some embodiments of formula (IV) or (IVA), n is less than 3, such as 2 or 1. In some cases, n is 2. In some cases, n is 1. In some cases, the target molecule of the subject modification is a compound of formula (IV). In some cases, the target molecule of the subject modification is a compound of formula (IVA).
[0204] In some embodiments, the target molecule modified by formula (IV), n is 1, and the compound is described by any one of formulas (IV1)-(IV3):
[0205]
[0206] In some embodiments, the target molecule modified by formula (IV) is n = 2, and the compound is described by any one of formulas (IV4)-(IV5):
[0207]
[0208] In some embodiments, the modified target molecule has the formula (IVA), n is 1, and the compound is described by any one of formulas (IVA1)-(IVA3):
[0209]
[0210] In some embodiments, the modified target molecule has the formula (IVA), n is 2, and the compound is described by any one of formulas (IVA4)-(IVA5):
[0211]
[0212] In some embodiments, the modified target molecule includes a linker (e.g., as described herein). Suitable linkers include, but are not limited to, carboxylic acids, alkyl esters, aryl esters, substituted aryl esters, aldehydes, amides, aryl amides, alkyl halides, thioesters, sulfonyl esters, alkyl ketones, aryl ketones, substituted aryl ketones, halosulfonyl groups, nitriles, and peptide linkers.
[0213] In some embodiments, it is used to link orthoquinones to biomolecules (Y). 1 An exemplary connector would include amides, such as –(CR 1 2) m NHC(O)-, where R 1 The group is selected from hydrogen or a substituent (e.g., as described herein) and m is an integer from 1 to 20. Exemplary connectors may also include PEG or substituted PEG connectors, for example, as described herein.
[0214] In some implementations, the connector is a detachable connector, for example, as described herein.
[0215] In some implementations, the modified target molecule is described by formula (IVB) or (IVC):
[0216]
[0217] Where Y 1 It is a biomolecule, which optionally comprises one or more groups selected from: active small molecules, affinity tags, fluorophores, and metal chelators; each R 1Independently selected from hydrogen, acyl, substituted acyl, alkyl, and substituted alkyl; Y 2 It is the second biomolecule; L 1 It is a linker selected from straight-chain or branched alkyl groups, straight-chain or branched substituted alkyl groups, polyethylene glycol (PEG), substituted PEG and one or more peptides; and n is an integer from 1 to 3.
[0218] In some embodiments of formula (IVB) or (IVC), n is less than 3, such as 2 or 1. In some cases, n is 2. In some cases, n is 1. In some cases, the target molecule of the subject modification is a compound of formula (IVB). In some cases, the target molecule of the subject modification is a compound of formula (IVC).
[0219] In some embodiments, the target molecule modified by formula (IVB), n is 1, and the compound is described by any one of formulas (IVB1)-(IVB3):
[0220]
[0221] In some embodiments, the target molecule modified by formula (IVB), n is 2, and the compound is described by any one of formulas (IVB4)-(IVB5):
[0222]
[0223] In some embodiments, the modified target molecule has the formula (IVC), n is 1, and the compound is described by any one of formulas (IVC1)-(IVC3):
[0224]
[0225] In some embodiments, the modified target molecule has the formula (IVC), n is 2, and the compound is described by any one of formulas (IVC4)-(IVC5):
[0226]
[0227] In some embodiments of any of formulas (IVB)-(IVC5), at least one R 1 It's hydrogen. In some cases, two R... 1 The groups are all hydrogen. In some cases, an R group... 1 The group is hydrogen, and another R 1 The group is selected from alkyl, substituted alkyl, acyl, and substituted acyl groups. In some cases, an R... 1 The group is hydrogen and another R 1 The group is alkyl. In some cases, an R 1 The group is hydrogen and another R1 The group is a substituted alkyl group. In some cases, an R 1 The group is hydrogen and another R 1 The group is an acyl group. In some cases, an R... 1 The group is hydrogen, and another R 1 The group is a substituted acyl group. In some cases, the acyl group has the formula -C(O)R 4 , where R 4 It is a lower alkyl group, such as methyl, ethyl, propyl, butyl, pentyl, or hexyl. In some cases, the substituted acyl group has the formula -C(O)R. 4 NH2, where R 4 It is a lower alkyl group, such as methyl, ethyl, propyl, butyl, pentyl, or hexyl. In some cases, the substituted acyl group has the formula -C(O)CH2NH2.
[0228] In some implementations of any of formulas (IVB)-(IVC5), L 1 It can be a straight-chain or branched alkyl group. In some cases, L 1 It is a lower alkyl group, such as methyl, ethyl, propyl, butyl, pentyl, or hexyl. In some cases, L... 1 It is a substituted alkyl group. In some cases, L 1 It is a substituted lower alkyl group. In some cases, L 1 It is a PEG or a substituted PEG (e.g., as described herein). In some other cases, L 1 It is a peptide. In some other cases, L... 1 It is a polypeptide. In some cases, L 1 These are linear connectors with lengths of 1 to 12 atoms, such as those with lengths of 1-10, 1-8, or 1-6 atoms, for example, linear connectors with lengths of 1, 2, 3, 4, 5, or 6 atoms. Connector L 1 It can be (C) 1-6 )alkyl connectors or substituted (C 1-6 The alkyl linker is optionally substituted with a heteroatom or linking functional group, such as ester (-CO2-), amide (CONH), urethane (OCONH), ether (-O-), thioether (-S-), and / or amino (-NR-, where R is H or alkyl). In some cases, the linker L 1 It may include a ketone group (C=O). In some cases, the ketone group, together with an amino, thiol, or ether group in the connector chain, can provide an amide, ester, or thioester group for attachment.
[0229] In some embodiments, the linking group L 1 It is a detachable connector, for example, as described herein.
[0230] In some implementations, the modified target molecule is described by any of the formulas (IVD)-(IVG):
[0231]
[0232] Where R 2 Selected from alkyl and substituted alkyl groups; R 3 It is selected from hydrogen, alkyl, substituted alkyl, peptide and polypeptide; and n is an integer from 1 to 3.
[0233] In some embodiments of any of formulas (IVD)-(IVG), n is less than 3, such as 2 or 1. In some cases, n is 2. In some cases, n is 1. In some cases, the target molecule of the subject modification is a compound of formula (IVD). In some cases, the target molecule of the subject modification is a compound of formula (IVE). In some cases, the target molecule of the subject modification is a compound of formula (IVF). In some cases, the target molecule of the subject modification is a compound of formula (IVG).
[0234] In some embodiments, any of the formulas (IVD)-(IVG) may have the relative stereochemistry shown in the following structures:
[0235]
[0236] .
[0237] In some embodiments, the target molecule modified by formula (IVD), n is 1, and the compound is described by any one of formulas (IVD1)-(IVD3):
[0238]
[0239] In some embodiments, the target molecule modified by formula (IVD) has n = 2, and the compound is described by any one of formulas (IVD4)-(IVD5):
[0240]
[0241] In some embodiments, the target molecule modified by formula (IVE), n is 1, and the compound is described by any one of formulas (IVE1)-(IVE3):
[0242]
[0243] In some embodiments, the target molecule modified by formula (IVE) has n = 2, and the compound is described by any one of formulas (IVE4)-(IVE5):
[0244]
[0245] In some embodiments, the target molecule modified by formula (IVF), n is 1, and the compound is described by any one of formulas (IVF1)-(IVF3):
[0246]
[0247] In some embodiments, the target molecule modified by formula (IVF), n is 2, and the compound is described by any one of formulas (IVF4)-(IVF5):
[0248]
[0249] In some embodiments, the target molecule modified by formula (IVG), n is 1, and the compound is described by any one of formulas (IVG1)-(IVG3):
[0250]
[0251] In some embodiments, the target molecule modified by formula (IVG), n is 2, and the compound is described by any one of formulas (IVG4)-(IVG5):
[0252]
[0253] In some embodiments of the target molecule described herein, R 2 It is an alkyl group. In some cases, R 2 It is a substituted alkyl group. In some cases, the alkyl group is a lower alkyl group, such as methyl, ethyl, propyl, butyl, pentyl, or hexyl.
[0254] In some embodiments of the target molecule described herein, R 3 It is hydrogen. In some cases, R 3 It is an alkyl group. In some cases, R 3 It is a substituted alkyl group. In some cases, the alkyl group is a lower alkyl group, such as methyl, ethyl, propyl, butyl, pentyl, or hexyl. In some cases, R... 3 It's a peptide. In some cases, R... 3 It is a polypeptide.
[0255] In some implementations, the modified target molecule is described by formula (IVH) or (IVJ):
[0256]
[0257] Where Y 1 It is a biomolecule, which optionally includes one or more groups selected from: active small molecules, affinity tags, fluorophores, and metal chelators; Y2 It is the second biomolecule; n is an integer from 1 to 3; and m is an integer from 0 to 20. In some cases, m is 10 or less, such as 9, 8, 7, 6, 5, 4, 3, 2, 1, or 0. In some cases, m is 5. In some cases, m is 4. In some cases, m is 3. In some cases, m is 2. In some cases, m is 1. In some cases, m is 0.
[0258] In some implementations of the subject approach, the modified target molecule is described by formula (IVK) or (IVL):
[0259]
[0260] In some embodiments of any of equations (IVH)-(IVJ), n is less than 3, such as 2 or 1. In some cases, n is 2. In some cases, n is 1.
[0261] In some embodiments, the target molecule modified by formula (IVH), n is 1, and the compound is described by any one of formulas (IVH1)-(IVH3):
[0262]
[0263] In some embodiments, the target molecule of formula (IVH) has n = 2, and the compound is described by any one of formulas (IVH4)-(IVH5):
[0264]
[0265] In some embodiments, the target molecule modified by formula (IVJ), n is 1, and the compound is described by any one of formulas (IVJ1)-(IVJ3):
[0266]
[0267] In some embodiments, the target molecule modified by formula (IVJ) has n = 2, and the compound is described by any one of formulas (IVJ4)-(IVJ5):
[0268]
[0269] In some embodiments, the target molecule containing a thiol group is a CRISPR-Cas effector peptide (e.g., as described herein).
[0270] In some embodiments of any of formulas (IV) to (IVJ5), Y 1 It is a polypeptide. In some cases, Y 1The peptides are selected from fluorescent proteins, antibodies, and enzymes. In some cases, the fluorescent protein is green fluorescent protein. Other suitable peptides are described elsewhere in this document.
[0271] Disintegral connector
[0272] Cleavable linkers that can be used for subject molecules of interest include electrophilic cleavable linkers, nucleophilic cleavable linkers, photocleavable linkers, metallocleavable linkers, electrolytically cleavable linkers, and linkers cleavable under reducing and oxidizing conditions. In some cases, cleavable linkers cleave under acidic conditions. In some cases, cleavable linkers are cleaved by enzymes. In some cases, cleavable linkers cleave under reducing conditions. In some cases, cleavable linkers cleave rapidly by glutathione reduction. In some cases, cleavable linkers involve disulfide bonds. In some cases, cleavable linkers cleave by physical stimulation. In some cases, cleavable linkers are photocleavable.
[0273] In some cases, L or L 1 It is an acid-labile connector. In some cases, the connector breaks down at pH 6 or lower, such as 6.0, 5.95, 5.9, 5.85, 5.8, 5.75, 5.7, 5.65, 5.6, 5.55, 5.5, 5.45, 5.4, 5.35, 5.3, 5.25, 5.2, 5.15, 5.1, 5.05, 5.0, 4.9, 4.85, 4.80, 4.75, 4.7, 4.65, 4.6, 4.55, 4.5 or even lower.
[0274] In some cases, L or L 1 It is a photodegradable connector. Suitable photodegradable connectors include o-nitrobenzyl connectors, benzoylmethyl connectors, alkoxybenzoin connectors, chromium aromatic complex connectors, NpSSMpact connectors, and neopentanoyl diol connectors, as described by Guillier et al. (Chem. Rev. 2000 1000:2091-2157).
[0275] In some cases, L or L 1 It is a linker that can be hydrolyzed and broken down into proteins.
[0276] This proteolytically cleavable linker may include a protease recognition sequence selected from the group consisting of: alanine carboxypeptidase, Armillaria mellea astaxanthin, bacterial leucylaminopeptidase, cancer-procoagulant, cathepsin B, clostridium protease, cytosol alanylaminopeptidase, elastase, endopeptidase Arg-C, enterokinase, gastric protease, gelatinase, Gly-X carboxypeptidase, glycyl endopeptidase, human rhinovirus 3C protease, ferrobinin C, IgA-specific serine endopeptidase, leucylaminopeptidase, leucyl endopeptidase, lysC, lysosomal pro-X carboxypeptidase, lysylaminopeptidase, methionylaminopeptidase, myxococcus, nardilysin, pancreatic endopeptidase E, and picornaein. 2A, Viral endopeptidase 3C, endopeptidase, prolyl aminopeptidase, protonase I, protonase II, Russelllysin, saccharopepsin, seminal coagulase, T-fibrinolysin activating factor, thrombin, tissue kallikrein, tobacco etching virus (TEV), togavirin, tryptophanyl aminopeptidase, U-fibrinolysin activating factor, V8, venombin A, venombin AB, and Xaa-pro aminopeptidase.
[0277] For example, a proteolytic cleavage linker may include matrix metalloproteinase cleavage sites, such as cleavage sites of MMPs selected from collagenase-1, collagenase-2 and collagenase-3 (MMP-1, MMP-8 and MMP-13), gelatinase A and B (MMP-2 and MMP-9), lysolysin 1, 2 and 3 (MMP-3, MMP-10 and MMP-11), matrix lysolysin (MMP-7) and membrane metalloproteinases (MT1-MMP and MT2-MMP). For example, the cleavage sequence of MMP-9 is Pro-XX-Hy (SEQ ID NO: 1054) (where X represents any residue; Hy, a hydrophobic residue), such as Pro-XX-Hy-(Ser / Thr) (SEQ ID NO: 847), or Pro-Leu / Gln-Gly-Met-Thr-Ser (SEQ ID NO: 848) or Pro-Leu / Gln-Gly-Met-Thr (SEQ ID NO: 849). Another example of a protease cleavage site is a plasminogen activator cleavage site, such as uPA or tissue plasminogen activator (tPA) cleavage site. In some cases, the cleavage site is a furin protease cleavage site. Specific examples of uPA and tPA cleavage sequences include sequences containing Val-Gly-Arg. Another example that may be included in a protease cleavage site in a proteolytically cleavable adapter is the tobacco etch virus (TEV) protease cleavage site, such as ENLYTQS (SEQ ID NO: 850), in which the protease cleaves between glutamine and serine. The TEV protease recognizes a linear amino acid sequence of the general formula EX1X2YX3Q(G / S) (SEQ ID NO: ), where each X1, X2, and X3 is any amino acid, and where cleavage occurs between Q and G or Q and S. TEV protease cleavable adapters may include ENLYFQG (SEQ ID NO: 957); ENLYTQS (SEQ ID NO: 958); ENLYFQGGY (SEQ ID NO: 959); ENLYFQS (SEQ ID NO: 960);Etc. Another example of a protease cleavage site that may be included in a proteolytic cleavage linker is an enterokinase cleavage site, such as DDDDK (SEQ ID NO:851) in which cleavage occurs after a lysine residue. Another example of a protease cleavage site that may be included in a proteolytic cleavage linker is a thrombin cleavage site, such as LVPR (SEQ ID NO:852). Suitable additional adapters containing protease cleavage sites include adapters containing one or more of the following amino acid sequences: LEVLFQGP (SEQ ID NO:853), cleaved by PreScission protease (a fusion protein containing human rhinovirus 3C protease and glutathione S-transferase; Walker et al. (1994) Biotechnol.12:601); thrombin cleavage sites, such as CGLVPAPGSGP (SEQ ID NO:854); SLLKSRMVPNFN (SEQ ID NO:855) or SLLIARRMPNFN (SEQ ID NO:856), cleaved by cathepsin B; SKLVQASASGVN (SEQ ID NO:857) or SSYLKASDAPDN (SEQ ID NO:858), cleaved by Epstein-Barr viral protease; RPKPQQFFGLMN (SEQ ID NO:859), cleaved by MMP-3 (lysosometin); SLRPLALWRSFN (SEQ ID NO:854) or ...APGSGP (SEQ ID NO:855) or SLLAPGSGP (SEQ ID NO:855) or SLLAPGSGP (SEQ ID NO:855) or SLLAPGSGP (SEQ ID NO:855) or SLLAPGSGP (SEQ ID NO:855) or SLLAPGSGP (SEQ ID NO:855) or SLLAP NO:860), cleaved by MMP-7 (stromalloin); SPQGIAGQRNFN (SEQ ID NO:861), cleaved by MMP-9; DVDERDVRGFASFL (SEQ ID NO:862), cleaved by thermophilic protease-like MMP; SLPLGLWAPNFN (SEQ ID NO:863), cleaved by matrix metalloproteinase 2 (MMP-2); SLLIFRSWANFN (SEQ ID NO:864), cleaved by cathepsin L; SGVVIATVIVIT (SEQ ID NO:865), cleaved by cathepsin D; SLGPQGIWGQFN (SEQ ID NO:866), cleaved by matrix metalloproteinase 1 (MMP-1); KKSPGRVVGGSV (SEQ ID NO:867), cleaved by urokinase-type plasminogen activator; PQGLLGAPGILG (SEQ ID NO:868), cleaved by type 1 membrane matrix metalloproteinase (MT-MMP);HGPEGLRVGFYESDVMGRGHARLVHVEEPHT (SEQ ID NO:869) is cleaved by lysosome 3 (or MMP-11), thermophilic protease, fibroblast collagenase, and lysosome 1; GPQGLAGQRGIV (SEQ ID NO:870) is cleaved by matrix metalloproteinase 13 (collagenase-3); GGSGQRGRKALE (SEQ ID NO:871) is cleaved by tissue-type plasminogen activator (tPA); SLSALLSSDIFN (SEQ ID NO:872) is cleaved by human prostate-specific antigen; SLPRFKIIGGFN (SEQ ID NO:873) is cleaved by kallikrein (hK3); SLLGIAVPGNFN (SEQ ID NO:874) is cleaved by neutrophil elastase; and FFKNIVTPRTPP (SEQ ID NO:875) is cleaved by calpain (calcium-activated neutral protease).
[0278] In some cases, the joint contains disulfide bonds and is cleavable under reducing conditions, such as with β-mercaptoethanol, cysteine-HCl, tris(2-carboxyethyl)phosphonic acid hydrochloride, or another reducing agent.
[0279] In some cases, the connector contains a dipeptide, such as a valine-citrulline dipeptide or a valine-lysine dipeptide.
[0280] biomolecules
[0281] Biomolecules suitable for the methods or conjugates of this disclosure include polypeptides, polynucleotides, carbohydrates, lipids, fatty acids, steroids, purines, pyrimidines, their derivatives, structural analogs, and combinations thereof.
[0282] Suitable biomolecules include, but are not limited to, polypeptides, nucleic acids, glycoproteins, small molecules, carbohydrates, lipids, glycolipids, lipoproteins, lipopolysaccharides, sugars, amino acids, organic dyes, and synthetic polymers.
[0283] Suitable lipids include, for example, 3-N-[(methoxypoly(ethylene glycol)2000)carbamoyl]-1,2-dimyristyloxy-propylamine (PEG-C-DMA), 1,2-dilinoleyloxy-N,N-dimethyl-3-aminopropane (DLinDMA), 1,2-distearyl-sn-glycero-3-phosphocholine (DSPC), cholesterol, dipalmitoylphosphatidylcholine, 3-N-[(w- ...dimyristyl-sn-glycero-3-phosphocholine (DSPC), cholesterol, dipalmitoylphosphatidylcholine, 3-N-[(w-methoxypoly(ethylene glycol)2000)carbamoyl]-1,2-dimyristyloxy-propylamine (PEG-C-DMA), 1,2-dimyristyloxy-N,N-dimethyl-3-aminopropane (DLinDMA), 1,2-dimyristyloxy-N,N-dimethyl-3-aminopropane (DLinDMA), 1,2-dimyristyloxy-N,N-dimethyl-3-aminopropane (DLinDMA), 1,2-dimyristyloxy-N,N-dimethyl-3-aminopropane (DLinDMA
[00] Carbamoyl]-1,2-dimyristyloxypropylamine, 1,2-dilinoleyloxy-3-N,N-dimethylaminopropane, 1,2-distearate-sn-glycero-3-phosphocholine, PEG-cDMA, 1,2-dilinoleyloxy-3-(N;N-dimethyl)aminopropane (DLinDMA), 2,2-dilinoleyl-4-dimethylaminoethyl-[1,3]-dioxolane (DLin-KC2-DMA), etc.
[0284] Suitable biomolecules include affinity moieties. Suitable affinity moieties include His5 (HHHHH) (SEQ ID NO:876); HisX6 (HHHHHH) (SEQ ID NO:877); c-myc (EQKLISEEDL) (SEQ ID NO:878); Flag (DYKDDDDK) (SEQ ID NO:879); StrepTag (WSHPQFEK) (SEQ ID NO:880); hemagglutinin, such as HA tag (YPYDVPDYA) (SEQ ID NO:881); glutathione S-transferase (GST); thioredoxin; cellulose-binding domain, RYIRS (SEQ ID NO:882); Phe-His-His-Thr (SEQ ID NO:883); chitin-binding domain; S-peptide; T7 peptide; SH2 domain; C-terminal RNA tag, WEAAAREACCRECCARA (SEQ ID NO:883). NO:884); metal-binding domains, such as zinc-binding or calcium-binding domains, as those derived from calcium-binding proteins, such as calmodulin, troponin C, calcineurin B, myosin light chain, recovery protein, S-regulatory protein, cone protein, VILIP, neurotrophin, hippocampal calcium-binding protein, aggregate protein, calcium chelate, calpain large subunit, S100 protein, parvalbumin, calcium-binding protein D9K, calcium-binding protein D28K, and calreticulin; biotin; streptavidin; MyoD; leucine zipper polypeptide; and maltose-binding protein. In some cases, the suitable biomolecule is biotin.
[0285] In some cases, a dimerization domain is a suitable biomolecule for conjugation to a target peptide. Non-limiting examples of suitable dimerization domains include peptides with the following dimerization pairs:
[0286] a) FK506-binding protein (FKBP) and FKBP;
[0287] b) FKBP and calcineurin catalytic subunit A (CnA);
[0288] c) FKBP and cyclic proteins;
[0289] d) FKBP and FKBP-rapamycin-associated protein (FRB);
[0290] e) Gyrase B (GyrB) and GyrB;
[0291] f) Dihydrofolate reductase (DHFR) and DHFR;
[0292] g) DmrB and DmrB;
[0293] h) PYL and ABI;
[0294] i) Cry2 and CIB1; and
[0295] j) GAI and GID1.
[0296] For example, in some cases, a biomolecule is a polypeptide containing at least about 85%, at least about 90%, at least about 95%, at least about 98%, or 100% amino acid sequence identity with the following amino acid FKBP amino acid sequence:
[0297] MGVQVETISPGDGRTFPKRGQTCVVHYTGMLEDGKKFDSSRDRNKPFKFMLGKQEVIRGWEEGVAQMSVGQRAKLTISPDYAYGATGHPGIIPPHATLVFDVELLKLE (SEQ ID NO: 885).
[0298] In some cases, the biomolecules suitable for conjugating to the target polypeptide are members of specific binding pairs. Specific binding pairs include, for example: i) antibody-antigen; ii) cell adhesion molecule-extracellular matrix; iii) ligand-receptor; iv) biotin-avidin; and so on.
[0299] Suitable synthetic polymers include, but are not limited to, polyalkylene compounds such as polyethylene and polypropylene and polyethylene glycol (PEG); polychloroprene; polyethylene ethers such as poly(vinyl acetate); polyhaloethylene compounds such as poly(vinyl chloride); polysiloxanes; polystyrene; polyurethanes; polyacrylates such as poly((meth)acrylate), poly((meth)acrylate), poly(((meth)acrylate) n-butyl acrylate), poly(((meth)acrylate) isobutyl acrylate), poly(((meth)acrylate) tert-butyl acrylate), poly(((meth)acrylate) hexyl acrylate), poly(((meth)acrylate) isodecyl acrylate), poly(((meth)acrylate) lauryl acrylate), poly(((meth)acrylate) phenyl acrylate), poly((meth)acrylate), poly(isopropyl acrylate), poly(isobutyl acrylate) and poly(octadecyl acrylate); polyacrylamides such as poly(acrylamide), poly(methacrylamide), poly(ethylacrylamide), poly(ethyl methacrylamide), poly(N-isopropylacrylamide), poly(n-, iso- and tert-butylacrylamide); and copolymers and mixtures thereof.
[0300] In some cases, the biomolecule to be conjugated to the target peptide is a peptide. Suitable peptides include, for example, fluorescent proteins; receptors; enzymes; structural proteins; affinity tags; and so on.
[0301] Suitable fluorescent proteins include, but are not limited to, green fluorescent protein (GFP) or its variants, blue fluorescent variant of GFP (BFP), cyan fluorescent variant of GFP (CFP), yellow fluorescent variant of GFP (YFP), enhanced GFP (EGFP), enhanced CFP (ECFP), enhanced YFP (EYFP), GFPS65T, Emerald, Topaz (TYFP), Venus, Citrine, mCitrine, GFPuv, destabilized EGFP (dEGFP), destabilized ECFP (dECFP), destabilized EYFP (dEYFP), mCFPm, Cerulean, T-Sapphire, CyPet, YPet, mKO, HcRed, t-HcRed, DsRed, DsRed2, DsRed-monomer, J-Red, dimer 2, t-dimer 2 (12), mRFP1, pocilloporin, Renilla GFP, Monster GFP, paGFP, Kaede protein and kindling protein, phycobiliproteins and phycobiliprotein conjugates (including B-phycoerythrin, R-phycoerythrin and allophycocyanin). Other examples of fluorescent proteins include mHoneydew, mBanana, mOrange, dTomato, tdTomato, mTangerine, mStrawberry, mCherry, mGrape1, mRaspberry, mGrape2, mPlum (Shaner et al. (2005) Nat. Methods 2:905-909). Any of the various fluorescent and colored proteins from coral species, as described, for example, in Matz et al. (1999) Nature Biotechnol. 17:969-973, is suitable for use.
[0302] In some cases, the biomolecule is an antibody. Suitable antibodies are described elsewhere in this article. Antibodies can be any antigen-binding antibody-based polypeptide, and their variety is well known in the art. In some cases, antibodies are single-chain Fv (scFv). Other antibody-based recognition domains (cAb VHH (cameloid antibody variable domain) and humanized forms, IgNAR VH (shark antibody variable domain) and humanized forms, sdAb VH (single-domain antibody variable domain) and "camelized" antibody variable domains) are also suitable. In some cases, T-cell receptor (TCR)-based recognition domains such as single-chain TCRs (scTv, single-chain double-domain TCRs containing Vα and Vβ) are also applicable.
[0303] Antibodies can be specific to antigens such as CD19, CD20, CD38, CD30, Her2 / neu, ERBB2, CA125, MUC-1, prostate-specific membrane antigen (PSMA), CD44 surface adhesion molecule, mesothelin, carcinoembryonic antigen (CEA), epidermal growth factor receptor (EGFR), EGFRvIII, vascular endothelial growth factor receptor-2 (VEGFR2), high molecular weight melanoma-associated antigen (HMW-MAA), MAGE-A1, IL-13R-a2, GD2, etc. In some cases, antibodies are specific to cytokines. In some cases, antibodies are specific to cytokine receptors. In some cases, antibodies are specific to growth factors. In some cases, antibodies are specific to growth factor receptors. In some cases, antibodies are specific to cell surface receptors. In some cases, antibodies are anti-CD3 antibodies.
[0304] In some cases, both the target molecule and the biomolecule are antibodies. In other cases, the target molecule is a first antibody specific to a first antigen, and the biomolecule is a second antibody specific to a second antigen. The first and second antigens can be completely separate molecules. For example, the first antigen can be a first polypeptide, and the second antigen can be a second polypeptide. The first antigen can be a first epitope displayed by the antigen, and the second antigen can be a second epitope displayed by the same antigen. The resulting conjugate can be a bispecific antibody.
[0305] In some cases, biomolecules endow target biomolecules with properties such as: i) increased serum half-life; ii) increased immunogenicity; iii) enhanced pharmacokinetic properties; iv) increased transport across the blood-brain barrier; and so on. For example, in some cases, the biomolecule that increases serum half-life is human serum albumin. In some cases, the biomolecule that increases serum half-life is the albumin-binding domain. In some cases, the biomolecule that increases serum half-life is a thyroxine transporter. In some cases, the biomolecule that increases serum half-life is a thyroxine-binding protein. In some cases, the biomolecule is an immunoglobulin Fc polypeptide. In some cases, the biomolecule that promotes transport across the blood-brain barrier is the transferrin receptor (TR), insulin receptor (HIR), insulin-like growth factor receptor (IGFR), low-density lipoprotein receptor-associated proteins 1 and 2 (LPR-1 and 2), diphtheria toxin receptor, llama single-domain antibody, protein transduction domain, TAT, penetrantin, or polyarginine peptide.
[0306] Suitable biomolecules include small molecules, such as cancer chemotherapeutic agents. Suitable cancer chemotherapeutic agents include, for example, alkylating agents such as nitrogen mustard (e.g., chlorambucil, nitrogen mustard, cyclophosphamide, ifosfamide, and melphalan); nitrosoureas (e.g., carmustine, formustine, lomustine, and streptozocin); platinum compounds (e.g., carboplatin, cisplatin, oxaliplatin, and BBR3464); busulfan; dacarbazine; nitrogen mustard; procarbazine; temozolomide; thiotepa; uramustine; antimetabolites. Examples of purines include folic acid (e.g., methotrexate, pemetrexed, and raltitrexed); purines (e.g., cladribine, clofarabine, fludarabine, mercaptopurine, and thioguanine); pyrimidines (e.g., capecitabine); vidarabine; fluorouracil; gemcitabine; and phytoalkaloids such as podophyllotoxin (e.g., etoposide and teniposide), taxanes (e.g., docetaxel and paclitaxel), and vinca alkaloids (e.g., vincristine, vinblastine, vinblastine, and vinblastine alkaloids). Vinorelbine; cytotoxic / antitumor antibiotics, such as anthracycline family members (e.g., daunorubicin, doxorubicin, epirubicin, idarubicin, mitoxantrone, and pentorubicin), bleomycin, rifampin, hydroxyurea, and mitomycin; topoisomerase inhibitors, such as topotecan and irinotecan; photosensitizers, such as aminolevulinic acid, methyl aminolevulinate, porphyrin sodium, and verteporfen; and other agents, such as alivitamin A. Acids, hexamethylmelamine, acridine, anagrelide, arsenic trioxide, asparaginase, axitinib, besalodin, bevacizumab, bortezomib, celecoxib, denitroleukin, erlotinib, estradiol, gefitinib, hydroxyurea, imatinib, lapatinib, pazopanib, pentostatin, masrophenol, mitotan, pegaspargase, tamoxifen, sorafenib, sunitinib, vemurafenib, vandetanib, and retinoic acid. For example, in some cases, the target molecule is an antibody; and the biomolecule is a cancer chemotherapeutic agent.
[0307] Suitable biomolecules include cytokines, chemokines, and peptide hormones. Examples of suitable biomolecules include, for instance, interferons (e.g., IFN-γ); interleukins (e.g., IL-1α, IL-1β, IL-2, IL-4, IL-5, IL-6, IL-7, IL-9, IL-10, IL-12p40, IL-12p70, IL-13, IL-15, IL-17, etc.); IP-10, KC, MCP-1, MIP-1α, MIP-1β, M-CSFMIP-2, MIG; α-chemokines (e.g., CXC chemokines; such as CXC-1 to CXC-17); β-chemokines (CC chemokines) such as RANTES or CCL20 (also known as MIP-3α); and tumor necrosis factor-α. (TNF-α); eosinophil chemotactic factor; granulocyte colony-stimulating factor (G-CSF); granulocyte-macrophage-colony-stimulating factor (GM-CSF); erythropoietin; insulin; Gro-α; Groβ; Gro-γ; stromal cell-derived factor; platelet-derived growth factor (PDGF); vascular endothelial growth factor (VEGF); insulin-like growth factor (IGF); fibroblast growth factor (FGF); epidermal growth factor (EGF); leukemia inhibitory factor (LIF); hepatocyte growth factor (HGF); thrombopoietin; etc.
[0308] Suitable biomolecules include nucleic acids. In some cases, nucleic acids are DNA molecules. In some cases, nucleic acids are RNA molecules. In some cases, nucleic acids contain both deoxyribonucleotides and ribonucleotides. In some cases, nucleic acids are single-stranded DNA molecules. In some cases, nucleic acids are double-stranded DNA molecules. In some cases, nucleic acids are single-stranded RNA molecules. Suitable nucleic acids include, for example, small interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, etc. Suitable nucleic acids include those that are siRNA or other RNA interference agents (RNAi agents or iRNA agents), shRNA, antisense oligonucleotides, self-cleaving RNA, ribozymes, fragments thereof and / or variants thereof (such as peptidyl transferase 23S rRNA, RNase P, type I and type II introns, GIR1 branched ribozymes, leadzymes, hairpin ribozymes, hammerhead ribozymes, HDV ribozymes, mammalian CPEB3 ribozymes, VS ribozymes, glmS ribozymes, CoTC ribozymes, etc.), microRNA, microRNA mimics, supermirs, aptamers, antimirs, antagomirs, Ul adaptors, triplet-forming oligonucleotides, RNA activators, long non-coding RNA, short non-coding RNA (e.g., piRNA), immunomodulatory oligonucleotides (e.g., immunostimulatory oligonucleotides, immunosuppressive oligonucleotides), GNA, LNA, ENA, PNA, TNA, HNA, TNA, XNA, HeNA, CeNA, morpholine, G-quadruplexes (RNA and DNA), antiviral oligonucleotides, and decoy oligonucleotides. Nucleic acids can have any length and can include one or more of the following: modified ribonucleotide bases, modified deoxyribonucleotide bases, modified deoxyribose, modified ribose, and modified backbone bonds (e.g., phosphate thioester bonds).
[0309] Biomolecules for conjugation with CRISPR-Cas effector peptides
[0310] In some cases, the biomolecule to be conjugated to the target peptide is a biomolecule suitable for conjugation to CRISPR-Cas effector peptides.
[0311] In some cases, biomolecules suitable for conjugation to CRISPR-Cas effector peptides are those that can regulate the transcription of target DNA (e.g., repress transcription, increase transcription). For example, in some cases, the biomolecule is a protein (or derived from a protein domain) that represses transcription (e.g., a transcription repressor, a protein that functions through the recruitment of transcription repressor proteins, modifications of target DNA such as methylation, recruitment of DNA modifiers, regulation of histones associated with the target DNA, recruitment of histone modifiers (such as those that modify histone acetylation and / or methylation). In other cases, the biomolecule is a protein (or derived from a protein domain) that increases transcription (e.g., a transcription activator, a protein that functions through the recruitment of transcription activator proteins, modifications of target DNA such as methylation, recruitment of DNA modifiers, regulation of histones associated with the target DNA, recruitment of histone modifiers (such as those that modify histone acetylation and / or methylation).
[0312] In some cases, biomolecules suitable for conjugation to CRISPR-Cas effector peptides are peptides with enzymatic activities that modify target nucleic acids (e.g., nuclease activities such as FokI nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deamination activity, superoxide dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, or glycosylation activity).
[0313] In some cases, biomolecules suitable for conjugation to CRISPR-Cas effector peptides are peptides that have enzymatic activities (e.g., methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristylation activity, or demyristylation activity) that modify peptides associated with target nucleic acids (e.g., histones).
[0314] Examples of proteins (or fragments thereof) that can be used to increase transcription and are suitable as biomolecules for conjugation to CRISPR-Cas effector peptides include, but are not limited to: transcription activators, such as VP16, VP64, VP48, VP160, p65 subdomains (e.g., from NFkB) and activation domains and / or TAL activation domains of EDLL (e.g., for activity in plants); histone lysine methyltransferases, such as SET1A, SET1B, MLL1 to 5, ASH1, SYMD2, NSD1, etc.; histone lysine demethylases, such as JHDM2a / b, UTX, JMJD3, etc.; histone acetyltransferases, such as GCN5, PCAF, CBP, p300, TAF1, TIP60 / PLIP, MOZ / MYST3, MORF / MYST4, SRC1, ACTR, P160, CLOCK, etc.; and DNA demethylases, such as 10-11 translocation (TET) dioxygenase 1. (TET1CD), TET1, DME, DML1, DML2, ROS1, etc.
[0315] Examples of proteins (or fragments thereof) that can be used to reduce transcription and are suitable as biomolecules for conjugation to CRISPR-Cas effector peptides include, but are not limited to: transcriptional repressors, such as Krüppel-associated boxes (KRAB or SKD); KOX1 repressor domains; Mad mSIN3 interaction domain (SID); ERF repressor domain (ERD), SRDX repressor domain (e.g., for repression in plants); histone lysine methyltransferases, such as Pr-SET7 / 8, SUV4-20H1, RIZ1, etc.; histone lysine demethylases, such as JMJD2A / JHDM3A, JMJD2B, JMJD2C / GASC1, JMJD2D, JARID1A / RBP2, JARID1B / PLU-1, JARID1C / SMCX, JARID1D / SMCY, etc.; histone lysine deacetylases, such as HDAC1, HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, SIRT1, SIRT2, HDAC11, etc.; DNA methyltransferases, such as HhaI DNA m5c-methyltransferase (M.HhaI), DNA methyltransferase 1, etc. DNA methyltransferase 3a (DNMT1), DNA methyltransferase 3b (DNMT3b), METI, DRM3 (plant), ZMET2, CMT1, CMT2 (plant), etc.; and peripheral recruitment elements, such as lamin A, lamin B, etc.
[0316] In some cases, the biomolecule to be conjugated to a CRISPR-Cas effector peptide possesses enzymatic activity that modifies the target nucleic acid (e.g., ssRNA, dsRNA, ssDNA, dsDNA). Examples of enzymatic activity that can be provided by a biomolecule include, but are not limited to: nuclease activity, such as that provided by restriction enzymes (e.g., FokI nuclease); methyltransferase activity, such as that provided by methyltransferases (e.g., HhaI DNA m5c-methyltransferase (M.HhaI), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3). Activities provided by (plants), ZMET2, CMT1, CMT2 (plants), etc.; demethylase activity, such as that provided by demethylases (e.g., 10-11 translocation (TET) dioxygenase 1 (TET1CD), TET1, DME, DML1, DML2, ROS1, etc.); DNA repair activity; DNA damage activity; deamination activity, such as that provided by deaminases (e.g., cytosine deaminases, such as rat APOBEC1); dismutase activity; alkylation activity; dehydrogenase ... Purine activity; oxidative activity; pyrimidine dimer-forming activity; integrase activity, such as the activity provided by integrase and / or dissociative enzymes (e.g., Gin convertases such as the overactive mutant GinH106Y of Gin convertase, human immunodeficiency virus type 1 integrase (IN), Tn3 dissociative enzyme, etc.); transposase activity; recombinase activity, such as the activity provided by recombinase (e.g., the catalytic domain of Gin recombinase); polymerase activity; ligase activity; helicase activity; photolyase activity and glycosylationase activity).
[0317] In some cases, the biomolecule to be conjugated to a CRISPR-Cas effector peptide has enzymatic activity that modifies proteins associated with target nucleic acids (e.g., histones, RNA-binding proteins, DNA-binding proteins, etc.). Examples of enzymatic activities (modification of proteins associated with target nucleic acids) that can be provided by biomolecules include, but are not limited to: methyltransferase activities, such as those provided by histone methyltransferases (HMTs) (e.g., mottled inhibitor 3-9 homolog 1 (SUV39H1, also known as KMT1A), autosomal histone lysine methyltransferase 2 (G9A, also known as KMT1C and EHMT2), SUV39H2, ESET / SETDB1, etc., SET1A, SET1B, MLL1 to 5, ASH1, SYMD2, NSD1, DOT1L, Pr-SET7 / 8, SUV4-20H1, EZH2, RIZ1); and demethylase activities, such as those provided by histone demethylases (e.g., lysine demethylase 1A). (KDM1A, also known as LSD1), JHDM2a / b, JMJD2A / JHDM3A, JMJD2B, JMJD2C / GASC1, JMJD2D, JARID1A / RBP2, JARID1B / PLU-1, JARID1C / SMCX, JARID1D / SMCY, UTX, JMJD3, etc.) activities; acetyltransferase activities, such as those provided by histone acetyltransferases (e.g., human acetyltransferase p300, GCN5, PCAF, CBP, TAF1, TIP60 / PLIP, MOZ / MYST3, MORF / MYST4, HBO1 / MYST2). Activities provided by catalytic cores / fragments such as HMOF / MYST1, SRC1, ACTR, P160, CLOCK, etc.; deacetylase activities, such as those provided by histone deacetylases (e.g., HDAC1, HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, SIRT1, SIRT2, HDAC11, etc.); kinase activities; phosphatase activities; ubiquitin ligase activities; deubiquitination activities; adenylation activities; deadenylation activities; SUMOylation activities; deSUMOylation activities; ribosylation activities; deribosylation activities; myristylation activities; and demyristylation activities.
[0318] In some cases, the biomolecule to be conjugated to the CRISPR-Cas effector peptide is a catalytically active endonuclease. For example, in some cases, the target peptide is a CRISPR-Cas effector peptide that is non-catalytically active (e.g., does not exhibit endonuclease activity) and retains target nucleic acid binding activity (when complexed with guide RNA); and the biomolecule to be conjugated to the CRISPR-Cas effector peptide is a catalytically active endonuclease. For example, in some cases, the catalytically active endonuclease is a FokI peptide. As a non-limiting example, in some cases, the biomolecule to be conjugated to the CRISPR-Cas effector peptide is a FokI nuclease comprising an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the FokI amino acid sequence provided below; wherein the FokI nuclease is about 195 amino acids to about 200 amino acids in length.
[0319] FokI nuclease amino acid sequence:
[0320] QLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPDGAIYTVGSPIDYGVIVDTKAYSGGYNLPIGQADEMQRYVEENQTRNKHINPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHITNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINF (SEQ ID NO:886).
[0321] In some cases, the biomolecule to be conjugated to a CRISPR-Cas effector peptide is a deaminase. In other cases, the target CRISPR-Cas effector peptide is non-catalytically active. Suitable deaminases include cytidine deaminase and adenosine deaminase.
[0322] A suitable adenosine deaminase is any enzyme capable of deaminating adenosine in DNA. In some cases, the deaminase is TadA deaminase.
[0323] In some cases, suitable adenosine deaminases contain an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following amino acid sequences:
[0324] MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD (SEQ ID NO:887)
[0325] In some cases, suitable adenosine deaminases contain an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following amino acid sequences:
[0326] MRRAFITGVFFLSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD (SEQ ID NO:888).
[0327] In some cases, suitable adenosine deaminases contain an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following Staphylococcus aureus TadA amino acid sequence:
[0328] MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAHAEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFK NLRANKKSTN: (SEQ ID NO:889)
[0329] In some cases, suitable adenosine deaminases contain an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following Bacillus subtilis TadA amino acid sequence:
[0330] MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEEQRSIAHAEMLVIDEACKALGTWRLEGATLYVTLEPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMNLLQEERFNHQAEVVSGVLEEECGMLSAFFRELRKKKKAARKNLSE (SEQ ID NO:890)
[0331] In some cases, suitable adenosine deaminases contain an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following Salmonella typhimurium TadA:
[0332] MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFFRMRRQEIKALKKADRAEGAGPAV (SEQ ID NO:891)
[0333] In some cases, suitable adenosine deaminases contain an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the TadA amino acid sequence of the following Shewanella putrefaciens:
[0334] MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEILCLRSAGKKLENYRLLDATLYITLEPCAMCAGAMVHSRIARVVYGARDEKTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQRAQQGIE (SEQ ID NO:892)
[0335] In some cases, suitable adenosine deaminases contain an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following Haemophilus influenzae F3031 TadA amino acid sequence:
[0336] MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTAHAEIIALRNGAKNIQNYRLLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLS TFFQKRREEKKIEKALLKSLSDK (SEQ ID NO:893)
[0337] In some cases, suitable adenosine deaminases contain an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following Caulobacter crescentus TadA amino acid sequence:
[0338] MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGPIAAHDPTAHAEIAAMRAAAAKLGNYRLTDLTLVVTLEPCAMCAGAISHARIGRVVFGADDPKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFFRARRKAKI (SEQ ID NO:894)
[0339] In some cases, suitable adenosine deaminases contain an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following Geobacter sulfurreducens TadA amino acid sequence:
[0340] MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSNDPSAHAEMIAIRQAARRSANWRLTGATLYVTLEPCLMCMGAIILARLERVVFGCYDPKGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRKKAKATPALFIDERKVPPEP (SEQ ID NO:895)
[0341] Cytidine deaminases suitable as biomolecules to be conjugated to CRISPR-Cas effector polypeptides include any enzyme capable of deaminating cytidine in DNA.
[0342] In some cases, cytidine deaminases are deaminases from the apolipoprotein B mRNA-editing complex (APOBEC) family of deaminases. In some cases, APOBEC family deaminases are selected from the following groups: APOBEC1 deaminase, APOBEC2 deaminase, APOBEC3A deaminase, APOBEC3B deaminase, APOBEC3C deaminase, APOBEC3D deaminase, APOBEC3F deaminase, APOBEC3G deaminase, and APOBEC3H deaminase. In some cases, cytidine deaminases are activation-induced deaminases (AIDs).
[0343] In some cases, suitable cytidine deaminases contain an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following amino acid sequences:
[0344] MDSLLMNRKFLYQFKNVRWAKGRRETYLCYVVKRRDSATSFSLDFGYLRNKNGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRGNPNLSLRIFTARLYFCEDRKAEPEGLRRLHRAGVQIAIMTFKDYFYCWNTFVENHERTFKAWEGLHENSVRLSRQLRRILLPLYEVDDLRDAFRTLGL (SEQ ID NO:896)
[0345] In some cases, a suitable cytidine deaminase is AID and contains an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following amino acid sequence: MDSLLMNRRK FLYQFKNVRW AKGRRETYLC YVVKRRDSAT SFSLDFGYLR NKNGCHVELLFLRYISDWDL DPGRCYRVTW FTSWSPCYDC ARHVADFLRG NPNLSLRIFT ARLYFCEDRK AEPEGLRRLHRAGVQIAIMT FKENHERTFK AWEGLHENSV RLSRQLRRIL LPLYEVDDLR DAFRTLGL (SEQ ID NO:897).
[0346] In some cases, a suitable cytidine deaminase is AID and contains an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following amino acid sequence: MDSLLMNRRK FLYQFKNVRW AKGRRETYLC YVVKRRDSAT SFSLDFGYLR NKNGCHVELLFLRYISDWDL DPGRCYRVTW FTSWSPCYDC ARHVADFLRG NPNLSLRIFT ARLYFCEDRK AEPEGLRRLHRAGVQIAIMT FKDYFYCWNT FVENHERTFK AWEGLHENSV RLSRQLRRIL LPLYEVDDLR DAFRTLGL (SEQ ID NO:898).
[0347] In some cases, the methods of this disclosure for conjugating biomolecules to CRISPR-Cas effector peptides are carried out in the presence of trehalose. The concentration of trehalose can be from 25 mM to about 100 mM (e.g., 25 mM to 50 mM, 50 mM to 100 mM). For example, in some cases, the methods of this disclosure for conjugating biomolecules to CRISPR-Cas effector peptides are carried out under the following conditions: 20 mM Tris HCl, 300 mM KCl, 50 mM trehalose, pH 7.0, 4°C for 1 hour; 10 μM CRISPR-Cas effector peptide.
[0348] target molecules
[0349] Suitable target molecules for modification include, but are not limited to, peptides, polynucleotides, carbohydrates, lipids, glycolipids, and glycopeptides. The target molecules to be modified according to the method of this disclosure contain or are modified to contain a phenolic or catechol moiety.
[0350] In some cases, the target molecule is a polypeptide (“target polypeptide”).
[0351] Target peptides that can be modified using the methods disclosed herein include, but are not limited to, enzymes, antibodies, structural peptides, receptor ligands, and receptors. Target peptides may include structural proteins; receptors; enzymes; cell surface proteins; proteins integrated with cellular function; proteins involved in catalytic activity; proteins involved in motility; proteins involved in helicase activity; proteins involved in metabolic processes (anabolism and catabolism); proteins involved in antioxidant activity; proteins involved in proteolysis; proteins involved in biosynthesis; proteins with kinase activity; proteins with oxidoreductase activity; proteins with transferase activity; proteins with hydrolytic enzyme activity; proteins with lyase activity; proteins with isomerase activity; proteins with ligase activity; proteins with enzyme regulatory activity; proteins with signal transduction activity; structural peptides; peptides with binding activity; receptor peptides; and so on. Proteins involved in cell movement; proteins involved in membrane fusion; proteins involved in cell communication; proteins involved in regulating biological processes; proteins involved in development; proteins involved in cell differentiation; proteins involved in stimulus responses; behavioral proteins; cell adhesion proteins; proteins involved in cell death; proteins involved in transport (including protein transporter activity, nuclear transport, ion transporter activity, channel transporter activity, etc.); proteins involved in secretion activity; proteins involved in electron transporter activity; proteins involved in pathogenesis; proteins involved in chaperone protein regulator activity; proteins with nucleic acid binding activity; proteins with transcriptional regulatory activity; proteins involved in extracellular tissues; proteins involved in biogenesis; proteins involved in translation regulation; and so on.
[0352] In some cases, the target peptide is an antibody. Antibodies can be any antigen-binding antibody-based peptide, and their variety is well-known in the art. In some cases, the antibody is a single-chain Fv (scFv). Other antibody-based recognition domains are suitable for use (cAb VHH (cameloid antibody variable domain) and humanized forms, IgNAR VH (shark antibody variable domain) and humanized forms, sdAb VH (single-domain antibody variable domain) and "camelized" antibody variable domains). In some cases, T-cell receptor (TCR)-based recognition domains such as single-chain TCRs (scTv, single-chain double-domain TCRs containing Vα and Vβ) are also applicable.
[0353] Antibodies can be specific to antigens such as CD19, CD20, CD38, CD30, Her2 / neu, ERBB2, CA125, MUC-1, prostate-specific membrane antigen (PSMA), CD44 surface adhesion molecule, mesothelin, carcinoembryonic antigen (CEA), epidermal growth factor receptor (EGFR), EGFRvIII, vascular endothelial growth factor receptor-2 (VEGFR2), high molecular weight melanoma-associated antigen (HMW-MAA), MAGE-A1, IL-13R-a2, GD2, etc. In some cases, antibodies are specific to cytokines. In some cases, antibodies are specific to cytokine receptors. In some cases, antibodies are specific to growth factors. In some cases, antibodies are specific to growth factor receptors. In some cases, antibodies are specific to cell surface receptors. In some cases, antibodies are anti-CD3 antibodies.
[0354] In some cases, the antibody is selected from: 806, 9E10, 3F8, 81C6, 8H9, abavoxib, abatacept, abcixib, arbiturumab, alilucumab, actosumarab, adalimumab, adenomyumab, adunumab, afemoxicillin, afutuzumab, pego-arasizumab, ALD518, afasicept, alenumab, alikumarab, pentiazem atomoxicillin, amatoxins, AMG. 102. Other names mentioned include: Maanmozumab, Rastar-Anettozumab, Anilurozumab, Anluzumab, Apozizumab, Asimozizumab, Avasuzumab, Asecizumab, Asceticip, Atezizumab, Atenumab, Tosizumab, Atemumab, AVE1642, Bapizumab, Baliximab, Baviximab, Betomozumab, Begolozumab, Belimumab, Benalizumab, Betemumab, Besoxumab, Bevacizumab, Belottosumab, Bisimab, Bimarumab, Bivatuzumab. Mertansine, Lantomoumab, Butoxetine, BMS-936559, Becocilizumab, Brentuximab, Brinumumab, Padarumab, Broxetine, Brinumumab, Canatumab, Mecanumab, Racantuzumab, Caracci, Carosumab, Pendivitin, Carrucizumab, Caputoxumab, cBR96-Doxorubicin Immunoconjugate, CC49, CDP791, Cilizumab, Pescizumab, Cetuximab, cG250, Ch.14.18, Poxita, Cetuximab, Crazazine, Kriximab, Titan-Krituzumab, Cotrastuzumab, Racotumumab, Canamumab, Concytidine, CP751871, CR6261, Crizotinib, CS-1008, Darcyline, Dacizumab, Darotto, Pego-Dapirizumab, Daratozumab, Detrickumab, Densizumab, Madenituzumab, Dinoxumab, Delotuzumab (Biotin), Democuzumab, Denutozumab, Delivocumab, Atordomulab, Trazituzumab, Dulitazumab, Dupilumab, Dvalilumab, Dustautumab, Emexici, Eculizumab, Ebazumab, Ezalozyme, Efazalumab, Efazalumab, Efazalumab, Efazalumab, Efazalumab, Efazalumab, Efazalumab, Efazalumab, Efazalumab, Efazalumab Imattocilizumab, Enartitumumab, Vitin-Entaftumab, Pevoxel, Enanttocilizumab, Enoxac, Enoxac, Enoxac, Entoxumab, Entoxumab, Ciepimibumab, Ipatizumab, Erlizumab, Ertoxumab, Enartitumab, Edarazine, Itraribumab, Ivexumab, Evoloxumab, Avira, F19, Fanolesomab, Faramolumab, Fatocilizumab, Fasinoxumab, FBTA05, Panvecilizumab, Fezanumab, Felatocilizumab, Fentotuzumab, Felixovtuzumab, Flantocilizumab, Flantocilizumab, Farantumab, Forlantocilizumab Furavirimumab, non-hematoxylin and oleanumab, furavirimumab, furavirimumab, vortoxicin, galiliximab, ganitinumab, gavimumab, givoxetine, givoxetine, givoxetine, veltin-gabatumumabvedotin, golimumab, goliximab, guseximab, HGS-ETR2, hu3S193, huA33, ibatumumab, teimoxetine, ilecurumab, idasarizumab, IGN101, IgN311, igovoxetine, IIIA4, IM-2C6, IMAAB362, imarumumab, IMC-A1 2. Inceximab, Imatrozumab, Iracurumab, Reintoximab, Vitin-Indutoximab, Inliximab, Enoximumab, Izazumab Ogamicin, Intolimumab, Ipilimumab, Itolimumab, Isaruzumab, Ilizumab, Icazumab, J591, KB004, Keliximumab, KW-2871, Labetizumab, Pembrolizumab, Lanpalizumab, Levofloxacin, Lemasoxumab, Lenzruzumab, Lesamumab, Riviremab, Vitin-Lifatoximab, Rigostazumab, Settatan Lilotomabsatetraxetan), lintozumab, lireruzumab, lodixazumab, logivir, mocin-lovotozumab, rukamumab, pego-rulizumab, rucizumab, rutozumab, mapumumab, magtoxizob, masmozumab, mastozumab, malfurizumab, MEDI4736, meprobamate, metitumumab, METMAB, melazomumab, sutetraxetan, sutetraxetan Cimetuzumab, Mitomumab, MK-0646, MK-3475, MM-121, Mogliflozin, MORAb-003, Moromumab, Morvezin, MOv18, Percetumab, MPDL33280A, Moromumab-CD3, Tanacozumab, Namerucumab, Eto-Naprotomumab, Nanatetumab, Natalizumab, Nebakumab, Netumumab, Nemolizumab, Neremomumab Nevasumab, Nitozumab, Nivolumab, Thionomumab, Otosacchari, Obintozumab, Ocarutuzumab, Orezumab, Odomozumab, Ofamumab, Olatozumab, Olozumab, Omaruzumab, Ontoxizolumab, Opinumab, Mortozumab, Ogovomab, Otetuzumab, Oxytozumab, Oletozumab, Oxytoxizolumab, Oxytoxizolumab, Ozanezumab, Olizumab Paxicillinumab, Palolizumab, Perlimumab, Pancomycinumab, Pebakumab, Persatutuzumab, Percoizumab, Pertuximab, Pertezolam, Pertrastuzumab, Pembrolizumab, Pemtumomab, Peracillin, Pertuzumab, Pexacillinumab, Pildizumab, Vitin-Pinatuzumab, Plinzamab, Pramupirocinumab, Vitin-Polatuzumab vedotin), penicillin, priliximab, retoxaximab, prilimumab, PRO 140, quinilizumab, R1507, rapomumab, raltrastuzumab, rividumab, ramucumab, ramucumab, ranibizumab, rapiccurumab, refanezumab, rengavir, ralliximab, rituximab, linusumab, rituximab, rotoximab, roletzumab, lomoxol, lonlizumab, rovezumab, rulizumab, gavitenac-sacituzumab, samazumab, sarilumab, pendimethalin, sarilumab, SCH900105, Sekunumab, Seretuzumab, Sertoxacillin, Sevemumab, SGN-CD19A, SGN-CD33A, Sirolizumab, Cifamumab, Cetoxicillin, Simutuzumab, Cializumab, Cilukumab, Vitin-Sofituzumab, Sulanzumab, Sorituzumab, Sonepizumab, Soltarizumab, Staluronumab, Thioxalumab, Suvemumab, Taberucizumab, Texazotuzumab Tetraxetan, Tadalafil, Talizumab, Tadalafil, Pertamozumab, Taretozumab, Tefavib, Atemozumab, Tetozumab, Tenexizumab, Telizumab, Tetozumab, Tetulomab, Tetulomab, TGN1412, Ticilimumab / Tremelimumab, Tegrazab, Tetrazumab, TNX-650, Tocilizumab, Tolizumab, Tosatosulimumab, Tosizumab, Tovitolumab, Traroluzumab, Traroluzumab Sizuzumab, TRBS07, trelizumab, trelimezumab, tregolumumab, simmo-interleukin monoclonal antibody, tovirumab, utuximab, urolumumab, urolumumab, utuzumab, ustekinumab, vetin-vantozumab, ventetrazumab, varenseluzumab, variliximab, variliximab, varilizumab, vetozumab, vetozumab, vepamomumab, vesenkumab, vexizumab, voloximumab, martin-vosertozumab, votomomumab, zalumumab, zamumab, zatuximab, ziramumab, and azomomumab.
[0355] In some cases, the target peptide is a CRISPR-Cas effector peptide. Suitable CRISPR-Cas effector peptides are class II CRISPR / Cas endonucleases, such as type II, type V, or type VI CRISPR-Cas effector peptides. In some cases, suitable RNA-directed endonucleases are class II CRISPR / Cas endonucleases. In some cases, suitable RNA-directed endonucleases are class II CRISPR / Cas endonucleases (e.g., Cas9 protein). In some cases, CRISPR-Cas effector peptides are class V CRISPR-Cas effector peptides (e.g., Cpf1, C2c1, or C2c3 proteins). In some cases, suitable CRISPR-Cas effector peptides are class VI CRISPR-Cas effector peptides (e.g., C2c2 protein; also known as "Cas13a" protein). CasX proteins are also suitable. CasY proteins are also suitable.
[0356] In some cases, the CRISPR / Cas effector peptide is a type II CRISPR / Cas effector peptide. In some cases, the CRISPR / Cas effector peptide is a Cas9 peptide. The Cas9 protein is guided to a target site (e.g., stabilized at the target site) within a target nucleic acid sequence (e.g., a chromosomal sequence or an extrachromosomal sequence, such as a free-body sequence, microcircular sequence, mitochondrial sequence, chloroplast sequence, etc.) by association with the protein-binding segment of the Cas9 guide RNA. In some cases, the Cas9 peptide contains an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 90%, at least 95%, at least 98%, at least 99%, or greater than 99% amino acid sequence identity with the *Streptococcus pyogenes* Cas9 shown in SEQ ID NO:753. In some cases, the Cas9 peptide contains an amino acid sequence shown in any one of SEQ ID NO:5-816. In some cases, the Cas9 polypeptide contains an amino acid sequence that is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or greater than 99% of the amino acid sequence shown in any of SEQ ID NO:5-816.
[0357] In some cases, the Cas9 polypeptide is the Staphylococcus aureus Cas9 (saCas9) polypeptide. In some cases, the saCas9 polypeptide contains an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the saCas9 amino acid sequence shown in SEQ ID NO:249.
[0358] In some cases, the Cas9 polypeptide is the Campylobacter jejuni Cas9 (CjCas9) polypeptide. CjCas9 recognizes 5′-NNNVRYM-3′ as a prespacer adjacent motif (PAM). The amino acid sequence of CjCas9 is given in SEQ ID NO:55. In some cases, a suitable Cas9 polypeptide comprises an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 90%, at least 95%, at least 98%, at least 99%, or greater than 99% amino acid sequence identity with the CjCas9 amino acid sequence shown in SEQ ID NO:55.
[0359] In some cases, a suitable Cas9 peptide is a high-fidelity (HF) Cas9 peptide. Kleinstiver et al. (2016) Nature 529:490. For example, amino acids N497, R661, Q695, and Q926 of the *Streptococcus pyogenes* Cas9 amino acid sequence (e.g., SEQ ID NO:5) are substituted with, for example, alanine. For example, an HF Cas9 peptide may contain an amino acid sequence having at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with *Streptococcus pyogenes* Cas9 (e.g., SEQ ID NO:5), wherein amino acids N497, R661, Q695, and Q926 are substituted with, for example, alanine. In some cases, a suitable Cas9 peptide exhibits altered PAM specificity. See, for example, Kleinstiver et al. (2015) Nature 523:481.
[0360] In some cases, a suitable Cas9 polypeptide contains an amino acid sequence that has at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following Cas9-HF1 sequences:
[0361]
[0362] In some cases, the suitable CRISPR / Cas effector peptide is a type V CRISPR / Cas effector peptide. In some cases, the type V CRISPR / Cas effector peptide is the Cpf1 protein. In some cases, the Cpf1 protein comprises an amino acid sequence that is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% identical to the amino acid sequence shown in any of SEQ ID NO: 818-822.
[0363] In some cases, suitable CRISPR / Cas effector peptides are CasX or CasY peptides. CasX and CasY peptides are described in Burstein et al. (2017) Nature 542:237.
[0364] In some cases, a suitable CRISPR / Cas effector peptide is a fusion protein comprising a CRISPR / Cas effector peptide fused to a heterologous peptide (also known as a "fusion partner"). In other cases, the CRISPR / Cas effector peptide is fused to an amino acid sequence that provides subcellular localization (the fusion partner), i.e., the fusion partner is a subcellular localization sequence (e.g., one or more nuclear localization signals (NLS), two or more NLS, three or more NLS, etc., for targeting the cell nucleus).
[0365] Nucleic acids that bind to two classes of CRISPR / Cas effector peptides (e.g., Cas9 protein; type V or VI CRISPR / Cas protein; Cpf1 protein) and target the complex to a specific location within the target nucleic acid are referred to herein as “guide RNA” or “CRISPR / Cas guided RNA”. Guide RNA provides target specificity to the complex (RNP complex) by including a targeting segment comprising a guide sequence (also referred to herein as the target sequence) that is a nucleotide sequence complementary to the target nucleic acid sequence.
[0366] In some cases, the guide RNA comprises two separate nucleic acid molecules: an "activator" and a "target," and is referred to herein as "double-molecule guide RNA," "two-molecule guide RNA," or "dgRNA." In other cases, the guide RNA is a single molecule (e.g., for some Class 2 CRISPR / Cas proteins, the corresponding guide RNA is a single molecule; and in some cases, the activator and target are covalently linked to each other, for example, by inserting nucleotides), and the guide RNA is referred to as "single-molecule guide RNA," "one-molecule guide RNA," or simply "sgRNA."
[0367] Two types of CRISPR / Cas effector peptides
[0368] In class 2 CRISPR systems, the function of effector complexes (e.g., cleavage of target DNA) is performed by a single endonuclease (e.g., see Zetsche et al. Cell. Oct 22, 2015; 163(3):759-71; Makarova et al., Nat Rev Microbiol. Nov 2015; 13(11):722-36; Shmakov et al., Mol Cell. Nov 5, 2015; 60(3):385-97); and Shmakov et al. (2017) Nature Reviews Microbiology 15:169. Therefore, the term “class 2 CRISPR / Cas protein” is used in this paper to encompass CRISPR / Cas effector peptides (e.g., target DNA cleavage proteins) from class 2 CRISPR systems. Therefore, as used herein, the term "class 2 CRISPR / Cas effector peptides" encompasses type II CRISPR / Cas effector peptides (e.g., Cas9); type VA CRISPR / Cas effector peptides (e.g., Cpf1 (also known as "Cas12a")); type VB CRISPR / Cas effector peptides (e.g., C2c1 (also known as "Cas12b")); and type VC CRISPR / Cas effector peptides (e.g., C2c3). (also known as "Cas12c")); V-U1 type CRISPR / Cas effector peptide (e.g., C2c4); V-U2 type CRISPR / Cas effector peptide (e.g., C2c8); V-U5 type CRISPR / Cas effector peptide (e.g., C2c5); V-U4 type CRISPR / Cas protein (e.g., C2c9); V-U3 type CRISPR / Cas effector peptide (e.g., C2c10); VI-A type CRISPR / Cas effector peptide (e.g., C2c2 (also known as "Cas13a")); VI-B type CRISPR / Cas effector peptide (e.g., Cas13b (also known as C2c4)); and VI-C type CRISPR / Cas effector peptide (e.g., Cas13c (also known as C2c7)). To date, the term "class 2 CRISPR / Cas effector peptides" encompasses type II, type V, and type VI CRISPR / Cas effector peptides, but it also refers to any class 2 CRISPR / Cas effector peptide suitable for binding to the corresponding guide RNA and forming an RNP complex.
[0369] Type II CRISPR / Cas endonucleases (e.g., Cas 9)
[0370] In the native type II CRISPR / Cas system, Cas9 acts as an RNA-guided endonuclease. This endonuclease uses a dual guide RNA with a crRNA and a trans-activating crRNA (tracrRNA) to target recognition and cleave via a mechanism involving two nuclease active sites in Cas9. These active sites together produce double-stranded DNA breaks (DSBs), or can produce single-stranded DNA breaks (SSBs) individually. The type II CRISPR endonuclease Cas9 and engineered dual guide RNA (dgRNA) or single guide RNA (sgRNA) form a ribonucleoprotein (RNP) complex that can target the desired DNA sequence. Guided by the dual RNA complex or the chimeric single guide RNA, Cas9 produces site-specific DSBs or SSBs within the target double-stranded DNA (dsDNA), which are repaired by non-homologous end joining (NHEJ) or homologous directed recombination (HDR).
[0371] Type II CRISPR / Cas effector polypeptides are a type of Class II CRISPR / Cas endonuclease. In some cases, the type II CRISPR / Cas endonuclease is the Cas9 protein. The Cas9 protein forms a complex with Cas9 guide RNA. The guide RNA provides target specificity to the Cas9 guide RNA complex by having a nucleotide sequence (guide sequence) complementary to the sequence (target site) of the target nucleic acid (as described elsewhere herein). The Cas9 protein of the complex provides site-specific activity. In other words, the Cas9 protein is guided to the target site (e.g., stabilized at the target site) within the target nucleic acid sequence (e.g., chromosomal sequence or extrachromosomal sequence, such as free sequence, microcircular sequence, mitochondrial sequence, chloroplast sequence, etc.) by association with the protein-binding segment of the Cas9 guide RNA.
[0372] Cas9 proteins can bind to and / or modify (e.g., cleavage, nicking, methylation, demethylation, etc.) target nucleic acids and / or peptides associated with the target nucleic acids (e.g., methylation or acetylation of histone tails) (e.g., when the Cas9 protein includes an active fusion partner). In some cases, Cas9 proteins are naturally occurring proteins (e.g., naturally occurring in bacterial and / or archaea cells). In other cases, Cas9 proteins are not naturally occurring peptides (e.g., Cas9 proteins are variant Cas9 proteins, chimeric proteins, etc.).
[0373] Examples of suitable Cas9 proteins include, but are not limited to, those shown in SEQ ID NO: 5-816. Naturally occurring Cas9 proteins bind to Cas9-guided RNA, thereby directing it to a specific sequence within a target nucleic acid (target site) and cleaving the target nucleic acid (e.g., cleaving dsDNA to produce a double-strand break, cleaving ssDNA, cleaving ssRNA, etc.). Chimeric Cas9 proteins are fusion proteins comprising a Cas9 polypeptide fused to a heterologous protein (called a fusion partner), wherein the heterologous protein provides the activity (e.g., activity not provided by the Cas9 protein). The fusion partner can provide activities such as enzymatic activities (e.g., nuclease activity, DNA and / or RNA methylation activity, DNA and / or RNA cleavage activity, histone acetylation activity, histone methylation activity, RNA modification activity, RNA binding activity, RNA splicing activity, etc.). In some cases, a portion of the Cas9 protein (e.g., the RuvC domain and / or HNH domain) exhibits reduced nuclease activity relative to the corresponding portion of the wild-type Cas9 protein (e.g., in some cases, the Cas9 protein is a nicking enzyme). In some cases, the Cas9 protein is enzymatically inactive or has reduced enzymatic activity relative to the wild-type Cas9 protein (e.g., relative to Streptococcus pyogenes Cas9).
[0374] In some cases, the fusion protein comprises: a) a catalytically inactive Cas9 protein (or other catalytically inactive CRISPR effector peptide); and b) a catalytically active endonuclease. For example, in some cases, the catalytically active endonuclease is a FokI peptide. As a non-limiting example, in some cases, the fusion protein comprises: a) a catalytically inactive Cas9 protein (or other catalytically inactive CRISPR effector peptide); and b) a FokI nuclease comprising an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the FokI amino acid sequence provided below; wherein the FokI nuclease is about 195 amino acids to about 200 amino acids in length.
[0375] FokI nuclease amino acid sequence:
[0376] QLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPDGAIYTVGSPIDYGVIVDTKAYSGGYNLPIGQADEMQRYVEENQTRNKHINPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHITNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINF (SEQ ID NO:900).
[0377] The analysis used to determine whether a given protein interacts with the Cas9 guide RNA can be any convenient binding assay that tests for the binding between the protein and the nucleic acid. Suitable binding assays (e.g., gel shift assays) will be known to those skilled in the art (e.g., assays involving the addition of Cas9 guide RNA and protein to the target nucleic acid).
[0378] The analysis used to determine whether a protein is active (e.g., to determine whether a protein has nuclease activity and / or some heterologous activity to cleave the target nucleic acid) can be any convenient analysis (e.g., any convenient nucleic acid cleavage analysis that tests nucleic acid cleavage). Suitable analyses (e.g., cleavage analyses) will be known to those skilled in the art and may include adding Cas9 guide RNA and protein to the target nucleic acid.
[0379] In some cases, a suitable Cas9 protein comprises amino acids 7-166 or 731-1003 of the Cas9 amino acid sequence shown in SEQ ID NO: 5, or an amino acid sequence having 60% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 99% or more, or 100% amino acid sequence identity with the corresponding portion of any of the amino acid sequences shown in SEQ ID NO: 6-816.
[0380] Examples of various Cas9 proteins (and Cas9 domain structures) and Cas9-guided RNAs (as well as information regarding requirements for prespacer adjacent motif (PAM) sequences present in target nucleic acids) can be found in the art, for example, see Jinek et al., Science. Aug 17, 2012; 337(6096):816-21; Chylinski et al., RNABiol. May 2013; 10(5):726-37; Ma et al., Biomed Res Int. 2013; 2013:270805; Hou et al., Proc Natl Acad Sci US A. Sep 24, 2013; 110(39):15644-9; Jinek et al., Elife. 2013; 2:e00471; Pattanayak et al., Nat Biotechnol. September 2013; 31(9):839-43; Qi et al., Cell. February 28, 2013; 152(5):1173-83; Wang et al., Cell. May 9, 2013; 153(4):910-8; Auer et al., Genome Res. October 31, 2013; Chen et al., Nucleic Acids Res. November 1, 2013; 41(20):e19; Cheng et al., Cell Res. October 2013; 23(10):1163-71; Cho et al., Genetics. November 2013; 195(3):1177-80; DiCarlo et al., Nucleic Acids Res. April 2013; 41(7):4336-43; Dickinson et al., Nat Methods. Oct 2013;10(10):1028-34; Ebina et al., Sci Rep. 2013;3:2510; Fujii et al., Nucleic Acids Res. Nov 1, 2013;41(20):e187; Hu et al., Cell Res. Nov 2013;23(11):1322-5; Jiang et al., Nucleic Acids Res. Nov 1, 2013;41(20):e188; Larson et al., Nat Protoc. Nov 2013;8(11):2180-96; Mali et al., Nat Methods. Oct 2013;10(10):957-63; Nakayama et al., Genesis.2013 Dec;51(12):835-43; Ran et al., Nat Protoc. 2013 Nov;8(11):2281-308; Ran et al., Cell. 2013 Sep 12;154(6):1380-9; Upadhyay et al., G3 (Bethesda). 2013 Dec 9;3(12):2233-8; Walsh et al., Proc Natl Acad Sci US A. 2013 Sep 24;110(39):15514-5; Xie et al., Mol Plant. 2013 Oct 9; Yang et al., Cell. 2013 Sep 12;154(6):1370-9; Briner et al., Mol Cell. October 23, 2014; 56(2):333-9; Shmakov et al., NatRev Microbiol.March 2017; 15(3):169-182; and U.S. patents and patent applications: 8,906,616; 8,895,308; 8,889,418; 8,889,356; 8,871,445; 8,865,406; 8,795,965; 8,771,945; 8,697,359; 20140068797; 20140170753; 20140179006; 20140179770; 201401868 43; 20140186919; 20140186958; 20140189896; 20140227787; 20140234972; 20140242664; 20140242699; 20140242700; 20140242702; 20140248702; 20140256046; 20140273037; 20140273226; 20140273230; 2014027323 1; 20140273232; 20140273233; 20140273234; 20140273235; 20140287938; 20140295556; 20140295557; 20140298547; 20140304853; 20140309487; 20140310828; 20140310830; 20140315985; 20140335063; 20140335620 References 20140342456; 20140342457; 20140342458; 20140349400; 20140349405; 20140356867; 20140356956; 20140356958; 20140356959; 20140357523; 20140357530; 20140364333; and 20140377868; each of these references is hereby incorporated by full reference.
[0381] Variant Cas9 protein - cleavage enzyme and dCas9
[0382] In some cases, Cas9 proteins are variant Cas9 proteins. When compared to the amino acid sequence of the corresponding wild-type Cas9 protein, variant Cas9 proteins have an amino acid sequence that differs by at least one amino acid (e.g., deletion, insertion, substitution, fusion). In some cases, variant Cas9 proteins have amino acid alterations that reduce the nuclease activity of the Cas9 protein (e.g., deletion, insertion, or substitution). For example, in some cases, variant Cas9 proteins have 50% or less, 40% or less, 30% or less, 20% or less, 10% or less, 5% or less, or 1% or less of the nuclease activity of the corresponding wild-type Cas9 protein. In some cases, variant Cas9 proteins are essentially devoid of nuclease activity. When a Cas9 protein is a variant Cas9 protein that is essentially devoid of nuclease activity, it may be referred to as a nuclease-deficient Cas9 protein or a “dCas9” of “dead” Cas9. Proteins that cleave one strand of a double-stranded target nucleic acid but not the other (e.g., class 2 CRISPR / Cas proteins, such as Cas9 proteins) are referred to herein as “cleavage enzymes” (e.g., “cleavage enzyme Cas9”).
[0383] In some cases, variant Cas9 proteins can cleave the complementary strand of the target nucleic acid (sometimes referred to in the art as the target strand), but have a reduced ability to cleave the non-complementary strand of the target nucleic acid (sometimes referred to in the art as the non-target strand). For example, variant Cas9 proteins may have mutations (amino acid substitutions) that reduce the function of the RuvC domain. Therefore, Cas9 proteins can be cleaving enzymes that cleave the complementary strand but not the non-complementary strand. As a non-limiting example, in some embodiments, the variant Cas9 protein has a mutation at the amino acid position corresponding to residue D10 (e.g., D10A, aspartic acid to alanine) of SEQ ID NO: 5 (or the corresponding position in any of the proteins shown in SEQ ID NO: 6-261 and 264-816), and thus can cleave the complementary strand of the double-stranded target nucleic acid, but has a reduced ability to cleave the non-complementary strand of the double-stranded target nucleic acid (therefore, when the variant Cas9 protein cleaves the double-stranded target nucleic acid, it results in a single-strand break (SSB) instead of a double-strand break (DSB)) (see, for example, Jinek et al., Science. 2012 Aug 17; 337(6096):816-21). See, for example, SEQ ID NO: 262.
[0384] In some cases, variant Cas9 proteins can cleave the non-complementary strand of the target nucleic acid, but have a reduced ability to cleave the complementary strand. For example, variant Cas9 proteins may have mutations (amino acid substitutions) that reduce the function of the HNH domain. Thus, Cas9 proteins can be cleaving enzymes that cleave the non-complementary strand but not the complementary strand. As a non-limiting example, in some embodiments, variant Cas9 proteins have mutations at the amino acid position corresponding to residue H840 of SEQ ID NO: 5 (e.g., the H840A mutation, histidine to alanine) (or the corresponding position in any protein as shown in SEQ ID NO: 6-261 and 264-816), and therefore can cleave the non-complementary strand of the target nucleic acid, but have a reduced ability to cleave (e.g., not cleave) the complementary strand of the target nucleic acid. Such Cas9 proteins have a reduced ability to cleave target nucleic acids (e.g., single-stranded target nucleic acids) but retain the ability to bind to target nucleic acids (e.g., single-stranded target nucleic acids). See, for example, SEQ ID NO: 263.
[0385] In some cases, variant Cas9 proteins exhibit reduced ability to cleave both the complementary and non-complementary strands of a double-stranded target nucleic acid. As a non-limiting example, in some cases, variant Cas9 proteins have mutations at amino acid positions corresponding to residues D10 and H840 of SEQ ID NO: 5 (e.g., D10A and H840A) (or corresponding residues of any protein as shown in SEQ ID NO: 6-261 and 264-816), resulting in a reduced ability of the polypeptide to cleave (e.g., not cleave) both the complementary and non-complementary strands of the target nucleic acid. Such Cas9 proteins have reduced ability to cleave target nucleic acids (e.g., single-stranded or double-stranded target nucleic acids) but retain the ability to bind to the target nucleic acid. Cas9 proteins that cannot cleave target nucleic acids (e.g., due to one or more mutations, such as in the catalytic domain of the RuvC and HNH domains) are referred to as “dead” Cas9 or simply “dCas9”. See, for example, SEQ ID NO: 264.
[0386] CRISPR / Cas endonucleases type V and VI
[0387] In some cases, suitable CRISPR / Cas effector peptides are type V or type VI CRISPR / Cas endonucleases (i.e., CRISPR / Cas effector peptides are type V or type VI CRISPR / Cas endonucleases) (e.g., Cpf1, C2c1, C2c2, C2c3). Type V and type VI CRISPR / Cas endonucleases are a type of class 2 CRISPR / Cas endonuclease. Examples of type V CRISPR / Cas endonucleases include, but are not limited to, Cpf1, C2c1, and C2c3. An example of a type VI CRISPR / Cas effector peptide is C2c2. In some cases, suitable CRISPR / Cas effector peptides are type V CRISPR / Cas endonucleases (e.g., Cpf1, C2c1, C2c3). In some cases, the type V CRISPR / Cas effector peptide is the Cpf1 protein. In some cases, suitable CRISPR / Cas effector peptides are type VI CRISPR / Cas endonucleases (e.g., Cas13a).
[0388] Similar to type II CRISPR / Cas endonucleases, type V and VI CRISPR / Cas endonucleases form complexes with corresponding guide RNAs. The guide RNA provides target specificity to the endonuclease-guide RNA RNP complex by having a nucleotide sequence (guide sequence) complementary to the target nucleic acid sequence (target site) (as described elsewhere herein). The endonuclease in the complex provides site-specific activity. In other words, the endonuclease is guided to its target site (e.g., stable at the target site) within the target nucleic acid sequence (e.g., chromosomal sequence or extrachromosomal sequence, such as free sequences, microcircular sequences, mitochondrial sequences, chloroplast sequences, etc.) by association with the protein-binding segment of the guide RNA.
[0389] Examples and guidelines related to type V and type VI CRISPR / Cas proteins (e.g., Cpf1, C2c1, C2c2, and C2c3 guide RNAs) can be found in the field, for example see Zetsche et al., Cell. Oct 22, 2015; 163(3):759-71; Makarova et al., Nat Rev Microbiol. Nov 2015; 13(11):722-36; Shmakov et al., MolCell. Nov 5, 2015; 60(3):385-97; and Shmakov et al. (2017) Nature Reviews Microbiology 15:169.
[0390] In some cases, type V or type VI CRISPR / Cas endonucleases (e.g., Cpf1, C2c1, C2c2, C2c3) possess enzymatic activity, for example, when type V or type VI CRISPR / Cas peptides cleave target nucleic acids upon binding to guide RNA. In other cases, type V or type VI CRISPR / Cas endonucleases (e.g., Cpf1, C2c1, C2c2, C2c3) exhibit reduced enzymatic activity relative to their corresponding wild-type type V or type VI CRISPR / Cas endonucleases (e.g., Cpf1, C2c1, C2c2, C2c3) while retaining DNA-binding activity.
[0391] In some cases, the type V CRISPR / Cas endonuclease is the Cpf1 protein. In some cases, the Cpf1 protein contains an amino acid sequence having at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% amino acid sequence identity with the Cpf1 amino acid sequence shown in any of SEQ ID NO: 818-822. In some cases, the Cpf1 protein comprises a continuous segment of 100 to 200 amino acids (aa), 200 to 400 aa, 400 to 600 aa, 600 to 800 aa, 800 to 1000 aa, 1000 to 1100 aa, 1100 to 1200 aa, or 1200 to 1300 aa having at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% amino acid sequence identical to the Cpf1 amino acid sequence shown in any of SEQ ID NO:818-822.
[0392] In some cases, the Cpf1 protein contains an amino acid sequence whose RuvCI domain has at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% amino acid sequence identity with the Cpf1 amino acid sequence shown in any of SEQ ID NO: 818-822. In some cases, the Cpf1 protein contains an amino acid sequence whose RuvCII domain has at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% amino acid sequence identity with the Cpf1 amino acid sequence shown in any of SEQ ID NO: 818-822. In some cases, the Cpf1 protein contains an amino acid sequence having at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% amino acid sequence identity with the RuvCI, RuvCII, and RuvCIII domains of the Cpf1 amino acid sequence shown in any of SEQ ID NO: 818-822.
[0393] In some cases, the Cpf1 protein exhibits reduced enzymatic activity relative to the wild-type Cpf1 protein (e.g., relative to the Cpf1 protein containing any of the amino acid sequences shown in SEQ ID NO: 818-822) while retaining DNA-binding activity. In some cases, the Cpf1 protein contains an amino acid sequence that is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% identical to the Cpf1 amino acid sequence shown in any of SEQ ID NO: 818; and contains an amino acid substitution (e.g., D→A substitution) at amino acid residue 917 corresponding to the Cpf1 amino acid sequence shown in SEQ ID NO: 818. In some cases, the Cpf1 protein comprises an amino acid sequence that is identical to the Cpf1 amino acid sequence shown in any one of SEQ ID NO:818-822 by at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% of the amino acid sequence; and contains an amino acid substitution (e.g., E→A substitution) at the amino acid residue corresponding to amino acid 1006 of the Cpf1 amino acid sequence shown in SEQ ID NO: 818. In some cases, the Cpf1 protein comprises an amino acid sequence that is identical to the Cpf1 amino acid sequence shown in any one of SEQ ID NO: 818-822 by at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% of the amino acid sequence; and contains an amino acid substitution (e.g., D→A substitution) at amino acid residue 1255 corresponding to the Cpf1 amino acid sequence shown in SEQ ID NO: 818.
[0394] In some cases, a suitable Cpf1 protein comprises an amino acid sequence having at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% amino acid sequence identity with the Cpf1 amino acid sequence shown in any of SEQ ID NO: 818-822.
[0395] In some cases, the type V CRISPR / Cas endonuclease is the C2c1 protein (examples include those shown in SEQ ID NO: 823-830). In some cases, the C2c1 protein comprises an amino acid sequence having at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% amino acid sequence identity with the C2c1 amino acid sequence shown in any of SEQ ID NO: 823-830. In some cases, the C2c1 protein comprises a continuous segment of 100 to 200 amino acids (aa), 200 to 400 aa, 400 to 600 aa, 600 to 800 aa, 800 to 1000 aa, 1000 to 1100 aa, 1100 to 1200 aa, or 1200 to 1300 aa having at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% amino acid sequence identity with the C2c1 amino acid sequence shown in any of SEQ ID NO: 823-830.
[0396] In some cases, the C2c1 protein comprises an amino acid sequence having at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% amino acid sequence identity with the RuvCII domain of the C2c1 amino acid sequence shown in any one of SEQ ID NO: 823-830. In some cases, the C2c1 protein comprises an amino acid sequence having at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% amino acid sequence identity with the RuvCII domain of the C2c1 amino acid sequence shown in any one of SEQ ID NO: 823-830. In some cases, the C2c1 protein comprises an amino acid sequence having at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% amino acid sequence identity with the RuvCI, RuvCII, and RuvCIII domains of the C2c1 amino acid sequence shown in any one of SEQ ID NO: 823-830.
[0397] In some cases, the type V CRISPR / Cas endonuclease is a C2c3 protein (examples include those shown in SEQ ID NO: 831-834). In some cases, the C2c3 protein comprises an amino acid sequence having at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% amino acid sequence identity with the C2c3 amino acid sequence shown in any of SEQ ID NO: 831-834. In some cases, the C2c3 protein comprises a continuous segment of 100 to 200 amino acids (aa), 200 to 400 aa, 400 to 600 aa, 600 to 800 aa, 800 to 1000 aa, 1000 to 1100 aa, 1100 to 1200 aa, or 1200 to 1300 aa having at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% amino acid sequence identity with the C2c3 amino acid sequence shown in any of SEQ ID NO: 831-834.
[0398] In some cases, the C2c3 protein comprises an amino acid sequence having at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% amino acid sequence identity with the RuvCII domain of the C2c3 amino acid sequence shown in any one of SEQ ID NO: 831-834. In some cases, the C2c3 protein comprises an amino acid sequence having at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% amino acid sequence identity with the RuvCI, RuvCII, and RuvCIII domains of the C2c3 amino acid sequence shown in any one of SEQ ID NO: 831-834.
[0399] In some cases, the C2c3 protein exhibits reduced enzymatic activity relative to the wild-type C2c3 protein (e.g., relative to the C2c3 protein containing any of the amino acid sequences shown in SEQ ID NO: 831-834) while retaining DNA-binding activity. In some cases, a suitable C2c3 protein contains an amino acid sequence having at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% amino acid sequence identity with the C2c3 amino acid sequence shown in any of SEQ ID NO: 831-834.
[0400] In some cases, type VI CRISPR / Cas endonucleases are C2c2 proteins (examples include those shown in SEQ ID NO: 835-846). In some cases, the C2c2 protein comprises an amino acid sequence having at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% amino acid sequence identity with the C2c2 amino acid sequence shown in any of SEQ ID NO: 835-846. In some cases, the C2c2 protein comprises a continuous segment of 100 to 200 amino acids (aa), 200 to 400 aa, 400 to 600 aa, 600 to 800 aa, 800 to 1000 aa, 1000 to 1100 aa, 1100 to 1200 aa, or 1200 to 1300 aa having at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% amino acid sequence identity with the C2c2 amino acid sequence shown in any of SEQ ID NO: 835-846.
[0401] In some cases, the C2c2 protein comprises an amino acid sequence having at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% amino acid sequence identity with the RuvCII domain of the C2c2 amino acid sequence shown in any one of SEQ ID NO: 835-846. In some cases, the C2c2 protein comprises an amino acid sequence having at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% amino acid sequence identity with the RuvCI, RuvCII, and RuvCIII domains of the C2c2 amino acid sequence shown in any one of SEQ ID NO: 835-846.
[0402] In some cases, the C2c2 protein exhibits reduced enzymatic activity relative to the wild-type C2c2 protein (e.g., relative to the C2c2 protein containing any of the amino acid sequences shown in SEQ ID NO: 835-846) while retaining DNA-binding activity. In some cases, a suitable C2c2 protein contains an amino acid sequence having at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 90%, or 100% amino acid sequence identity with the C2c2 amino acid sequence shown in any of SEQ ID NO: 835-846.
[0403] Examples and instructions relating to type V or VI CRISPR / Cas endonucleases (including domain structures) and guide RNA (as well as information regarding requirements relating to prespacer adjacent motif (PAM) sequences present in target nucleic acids) can be found in the art, for example see Zetsche et al., Cell. Oct 22, 2015; 163(3):759-71; Makarova et al., Nat Rev Microbiol. Nov 2015; 13(11):722-36; Shmakov et al., MolCell. Nov 5, 2015; 60(3):385-97; and Shmakov et al., Nat Rev Microbiol. March 2017; 15(3):169-182; and U.S. Patent and Patent Application: 9,580,701; 20170073695, 20170058272, 20160362668, 20160362667, 20160298078, 20160289637, 20160215300, 20160208243 and 20160208241, each of which is hereby incorporated by reference in its entirety.
[0404] CasX and CasY proteins
[0405] Suitable CRISPR / Cas effector peptides include CasX and CasY peptides. See, for example, Burstein et al. (2017) Nature 542:237. Suitable CasX peptides include those described in WO 2018 / 064371. Suitable CasY peptides include those described in WO 2018 / 064352.
[0406] CRISPR / Cas effector fusion peptide
[0407] In some cases, the CRISPR / Cas effector peptide is a CRISPR / Cas effector fusion peptide, which comprises: i) a CRISPR / Cas effector peptide; and ii) a hetero fusion partner.
[0408] In some cases, fusion couplers can regulate the transcription of target DNA (e.g., repress transcription, increase transcription). For example, in some cases, the fusion coupler is a protein (or a domain of a protein) that represses transcription (e.g., a transcription repressor, a protein that functions through the recruitment of transcription repressor proteins, modifications of target DNA such as methylation, recruitment of DNA modifiers, regulation of histones associated with the target DNA, recruitment of histone modifiers (such as those that modify histone acetylation and / or methylation). In other cases, the fusion coupler is a protein (or a domain of a protein) that increases transcription (e.g., a transcription activator, a protein that functions through the recruitment of transcription activator proteins, modifications of target DNA such as methylation, recruitment of DNA modifiers, regulation of histones associated with the target DNA, recruitment of histone modifiers (such as those that modify histone acetylation and / or methylation).
[0409] In some cases, CRISPR / Cas effector fusion peptides include heteropeptides with enzymatic activities that modify target nucleic acids (e.g., nuclease activities such as FokI nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deamination activity, superoxide dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, or glycosylation activity).
[0410] In some cases, CRISPR / Cas effector fusion peptides include heterologous peptides that have enzymatic activities (e.g., methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristylation activity, or demyristylation activity) that modify peptides associated with target nucleic acids (e.g., histones).
[0411] Examples of proteins (or fragments thereof) that can be used to increase transcription and are suitable as heterofusion partners include, but are not limited to: transcription activators, such as VP16, VP64, VP48, VP160, p65 subdomains (e.g., from NFkB) and the activation domains and / or TAL activation domains of EDLL (e.g., for activity in plants); histone lysine methyltransferases, such as SET1A, SET1B, MLL1 to 5, ASH1, SYMD2, NSD1, etc.; histone lysine demethylases, such as JHDM2a / b, UTX, JMJD3, etc.; histone acetyltransferases, such as GCN5, PCAF, CBP, p300, TAF1, TIP60 / PLIP, MOZ / MYST3, MORF / MYST4, SRC1, ACTR, P160, CLOCK, etc.; and DNA demethylases, such as 10-11 translocation (TET) dioxygenase 1. (TET1CD), TET1, DME, DML1, DML2, ROS1, etc.
[0412] Examples of proteins (or fragments thereof) that can be used to reduce transcription and are suitable as heterofusion partners include, but are not limited to: transcriptional repressors, such as Krüppel-associated boxes (KRAB or SKD); KOX1 repressor domains; Mad mSIN3 interaction domain (SID); ERF repressor domain (ERD), SRDX repressor domain (e.g., for repression in plants); histone lysine methyltransferases, such as Pr-SET7 / 8, SUV4-20H1, RIZ1, etc.; histone lysine demethylases, such as JMJD2A / JHDM3A, JMJD2B, JMJD2C / GASC1, JMJD2D, JARID1A / RBP2, JARID1B / PLU-1, JARID1C / SMCX, JARID1D / SMCY, etc.; histone lysine deacetylases, such as HDAC1, HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, SIRT1, SIRT2, HDAC11, etc.; DNA methyltransferases, such as HhaI DNA m5c-methyltransferase (M.HhaI), DNA methyltransferase 1, etc. DNA methyltransferase 3a (DNMT1), DNA methyltransferase 3b (DNMT3b), METI, DRM3 (plant), ZMET2, CMT1, CMT2 (plant), etc.; as well as peripheral recruitment elements, such as lamin A, lamin B, etc.
[0413] In some cases, the fusion partner possesses enzymatic activity for modifying target nucleic acids (e.g., ssRNA, dsRNA, ssDNA, dsDNA). Examples of enzymatic activities that can be provided by the fusion partner include, but are not limited to: nuclease activity, such as that provided by restriction enzymes (e.g., FokI nuclease); methyltransferase activity, such as that provided by methyltransferases (e.g., HhaI DNAm5c-methyltransferase (M.HhaI), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3 (plant), ZMET2, CMT1, CMT2 (plant), etc.); and demethylase activity, such as that provided by demethylases (e.g., 10-11 translocation (TET) dioxygenase 1). Activities provided by (TET1CD), TET1, DME, DML1, DML2, ROS1, etc.); DNA repair activity; DNA damage activity; deamination activity, such as that provided by deaminases (e.g., cytosine deaminases, such as rat APOBEC1); dismutase activity; alkylation activity; depurinase activity; oxidation activity; pyrimidine dimer formation activity; integrase activity, such as that provided by integrase and / or dissociation enzymes (e.g., Gin convertases such as the overactive mutant GinH106Y of Gin convertase, human immunodeficiency virus type 1 integrase (IN), Tn3 dissociation enzyme, etc.); transposase activity; recombinase activity, such as that provided by recombinases (e.g., the catalytic domain of Gin recombinase); polymerase activity; ligase activity; helicase activity; photolyase activity and glycosylation enzyme activity).
[0414] In some cases, the fusion partner possesses enzymatic activity that modifies proteins associated with target nucleic acids (e.g., histones, RNA-binding proteins, DNA-binding proteins, etc.). Examples of enzymatic activities (modifying proteins associated with target nucleic acids) that can be provided by the fusion partner include, but are not limited to: methyltransferase activity, such as that provided by histone methyltransferases (HMTs) (e.g., mottled inhibitor 3-9 homolog 1 (SUV39H1, also known as KMT1A), autosomal histone lysine methyltransferase 2 (G9A, also known as KMT1C and EHMT2), SUV39H2, ESET / SETDB1, etc., SET1A, SET1B, MLL1 to 5, ASH1, SYMD2, NSD1, DOT1L, Pr-SET7 / 8, SUV4-20H1, EZH2, RIZ1); and demethylase activity, such as that provided by histone demethylases (e.g., lysine demethylase 1A). (KDM1A, also known as LSD1), JHDM2a / b, JMJD2A / JHDM3A, JMJD2B, JMJD2C / GASC1, JMJD2D, JARID1A / RBP2, JARID1B / PLU-1, JARID1C / SMCX, JARID1D / SMCY, UTX, JMJD3, etc.) activities; acetyltransferase activities, such as those provided by histone acetyltransferases (e.g., human acetyltransferase p300, GCN5, PCAF, CBP, TAF1, TIP60 / PLIP, MOZ / MYST3, MORF / MYST4, HBO1 / MYST2). Activities provided by catalytic cores / fragments such as HMOF / MYST1, SRC1, ACTR, P160, CLOCK, etc.; deacetylase activities, such as those provided by histone deacetylases (e.g., HDAC1, HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, SIRT1, SIRT2, HDAC11, etc.); kinase activities; phosphatase activities; ubiquitin ligase activities; deubiquitination activities; adenylation activities; deadenylation activities; SUMOylation activities; deSUMOylation activities; ribosylation activities; deribosylation activities; myristylation activities; and demyristylation activities.
[0415] In some cases, the fusion protein comprises: a) a non-catalytically active CRISPR / Cas effector peptide (e.g., a non-catalytically active Cas9 peptide); and b) a catalytically active endonuclease. For example, in some cases, the catalytically active endonuclease is a FokI peptide. As a non-limiting example, in some cases, the fusion protein comprises: a) a non-catalytically active Cas9 protein (or other non-catalytically active CRISPR effector peptide); and b) a FokI nuclease comprising an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the FokI amino acid sequence provided below; wherein the FokI nuclease has a length of about 195 amino acids to about 200 amino acids.
[0416] FokI nuclease amino acid sequence:
[0417] QLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPDGAIYTVGSPIDYGVIVDTKAYSGGYNLPIGQADEMQRYVEENQTRNKHINPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHITNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINF (SEQ ID NO:901).
[0418] In some cases, the fusion partner is a deaminase. Therefore, in some cases, the CRISPR / Cas effector peptide fusion peptide comprises: a) a CRISPR / Cas effector peptide; and b) a deaminase. In some cases, the CRISPR / Cas effector peptide is non-catalytically active. Suitable deaminases include cytidine deaminase and adenosine deaminase.
[0419] A suitable adenosine deaminase is any enzyme capable of deaminating adenosine in DNA. In some cases, the deaminase is TadA deaminase.
[0420] In some cases, suitable adenosine deaminases contain an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following amino acid sequences:
[0421] MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD (SEQ ID NO:902)
[0422] In some cases, suitable adenosine deaminases contain an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following amino acid sequences:
[0423] MRRAFITGVFFLSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD (SEQ ID NO:903).
[0424] In some cases, suitable adenosine deaminases contain an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following Staphylococcus aureus TadA amino acid sequence:
[0425] MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAHAEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFK NLRANKKSTN: (SEQ ID NO:904)
[0426] In some cases, suitable adenosine deaminases contain an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following Bacillus subtilis TadA amino acid sequence:
[0427] MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEEQRSIAHAEMLVIDEACKALGTWRLEGATLYVTLEPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMNLLQEERFNHQAEVVSGVLEEECGMLSAFFRELRKKKKAARKNLSE (SEQ ID NO:905)
[0428] In some cases, suitable adenosine deaminases contain an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following Salmonella typhimurium TadA:
[0429] MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFFRMRRQEIKALKKADRAEGAGPAV (SEQ ID NO:906)
[0430] In some cases, suitable adenosine deaminases contain an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following Shewanella putrefactive bacteria TadA amino acid sequence:
[0431] MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEILCLRSAGKKLENYRLLDATLYITLEPCAMCAGAMVHSRIARVVYGARDEKTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQRAQQGIE (SEQ ID NO:907)
[0432] In some cases, suitable adenosine deaminases contain an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following Haemophilus influenzae F3031 TadA amino acid sequence:
[0433] MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTAHAEIIALRNGAKNIQNYRLLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLS TFFQKRREEKKIEKALLKSLSDK (SEQ ID NO:908)
[0434] In some cases, suitable adenosine deaminases contain an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following *Stylosus* TadA amino acid sequence:
[0435] MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGPIAAHDPTAHAEIAAMRAAAAKLGNYRLTDLTLVVTLEPCAMCAGAISHARIGRVVFGADDPKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRKAKI (SEQ ID NO:909)
[0436] In some cases, suitable adenosine deaminases contain an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following Geobacter sulfurreducens TadA amino acid sequence:
[0437] MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSNDPSAHAEMIAIRQAARRSANWRLTGATLYVTLEPCLMCMGAIILARLERVVFGCYDPKGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRKKAKATPALFIDERKVPPEP (SEQ ID NO:910)
[0438] Cytidine deaminases suitable for inclusion in CRISPR / Cas effector peptide fusion peptides include any enzyme capable of deaminating cytidine in DNA.
[0439] In some cases, cytidine deaminases are deaminases from the apolipoprotein B mRNA-editing complex (APOBEC) family of deaminases. In some cases, APOBEC family deaminases are selected from the following groups: APOBEC1 deaminase, APOBEC2 deaminase, APOBEC3A deaminase, APOBEC3B deaminase, APOBEC3C deaminase, APOBEC3D deaminase, APOBEC3F deaminase, APOBEC3G deaminase, and APOBEC3H deaminase. In some cases, cytidine deaminases are activation-induced deaminases (AIDs).
[0440] In some cases, suitable cytidine deaminases contain an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following amino acid sequences:
[0441] MDSLLMNRKFLYQFKNVRWAKGRRETYLCYVVKRRDSATSFSLDFGYLRNKNGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRGNPNLSLRIFTARLYFCEDRKAEPEGLRRLHRAGVQIAIMTFKDYFYCWNTFVENHERTFKAWEGLHENSVRLSRQLRRILLPLYEVDDLRDAFRTLGL (SEQ ID NO:911)
[0442] In some cases, a suitable cytidine deaminase is AID and contains an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following amino acid sequence: MDSLLMNRRK FLYQFKNVRW AKGRRETYLC YVVKRRDSAT SFSLDFGYLR NKNGCHVELLFLRYISDWDL DPGRCYRVTW FTSWSPCYDC ARHVADFLRG NPNLSLRIFT ARLYFCEDRK AEPEGLRRLHRAGVQIAIMT FKENHERTFK AWEGLHENSV RLSRQLRRIL LPLYEVDDLR DAFRTLGL (SEQ ID NO:912).
[0443] In some cases, a suitable cytidine deaminase is AID and contains an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity with the following amino acid sequence: MDSLLMNRRK FLYQFKNVRW AKGRRETYLC YVVKRRDSAT SFSLDFGYLR NKNGCHVELLFLRYISDWDL DPGRCYRVTW FTSWSPCYDC ARHVADFLRG NPNLSLRIFT ARLYFCEDRK AEPEGLRRLHRAGVQIAIMT FKDYFYCWNT FVENHERTFK AWEGLHENSV RLSRQLRRIL LPLYEVDDLR DAFRTLGL (SEQ ID NO: 913).
[0444] In some cases, CRISPR / Cas effector peptide fusion peptides contain CRISPR / Cas effector peptides that exhibit nicking enzyme activity. Suitable nicking enzymes are described elsewhere in this article.
[0445] In some cases, fusion CRISPR / Cas effector peptides contain one or more localization signal peptides. Suitable localization signals (“subcellular localization signals”) include, for example, nuclear localization signals (NLS) for targeting the cell nucleus; sequences for retaining the fusion protein outside the cell nucleus, such as nuclear export sequences (NES); sequences for retaining the fusion protein in the cytoplasm; mitochondrial localization signals for targeting mitochondria; chloroplast localization signals for targeting chloroplasts; endoplasmic reticulum (ER) retention signals; and ER export signals; and so on. In some cases, the fusion peptide does not include an NLS, such that the protein does not target the cell nucleus (this can be advantageous, for example, when the target nucleic acid is RNA present in the cytosol).
[0446] In some cases, the fusion polypeptide includes (fused to) a nuclear localization signal (NLS) (e.g., in some cases, 2 or more, 3 or more, 4 or more, or 5 or more NLS). Therefore, in some cases, the fusion polypeptide includes one or more NLS (e.g., 2 or more, 3 or more, 4 or more, or 5 or more NLS). In some cases, one or more NLS (2 or more, 3 or more, 4 or more, or 5 or more NLS) are localized at or near the N-terminus and / or C-terminus (e.g., within 50 amino acids). In some cases, one or more NLS (2 or more, 3 or more, 4 or more, or 5 or more NLS) are localized at or near the N-terminus (e.g., within 50 amino acids). In some cases, one or more NLS (2 or more, 3 or more, 4 or more, or 5 or more NLS) are localized at or near the C-terminus (e.g., within 50 amino acids). In some cases, one or more NLS (three or more, four or more, or five or more NLS) are located at or near both the N-terminus and the C-terminus (e.g., within 50 amino acids). In other cases, NLS are located at the N-terminus and NLS are located at the C-terminus.
[0447] In some cases, the fusion peptide comprises (fused to) 1 to 10 NLSs (e.g., 1-9, 1-8, 1-7, 1-6, 1-5, 2-10, 2-9, 2-8, 2-7, 2-6, or 2-5 NLSs). In other cases, the fusion peptide comprises (fused to) 2 to 5 NLSs (e.g., 2-4 or 2-3 NLSs).
[0448] Non-limiting examples of NLS include NLS sequences derived from: NLS of the SV40 viral large T antigen having the amino acid sequence PKKKRKV (SEQ ID NO:914); NLS from nucleoplasmic proteins (e.g., nucleoplasmic protein bipartite NLS having the sequence KRPAATKKAGQAKKKK (SEQ ID NO:915)); c-myc NLS having the amino acid sequence PAAKRVKLD (SEQ ID NO:916) or RQRRNELKRSP (SEQ ID NO:917); hRNPA1 M9 NLS having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO:918); the sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO:919) from the IBB domain of the input protein-α; and the sequence VSRKRPRP of the fibroid T protein (SEQ ID NO:919). Sequences of human p53 (SEQ ID NO:920) and PPKKARED (SEQ ID NO:921); sequence of human p53 (SEQ ID NO:922); sequence of mouse c-abl IV (SEQ ID NO:923); sequence of influenza virus NS1 (SEQ ID NO:924) and PKQKKRK (SEQ ID NO:925); sequence of hepatitis virus delta antigen (SEQ ID NO:926); sequence of mouse Mx1 protein (SEQ ID NO:927); sequence of human poly(ADP-ribose) polymerase (SEQ ID NO:928); and sequence of steroid hormone receptor (human) glucocorticoid (SEQ ID NO:929). In some cases, the NLS contains the amino acid sequence MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 930). Generally, the NLS (or multiple NLSs) has sufficient strength to drive the fusion peptide to accumulate in a detectable amount in the nucleus of a eukaryotic cell. Detection of the accumulation in the nucleus can be performed using any suitable technique. For example, a detectable marker can be fused to the fusion peptide, allowing visualization of its intracellular location. The nucleus can also be isolated from the cell, and its contents can then be analyzed using any suitable method for detecting proteins, such as immunohistochemistry, Western blotting, or enzyme activity assays. Accumulation in the nucleus can also be determined indirectly.
[0449] In some cases, CRISPR / Cas effector peptide fusion peptides include a “protein transduction domain” or PTD (also known as a CPP – cell-penetrating peptide), which refers to a peptide, polynucleotide, carbohydrate, or organic or inorganic compound that facilitates traversal of lipid bilayers, micelles, cell membranes, organelle membranes, or vesicle membranes. A PTD attached to another molecule (which can range from small polar molecules to large polymers and / or nanoparticles) facilitates membrane traversal, e.g., from extracellular space to intracellular space or from cytosol to organelles. In some embodiments, the PTD is covalently attached to the amino terminus of the peptide. In some embodiments, the PTD is covalently attached to the carboxyl terminus of the peptide. In some cases, the PTD is intercalated into the fusion peptide at a suitable insertion site (i.e., not at the N-terminus or C-terminus of the fusion peptide). In some cases, the subject fusion peptide includes (conjugated to, fused to) one or more PTDs (e.g., two or more, three or more, four or more PTDs). In some cases, the PTD includes nuclear localization signals (NLS) (e.g., in some cases, two or more, three or more, four or more, or five or more NLS). Therefore, in some cases, the fusion polypeptide includes one or more NLS (e.g., two or more, three or more, four or more, or five or more NLS). In some embodiments, the PTD is covalently linked to a nucleic acid (e.g., a guide nucleic acid, a polynucleotide encoding a guide nucleic acid, a polynucleotide encoding a fusion polypeptide, a donor polynucleotide, etc.).Examples of PTDs include, but are not limited to, a minimum eleven-amino acid polypeptide protein transduction domain (corresponding to residues 47-57 of HIV-1 TAT containing YGRKKRRQRRR; SEQ ID NO: 931); polyarginine sequences containing multiple arginines sufficient for direct cell entry (e.g., 3, 4, 5, 6, 7, 8, 9, 10, or 10-50 arginines); the VP22 domain (Zender et al. (2002) Cancer Gene Ther. 9(6):489-96); the Drosophila antennal foot gene protein transduction domain (Noguchi et al. (2003) Diabetes 52(7):1732-1737); a truncated human calcitonin peptide (Trehin et al. (2004) Pharm. Research 21:1248-1256); and polylysine (Wender et al. (2000) Proc. Natl. Acad. Sci.). USA 97:13003-13008); RRQRRTSKLMKR (SEQ ID NO:932); Transportan GWTLNSAGYLLGKINLKALAALAKKIL (SEQ ID NO:933); KALAWEAKLAKALAKALAKHLAKALAKALKCEA (SEQ ID NO:934); and RQIKIWFQNRRMKWKK (SEQ ID NO:934) NO:935). Exemplary PTDs include, but are not limited to: YGRKKRRQRRR (SEQ ID NO: 936); RKKRRQRRR (SEQ ID NO: 937); arginine homopolymers having 3 to 50 arginine residues; exemplary PTD domain amino acid sequences include, but are not limited to, any one of the following sequences: YGRKKRRQRRR (SEQ ID NO: 938); RKKRRQRR (SEQ ID NO: 939); YARAAARQARA (SEQ ID NO: 940); THRLPRRRRRR (SEQ ID NO: 941); and GGRRARRRRRR (SEQ ID NO: 942). In some embodiments, the PTD is an activatable CPP (ACPP) (Aguilera et al. (2009) Integr Biol (Camb) 6 ;1(5-6): 371-381). ACPPs comprise polycationic CPPs (e.g., Arg9 or "R9") connected to matching polyanions (e.g., Glu9 or "E9") via a cuttable connector, which reduces the net charge to near zero and thereby inhibits adhesion and uptake into cells.When the joint is cut, polyanions are released, locally exposing polyarginine and its inherent adhesiveness, thereby “activating” ACPP to traverse the membrane.
[0450] Guide RNA
[0451] When the target peptide is a CRISPR / Cas effector peptide, in some cases, the CRISPR / Cas effector peptide complexes with the CRISPR / Cas effector peptide guide RNA (also known as "CRISPR-Cas guide RNA").
[0452] Nucleic acid molecules that bind to CRISPR / Cas effector peptides and target the complex to a specific location within the target nucleic acid are referred to in this paper as “CRISPR / Cas effector peptide guide RNA” or simply “guide RNA”.
[0453] Guide RNA (can be said to consist of two segments: a first segment (referred to herein as the “targeting segment”) and a second segment (referred to herein as the “protein-binding segment”). A “segment” refers to a section / part / region of a molecule, such as a continuous nucleotide fragment in a nucleic acid molecule. A segment can also refer to a region / part of a complex such that it can contain more than one molecule. The “targeting segment” is also referred to herein as the “variable region” of the guide RNA. The “protein-binding segment” is also referred to herein as the “constant region” of the guide RNA. In some cases, the guide RNA is a Cas9 guide RNA.
[0454] The first segment (target segment) of the guide RNA comprises a nucleotide sequence (guide sequence) that is complementary (and thus hybridizes) to a specific sequence (target site) within the target nucleic acid (e.g., target ssRNA, target ssDNA, the complementary strand of a double-stranded target DNA, etc.). The protein-binding segment (or “protein-binding sequence”) interacts (binds) with a CRISPR / Cas effector polypeptide. The protein-binding segment of the guide RNA comprises two complementary nucleotide segments that hybridize to form a double-stranded RNA duplex (dsRNA duplex). Site-specific binding and / or cleavage of the target nucleic acid (e.g., genomic DNA) can occur at a location (e.g., the target sequence at a target locus) determined by the base pairing complementarity between the guide RNA (guide sequence of the guide RNA) and the target nucleic acid.
[0455] The guide RNA and the CRISPR / Cas effector peptide form a complex (e.g., via non-covalent interactions). The guide RNA provides target specificity to the complex by including a targeting segment comprising a guide sequence (a nucleotide sequence complementary to the target nucleic acid sequence). The CRISPR / Cas effector peptide of the complex provides site-specific activity (e.g., cleavage activity or activity provided by the CRISPR / Cas effector peptide when it is a CRISPR / Cas effector peptide fusion peptide, i.e., having a fusion partner). In other words, the CRISPR / Cas effector peptide is guided to the target nucleic acid sequence (e.g., target nucleic acid in chromosomal nucleic acids, such as chromosomes; target sequence in extrachromosomal nucleic acids, such as free nucleic acids, microcircles, ssRNA, ssDNA, etc.; target sequence in mitochondrial nucleic acids; target sequence in chloroplast nucleic acids; target sequence in plasmids; target sequence in viral nucleic acids; etc.) by virtue of its association with the guide RNA.
[0456] The "guide sequence," also known as the "target sequence" of the guide RNA, can be modified so that the guide RNA can target any desired sequence of any desired target nucleic acid with the CRISPR / Cas effector polypeptide, except for the prespacer adjacent motif (PAM) sequence, which may be considered. Thus, for example, the guide RNA may have a targeting segment that has a sequence complementary to (e.g., capable of hybridizing with) a sequence in nucleic acids in eukaryotic cells (e.g., viral nucleic acids, eukaryotic nucleic acids (e.g., eukaryotic chromosomes, chromosomal sequences, eukaryotic RNA, etc.)).
[0457] In some embodiments, the guide RNA comprises two separate nucleic acid molecules: an "activator" and a "target," and is referred to herein as "dual guide RNA," "bimolecular guide RNA," or "dgRNA." In some embodiments, the activator and target are covalently linked to each other (e.g., by inserting nucleotides), and the guide RNA is referred to as "single guide RNA," "Cas9 single guide RNA," "single-molecule Cas9 guide RNA," or simply "sgRNA."
[0458] The guide RNA comprises a crRNA-like molecule (“CRISPR RNA” / “target” / “crRNA” / “crRNA repeat”) and a corresponding tracrRNA-like molecule (“trans-acting CRISPR RNA” / “activator” / “tracrRNA”). The crRNA-like molecule (target) contains the targeting region (single-stranded) of the guide RNA and one half of the dsRNA double helix that forms the protein-binding region of the guide RNA (“double-strand forming region”). The corresponding tracrRNA-like molecule (activator / tracrRNA) contains the other half of the dsRNA double helix that forms the protein-binding region of the guide RNA (double-strand forming region). In other words, the nucleotide segment of the crRNA-like molecule is complementary to and hybridizes with the nucleotide segment of the tracrRNA-like molecule to form the dsRNA double helix of the protein-binding domain of the guide RNA. Therefore, it can be said that each target molecule has a corresponding activator molecule (which has a region that hybridizes with the target). The target molecule also provides the targeting region. Thus, the target molecule and the activator molecule (as corresponding pairs) hybridize to form the guide RNA. The precise sequence of a given crRNA or tracrRNA molecule is characteristic of the species in which the RNA molecule is found. Dual guide RNAs can include any corresponding activator and target pair.
[0459] As used herein, the term "activator" or "activator RNA" refers to a tracrRNA-like molecule with two guide RNAs (tracrRNA: "trans-acting CRISPR RNA") (and therefore, when "activator" and "target" are linked together by, for example, nucleotide insertion, it refers to a tracrRNA-like molecule with a single guide RNA). Thus, for example, a guide RNA (dgRNA or sgRNA) contains an activator sequence (e.g., a tracrRNA sequence). A tracr molecule (tracrRNA) is a naturally occurring molecule that hybridizes with a CRISPR RNA molecule (crRNA) to form two guide RNAs. The term "activator" is used herein to cover naturally occurring tracrRNAs, but also to tracrRNAs with modifications (e.g., truncation, sequence changes, base modifications, backbone modifications, linker modifications, etc.) in which the activator retains at least one function of the tracrRNA (e.g., a dsRNA duplex that facilitates binding to the Cas9 protein). In some cases, the activator provides one or more stem-loops that can interact with the Cas9 protein. Activators can be referred to as having a tracr sequence (tracrRNA sequence) and in some cases tracrRNA, but the term "activator" is not limited to naturally occurring tracrRNA.
[0460] As used herein, the term "target" or "target RNA" refers to a crRNA-like molecule with two guide RNAs (crRNA: "CRISPR RNA") (and therefore, when "activator" and "target" are linked together, for example, by an insertion nucleotide, it refers to a crRNA-like molecule with a single guide RNA). Thus, for example, a guide RNA (dgRNA or sgRNA) comprises a target region (which includes nucleotides that hybridize with (and are complementary to) the target nucleic acid) and a double-strand-forming region (e.g., the double-strand-forming region of crRNA, which may also be referred to as a crRNA repeat sequence). Because the user modifies the sequence of the target region (the region that hybridizes with the target sequence of the target nucleic acid) of the target RNA to hybridize with the desired target nucleic acid, the sequence of the target RNA is often a non-naturally occurring sequence. However, the double-strand-forming region of a target RNA that hybridizes with the double-strand-forming region of an activator (described in more detail below) may include a naturally occurring sequence (e.g., a sequence that may include the naturally occurring double-strand-forming region of crRNA, which may also be referred to as a crRNA repeat sequence). Therefore, the term "target" is used in this document to distinguish it from naturally occurring crRNA, although in fact, portions of the target (e.g., double-strand-forming segments) often include sequences derived from naturally occurring crRNA. However, the term "target" encompasses all naturally occurring crRNA.
[0461] Guide RNA can also be said to consist of three parts: (i) a target sequence (a nucleotide sequence that hybridizes with the target nucleic acid sequence); (ii) an activator sequence (as described above) (in some cases, called a tracr sequence); and (iii) a sequence that hybridizes with at least a portion of the activator sequence to form a double-stranded RNA. The target RNA has (i) and (iii); while the activator RNA has (ii).
[0462] Guide RNAs (e.g., dual or single guide RNAs) can include any corresponding activator and target pair. In some cases, the double-strand-forming region can be exchanged between the activator and the target. In other words, in some cases, the target consists of a nucleotide sequence from the double-strand-forming region of the tracrRNA (which is often part of the activator), while the activator consists of a nucleotide sequence from the double-strand-forming region of the crRNA (which is often part of the target).
[0463] As described above, a target RNA comprises a targeting segment (single strand) of the guide RNA and a nucleotide segment of one half of a dsRNA double strand that forms the protein-binding segment of the guide RNA (“double strand forming segment”). A corresponding tracrRNA-like molecule (activator) comprises the other half of a dsRNA double strand that forms the protein-binding segment of the guide RNA (double strand forming segment). In other words, the nucleotide segment of the target RNA is complementary to and hybridizes with the nucleotide segment of the activator to form a dsRNA double strand of the protein-binding segment of the guide RNA. Thus, it can be said that each target RNA has a corresponding activator (which has a region that hybridizes with the target RNA). The target molecule additionally provides a targeting segment. Therefore, the target RNA and the activator (as corresponding pairs) hybridize to form the guide RNA. The specific sequence of a given naturally occurring crRNA or tracrRNA molecule is characteristic of the species in which the RNA molecule is found. Examples of suitable activators and targets are well known in the art.
[0464] Nucleic acid modification
[0465] In some cases, CRISPR-Cas guide RNAs have one or more modifications (e.g., base modifications, backbone modifications, sugar modifications, etc.) to provide nucleic acids with new or enhanced characteristics (e.g., improved stability).
[0466] Suitable nucleic acid modifications include, but are not limited to: nucleotides modified with 2'O-methyl, nucleotides modified with 2'fluorine, nucleotides modified with locked nucleic acid (LNA), nucleotides modified with peptide nucleic acid (PNA), nucleotides with phosphate thioester bonds, and 5' caps (e.g., 7-methylguanylic acid cap (m7G)).
[0467] Suitable modified nucleic acid backbones containing phosphorus atoms include, for example, thiophosphates, chiral thiophosphates, dithiophosphates, phosphate triesters, aminoalkyl phosphate triesters, methyl and other alkyl phosphates including 3'-alkylene phosphates, 5'-alkylene phosphates and chiral phosphates, phosphonates, aminophosphates including 3'-aminoaminophosphates and aminoalkylaminophosphates, diaminophosphates, thiophosphoramides, thioalkyl phosphates, thioalkyl phosphate triesters, selenophosphates and borophosphates with normal 3'-5' bonds, their 2'-5' linked analogs, and those with antipolarity, wherein one or more nucleotide internucleotide bonds are 3'-3', 5'-5', or 2'-2' bonds. Suitable antipolar oligonucleotides contain a single 3'-3' bond at the 3' internucleotide bond, which can be a basic (nucleobase loss or substitution by a hydroxyl group) single antinucleoside residue. Various salts (e.g., potassium or sodium), mixed salts, and free acid forms are also included.
[0468] In some cases, CRISPR-Cas guide RNAs have one or more nucleotides linked by phosphate-thioate bonds (i.e., the subject nucleic acid has one or more phosphate-thioate bonds). The sulfur atom in a phosphate-thioate (PS) bond (i.e., phosphorothioate linkage) replaces a non-bridging oxygen atom in the phosphate backbone of the nucleic acid (e.g., an oligonucleotide). This modification makes the internucleotide bonds resistant to nuclease degradation. Phosphothioate bonds can be introduced between the last 3-5 nucleotides at the 5' or 3' end of the oligonucleotide to inhibit exonuclease degradation. Including phosphate-thioate bonds within the oligonucleotide (e.g., throughout the entire oligonucleotide) can also help reduce endonuclease attack.
[0469] Also suitable are CRISPR-Cas guide RNAs having a morpholino backbone structure as described, for example, in U.S. Patent No. 5,034,506. For example, in some embodiments, the CRISPR-Cas guide RNA comprises a 6-membered morpholino ring instead of a ribose ring. In some embodiments of these embodiments, a diaminophosphate or other non-phosphodiester nucleoside internucleotide bond replaces the phosphodiester bond.
[0470] CRISPR-Cas guide RNA may also include one or more substituted sugar moieties. Suitable polynucleotides contain sugar substituents selected from the following: OH; F; O-, S-, or N-alkyl; O-, S-, or N-alkenyl; O-, S-, or N-alkynyl; or O-alkyl-O-alkyl, wherein the alkyl, alkenyl, and alkynyl groups may be substituted or unsubstituted C1 to C2 groups. 10 Alkyl or C2 to C 10 Alkenyl and ynyl groups. Particularly suitable is: O((CH2) n O) m CH3, O(CH2) n OCH3, O(CH2) n NH2, O(CH2) n CH3, O(CH2) n ONH2 and O(CH2) n ON((CH2) n CH3)2, where n and m are 1 to approximately 10. Other suitable polynucleotides contain sugar substituents selected from the following: C1 to C2. 10Lower alkyl groups, substituted lower alkyl groups, alkenyl groups, alkynyl groups, aryl groups, O-aryl or O-aryl groups, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2, heterocyclic alkyl groups, heterocyclic aryl groups, aminoalkylamino groups, polyalkylamino groups, substituted silyl groups, RNA cleaving groups, reporter groups, intercalators, groups that improve the pharmacokinetic properties of oligonucleotides, or groups that improve the pharmacodynamic properties of oligonucleotides, and other substituents with similar properties. Suitable modifications include 2'-methoxyethoxy (2'-O-CH2CH2OCH3, also known as 2'-O-(2-methoxyethyl) or 2'-MOE) (Martin et al., Helv. Chim. Acta, 1995, 78, 486-504), i.e. alkoxyalkoxy. Other suitable modifications include 2'-dimethylaminoethoxy, i.e., the O(CH2)2ON(CH3)2 group, also known as 2'-DMAOE, as described in the examples below; and 2'-dimethylaminoethoxyethoxy (also known in the art as 2'-O-dimethyl-amino-ethoxy-ethyl or 2'-DMAEOE), i.e., 2'-O-CH2-O-CH2-N(CH3)2.
[0471] Method of coupling two proteins using a coupling protein
[0472] This disclosure provides a method for chemically selectively coupling a first polypeptide and a second polypeptide via a coupling peptide. The chemically selectively coupled product may comprise, in order from N-terminus to C-terminus: i) the first polypeptide; ii) the coupling polypeptide; and iii) the second polypeptide. As described above, this method utilizes the substrate preference of tyrosinase polypeptides.
[0473] For example, in some cases, this disclosure provides a method for chemically selectively coupling a first polypeptide and a second polypeptide to a coupling polypeptide, the method comprising: a) contacting the first polypeptide with the coupling polypeptide to produce a first polypeptide-coupling polypeptide conjugate, wherein the first polypeptide comprises a thiol moiety (e.g., Cys, wherein the Cys may be located at any solvent-accessible site within the first polypeptide), wherein the coupling polypeptide comprises an N-terminal reactive moiety covalently bonded to the thiol moiety present in the first polypeptide, wherein the coupling polypeptide comprising the N-terminal reactive moiety is produced by reacting a polypeptide comprising an N-terminal phenolic or catechol moiety and a C-terminal phenolic or catechol moiety (“coupling precursor polypeptide”) with a first enzyme capable of oxidizing the N-terminal phenolic or catechol moiety but not the C-terminal phenolic or catechol moiety to produce the N-terminal reactive moiety; and wherein the coupling polypeptide comprises ...e.g., Cys) The first polypeptide-coupled polypeptide conjugate comprises two or more positively charged or neutral amino acids within the ten amino acids of the catechin moiety and two or more negatively charged amino acids within the ten amino acids of the C-terminal phenol or catechol moiety; and b) contacting the second polypeptide with the first polypeptide-coupled polypeptide conjugate, wherein the second polypeptide comprises a thiol moiety (e.g., Cys, wherein the Cys may be located at any solvent-accessible position within the second polypeptide), wherein the first polypeptide-coupled polypeptide conjugate comprises a C-terminal reactive moiety covalently bonded to the thiol moiety present in the second polypeptide, wherein the first polypeptide-coupled polypeptide conjugate comprising the C-terminal reactive moiety is produced by reacting the first polypeptide-coupled polypeptide conjugate with a second enzyme capable of oxidizing the C-terminal phenol or catechol moiety to produce the C-terminal reactive moiety; and wherein the contact produces a first polypeptide-coupled polypeptide-second polypeptide conjugate. In some cases, the first enzyme is composed of... Figure 8 or Figure 9 The abTYR amino acid sequence described herein is a tyrosinase polypeptide with at least 75% amino acid sequence identity; and the second enzyme is a tyrosinase polypeptide containing amino acids that are identical to the amino acid sequence described herein. Figures 10A to 10Z and Figures 10AA to 10VV A tyrosinase polypeptide whose amino acid sequence has at least 75% amino acid sequence identity for any one of the amino acid sequences described by either of the above.
[0474] As another example, in some cases, this disclosure provides a method for chemically selectively coupling a first polypeptide and a second polypeptide to a coupling polypeptide, the method comprising: a) contacting the first polypeptide with the coupling polypeptide to produce a first polypeptide-coupling polypeptide conjugate, wherein the first polypeptide comprises a thiol moiety, wherein the coupling polypeptide comprises an N-terminal reactive moiety covalently bonded to the thiol moiety present in the first polypeptide, wherein the coupling polypeptide comprising the N-terminal reactive moiety is produced by reacting a polypeptide comprising an N-terminal phenolic or catechol moiety and a C-terminal phenolic or catechol moiety with a first enzyme capable of oxidizing the N-terminal phenolic or catechol moiety but not the C-terminal phenolic or catechol moiety to produce the N-terminal reactive moiety; wherein the coupling polypeptide comprises the N-terminal phenolic or catechol moiety in the ten The first polypeptide contains two or more negatively charged amino acids and two or more positively charged or neutral amino acids within ten amino acids of the C-terminal phenolic or catechol moiety; and b) contacting the second polypeptide with the first polypeptide-coupled polypeptide conjugate, wherein the second polypeptide contains a thiol moiety, wherein the first polypeptide-coupled polypeptide conjugate contains a C-terminal reactive moiety, the C-terminal reactive moiety forming a covalent bond with the thiol moiety present in the second polypeptide, wherein the first polypeptide-coupled polypeptide conjugate containing the C-terminal reactive moiety is produced by reacting the first polypeptide-coupled polypeptide conjugate with a second enzyme capable of oxidizing the C-terminal phenolic or catechol moiety to produce the C-terminal reactive moiety; and wherein the contact produces a first polypeptide-coupled polypeptide-second polypeptide conjugate. In some cases, the first enzyme is composed of... Figures 10A to 10Z and Figures 10AA to 10VV A tyrosinase polypeptide having at least 75% amino acid sequence identity for any one of the amino acid sequences described by either of the two enzymes; and b) the second enzyme is a tyrosinase polypeptide containing amino acid sequences that are identical to those described by either of the two enzymes. Figure 8 or Figure 9 The abTYR amino acid sequence depicted is a tyrosinase polypeptide with at least 75% amino acid sequence identity.
[0475] Coupled peptides can have a length of 10 to 100 amino acids, or more than 100 amino acids. In some cases, coupled peptides have a length of 10 to 25 amino acids. In some cases, coupled peptides have a length of 25 to 50 amino acids. In some cases, coupled peptides have a length of 50 to 100 amino acids. In some cases, coupled peptides have a length of more than 100 amino acids; for example, in some cases, coupled peptides have a length of 100 to 200 amino acids, 200 to 500 amino acids, or more than 500 amino acids (e.g., 500 to 1000, 1000 to 2000, or more than 2000 amino acids). In some cases, both the N-terminal and C-terminal phenolic moieties are tyrosine residues, and the enzyme that produces the reactive moieties is a tyrosinase.
[0476] As described above, in some cases, the coupled polypeptide comprises: a) two or more negatively charged amino acids within ten amino acids of the N-terminal phenolic or catechol moiety; or b) two or more negatively charged amino acids within ten amino acids of the C-terminal phenolic or catechol moiety. In this case, the coupled polypeptide may comprise: a) 2, 3, 4, 5, 6, 7, 8, 9, or 10 negatively charged amino acids within ten amino acids of the N-terminal phenolic or catechol moiety; or b) 2, 3, 4, 5, 6, 7, 8, 9, or 10 negatively charged amino acids within ten amino acids of the C-terminal phenolic or catechol moiety. As a non-limiting example, the coupled polypeptide comprises the following amino acid sequence: YEEEE(X) n RRRRY (SEQ ID NO: 961), where X is any amino acid, and where n is an integer from 0 to 40 (e.g., where n is an integer from 0 to 5, 5 to 10, 10 to 15, 15 to 20, 20 to 25, 25 to 30, 30 to 35, or 35 to 40). As another non-limiting example, the coupled polypeptide comprises the following amino acid sequence: YDDDD(X) n KKKKY (SEQ ID NO: 962), where X is any amino acid, and where n is an integer from 0 to 40 (e.g., where n is an integer from 0 to 5, 5 to 10, 10 to 15, 15 to 20, 20 to 25, 25 to 30, 30 to 35, or 35 to 40).
[0477] As described above, in some cases, the coupled polypeptide comprises: a) two or more positively charged amino acids within ten amino acids of the N-terminal phenolic or catechol moiety; or b) two or more positively charged amino acids within ten amino acids of the C-terminal phenolic or catechol moiety. In this case, the coupled polypeptide may comprise: a) 2, 3, 4, 5, 6, 7, 8, 9, or 10 positively charged amino acids within ten amino acids of the N-terminal phenolic or catechol moiety; or b) 2, 3, 4, 5, 6, 7, 8, 9, or 10 positively charged amino acids within ten amino acids of the C-terminal phenolic or catechol moiety. As a non-limiting example, the coupled polypeptide comprises the following amino acid sequence: YKKKK(X) n DDDDY (SEQ ID NO: 963), where X is any amino acid, and where n is an integer from 0 to 40 (e.g., where n is an integer from 0 to 5, 5 to 10, 10 to 15, 15 to 20, 20 to 25, 25 to 30, 30 to 35, or 35 to 40). As another non-limiting example, the coupled polypeptide comprises the following amino acid sequence: YRRRR(X) n EEEEY (SEQ ID NO: 964), where X is any amino acid, and where n is an integer from 0 to 40 (e.g., where n is an integer from 0 to 5, 5 to 10, 10 to 15, 15 to 20, 20 to 25, 25 to 30, 30 to 35, or 35 to 40).
[0478] As described above, this disclosure provides a conjugated polypeptide. This disclosure provides a composition comprising the conjugated polypeptide of this disclosure. This disclosure provides a composition comprising: a) the conjugated polypeptide of this disclosure; and b) a buffer solution.
[0479] Suitable first and second peptides include any of the peptides described above. For example, in some cases, the first and / or second peptide is an antibody (e.g., a single-chain antibody). As another example, in some cases, the first and / or second peptide is a CRISPR / Cas effector peptide. As another example, in some cases, the first peptide is a CRISPR / Cas effector peptide, and the second peptide is an IgFc peptide. As another example, in some cases, the first peptide is a CRISPR / Cas effector peptide, and the second peptide is a nanobody. As another example, in some cases, the first peptide is a CRISPR / Cas effector peptide, and the second peptide is an scFv peptide.
[0480] Methods for conjugating two or more polypeptides
[0481] This disclosure provides a method for sequentially coupling two or more polypeptides to each other. As described above, this method utilizes the substrate preference of tyrosinase polypeptides. This method can be performed on an insoluble matrix, i.e., a fixed surface, such as beads. Figures 39A to 39G The present disclosure schematically describes a method for sequentially coupling two or more polypeptides to each other.
[0482] Therefore, this disclosure provides a method for covalently linking a first polypeptide to a second polypeptide, the method comprising: a) contacting the first polypeptide with a fixed reactive portion, wherein the fixed reactive portion is generated by reacting a fixed phenolic or catechol portion with a first enzyme, wherein the first enzyme is capable of oxidizing the fixed phenolic or catechol portion to generate the fixed reactive portion, wherein the first polypeptide comprises: i) a thiol portion; and ii) a phenolic or catechol portion, wherein the first polypeptide comprises two or more negatively charged amino acids within ten amino acids of the phenolic or catechol portion, and wherein the fixed reactive portion forms a covalent bond with the thiol portion present in the first polypeptide to generate the fixed first polypeptide; b) Contacting the fixed first polypeptide with a second enzyme, wherein the second enzyme is capable of oxidizing the phenolic or catechol moiety present in the first polypeptide to produce a fixed first polypeptide comprising a reactive moiety; and c) Contacting the fixed first polypeptide comprising a reactive moiety with a second polypeptide, wherein the second polypeptide comprises: i) a thiol moiety; and ii) a phenolic or catechol moiety, wherein the second polypeptide comprises two or more neutral or positively charged amino acids within ten amino acids of the phenolic or catechol moiety, wherein the reactive moiety present in the fixed first polypeptide forms a covalent bond with the thiol moiety present in the second polypeptide, thereby producing a fixed conjugate comprising the first polypeptide covalently linked to the second polypeptide. In some cases, the first enzyme is a conjugate comprising... Figure 8 , Figure 9 , Figures 10A to 10Z and Figures 10AA to 10VV A tyrosinase polypeptide having at least 75% amino acid sequence identity for any amino acid sequence described by either of the two enzymes. In some cases, the thiol moiety present in the first polypeptide is present in Cys (e.g., solvent-accessible Cys; e.g., N-terminal Cys), and the phenol moiety present in the first polypeptide is present in a Tyr residue. In some cases, the Tyr residue is present in an amino acid segment containing EEEY (SEQ ID NO: 953), EEEEY (SEQ ID NO: 955), DDDDY (SEQ ID NO: 965), or DDDDY (SEQ ID NO: 965). In some cases, the second enzyme is a tyrosinase polypeptide containing amino acids with the same amino acid sequence as described by either of the two enzymes. Figures 10A to 10Z and Figures 10AA to 10VV A tyrosinase polypeptide having at least 75% amino acid sequence identity with any one of the amino acid sequences described by either of the two enzymes. In some cases, the second enzyme is a polypeptide containing amino acid sequences that are identical to those described by either of the two enzymes. Figure 10M The amino acid sequence depicted herein has amino acid sequence identity of at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%, wherein the tyrosinase polypeptide contains a D55 substitution (e.g., a D55K substitution).
[0483] This method can be used to sequentially link any number of polypeptides. For example, in some cases, the method further includes c) contacting the fixed conjugate with a third enzyme, wherein the third enzyme is capable of oxidizing the phenolic or catechol moiety present in the second polypeptide to produce a fixed conjugate comprising a reactive moiety; and d) contacting the fixed conjugate comprising the reactive moiety with a third polypeptide, wherein the third polypeptide comprises: i) a thiol moiety; and ii) a phenolic or catechol moiety, wherein the third polypeptide comprises two or more negatively charged amino acids within ten amino acids of the phenolic or catechol moiety, and wherein the reactive moiety present in the fixed conjugate forms a covalent bond with the thiol moiety present in the second polypeptide, thereby producing a fixed conjugate comprising the third polypeptide covalently linked to the second polypeptide. In some cases, the third enzyme is a conjugate containing a reactive moiety. Figure 8 or Figure 9 The amino acid sequence described herein is a tyrosinase polypeptide with at least 75% amino acid sequence identity.
[0484] Alternating use of: a) a tyrosinase that preferentially modifies Tyr residues present in a negatively charged environment (e.g., when the polypeptide contains two or more negatively charged amino acids within ten amino acids of the Tyr residues); and b) a tyrosinase that preferentially modifies Tyr residues present in a neutral or positively charged environment (e.g., when the polypeptide contains two or more neutral or positively charged amino acids within ten amino acids of the Tyr residues), can sequentially add polypeptide substrates to a fixed conjugate containing one, two, three, or more polypeptides. The above method can be modified, for example, such that the first polypeptide comprises: i) a thiol moiety; and ii) a phenolic or catechol moiety, wherein the first polypeptide comprises two or more neutral or positively charged amino acids within ten amino acids of the phenolic or catechol moiety; in such cases, the second polypeptide will comprise: i) a thiol moiety; and ii) a phenolic or catechol moiety, wherein the first polypeptide comprises two or more negatively charged amino acids within ten amino acids of the phenolic or catechol moiety.
[0485] In some cases, the tyrosinase is inactivated or removed between any two steps of the method and before the addition of an additional tyrosinase. For example, in some cases, the second enzyme is inactivated or removed between steps (b) and (c) of the method described above. In some cases, the thiol moiety present in the second polypeptide is present in Cys, and the phenolic moiety present in the second polypeptide is present in the Tyr residue. In some cases, the Tyr residue is present in an amino acid segment containing RRRY (SEQ ID NO: 949), RRRRY (SEQ ID NO: 951), KKKY (SEQ ID NO: 966), or KKKKY (SEQ ID NO: 967).
[0486] like Figure 39A The diagram schematically depicts abTYR used to link biotin-phenol and a first polypeptide (“protein A”) containing a thiol group and the sequence EEEEY (SEQ ID NO: 955) to produce a biotin-first polypeptide conjugate. The biotin-first polypeptide conjugate can be contacted with streptavidin beads to immobilize the biotin-first polypeptide conjugate. A second polypeptide (“protein B”) containing a thiol group and the sequence RRRRY (SEQ ID NO: 951) can be connected via bmTYR (D55K) (e.g., Figure 10M The bmTYR (D55K) described in the diagram conjugates to the first polypeptide of a fixed biotin-first polypeptide conjugate to produce a fixed first polypeptide-second polypeptide conjugate. In some cases, two different polypeptides (e.g., "protein A" and "protein B") are added alternately to the polyconjugate, such as... Figure 39A , Figure 39B and Figure 39C As shown in the diagram. Alternatively, multiple copies of a single polypeptide can be linked, as shown in the diagram. Figure 39D and Figure 39E As shown in the diagram. As another possibility, in some cases, each linked polypeptide is different from the others, such as... Figure 39F and Figure 39G As shown in the image (e.g., “protein A”; “protein B”; and “protein C”).
[0487] Composition
[0488] This disclosure also provides compositions comprising pharmaceutical compositions comprising: a thiol-containing target molecule of formula (III), and a biomolecule of formula (I) comprising a phenolic or catechol moiety.
[0489] (III) (I)
[0490] Wherein Y1 is a biomolecule, which optionally includes one or more portions selected from: active small molecules, affinity tags, fluorophores, and metal chelators; L is an optional linker; X1 is selected from hydrogen and hydroxyl groups; and Y2 is a second biomolecule.
[0491] In some embodiments, a composition of formula (III) comprising a target molecule of thiol and a pharmaceutically acceptable excipient is provided.
[0492] In some embodiments, a composition of formula (I) comprising a biomolecule of phenolic or catechol moiety and a pharmaceutically acceptable excipient is provided.
[0493] In some embodiments of the subject composition, Y 2 It is a CRISPR-Cas effector polypeptide, for example, as described herein.
[0494] In certain embodiments of the subject composition, formula (I) is described by any of the formulas (IA), (IAa), (IB), (IC), (ID), (IDa), and (IDb) disclosed herein.
[0495] The subject composition typically comprises: a thiol-containing subject target molecule of formula (III); a biomolecule of formula (I) containing a phenolic or catechol moiety; and at least one additional compound. Suitable additional compounds include, but are not limited to: salts, such as magnesium salts, sodium salts, etc., such as NaCl, MgCl2, KCl, MgSO4, etc.; buffers, such as Tris buffer, N-(2-hydroxyethyl)piperazine-N'-(2-ethanesulfonic acid) (HEPES), 2-(N-morpholino)ethanesulfonic acid (MES), sodium 2-(N-morpholino)ethanesulfonate (MES), 3-(N-morpholino)propanesulfonic acid (MOPS), N-tris[hydroxymethyl]methyl-3-aminopropanesulfonic acid (TAPS), etc.; solubilizers; detergents, such as nonionic detergents, such as Tween-20, etc.; protease inhibitors; and so on.
[0496] In some embodiments, the subject composition comprises: a thiol-containing subject target molecule of formula (III); a biomolecule of formula (I) comprising a phenolic or catechol moiety; and a pharmaceutically acceptable excipient. A variety of pharmaceutically acceptable excipients are known in the art and need not be discussed in detail herein. Pharmaceutically acceptable excipients have been well described in various publications, including, for example, A. Gennaro (2000) “Remington: The Science and Practice of Pharmacy,” 20th edition, Lippincott, Williams, & Wilkins; Pharmaceutical Dosage Forms and Drug Delivery Systems (1999) HC Ansel et al., 7th edition, Lippincott, Williams, & Wilkins; and Handbook of Pharmaceutical Excipients (2000) AH Kibbe et al., 3rd edition, Amer. PharmaceuticalAssoc.
[0497] Pharmaceutically acceptable excipients, such as mediators, adjuvants, carriers, or diluents, are readily available to the public. Furthermore, pharmaceutically acceptable auxiliary substances, such as pH adjusters and buffers, tension modifiers, stabilizers, and humectants, are readily available to the public.
[0498] Reagent test kit
[0499] The compounds and compositions described herein may be packaged as kits, which may optionally include instructions for use of the compounds or compositions in various exemplary applications. Non-limiting examples include kits containing compounds or compositions in, for example, powder or lyophilized form, along with instructions for use in the subject methods, including reconstitution, dosage information, and storage information. Kits may optionally contain containers of the compounds or compositions in a spare liquid form, or require further mixing with a solution for administration.
[0500] This disclosure includes a kit comprising a first container containing a composition comprising a thiol-containing subject target molecule of formula (III) and a phenolic moiety of formula (I); and a second container containing an enzyme capable of oxidizing a phenolic or catechol moiety. In some cases, the enzyme is a tyrosinase.
[0501] In some embodiments, the subject kit includes a first container containing a thiol-containing subject target molecule of formula (III); a second container containing a biomolecule of formula (I) containing a phenolic or catechol moiety; and a third container containing an enzyme capable of oxidizing the phenolic or catechol moiety. In some cases, the enzyme is a tyrosinase.
[0502] The kit may include optional components that aid in the main methodology, such as vials for reconstituted powder form. The kit is available in a sealed container suitable for single or multiple punctures with a hypodermic needle (e.g., a coiled septum seal) while maintaining sterile integrity. Kit components can be assembled into cartons, blister packs, bottles, tubes, etc.
[0503] In addition to the components described above, the subject kit may also include instructions for implementing the subject method. These instructions may exist in various forms within the subject kit, with one or more of these forms present. One form of these instructions may be information printed on a suitable medium or substrate, such as one or more sheets of paper on which information is printed, in the kit packaging, in the package insert, etc. Another form would be computer-readable media, such as CDs, DVDs, Blu-ray discs, computer-readable storage devices (e.g., flash memory), etc. Yet another form may be a website address, such as a link to a website for downloading a suitable smartphone app for detecting functional dyes, which can be accessed via the Internet to remove information from the site. Any convenient means may be present in the kit.
[0504] practicality
[0505] The subject compounds, compositions, kits, and subject modification methods can be used in a variety of applications, including research and diagnostic applications.
[0506] Research applications of interest include any application that focuses on the selective manipulation of target molecules, biomolecules, cells, particles, and surfaces, including in vitro manipulation, labeling, and tracking of biomolecules (e.g., proteins).
[0507] The subject methods and compositions can also be used for therapeutic applications, such as antibody-drug conjugates (ADCs) of interest, which can be used (e.g., in novel immunotherapies), protein delivery for gene therapy, and applications in vaccine development.
[0508] Methods for screening tyrosinase variants
[0509] This disclosure provides a method for identifying tyrosinase variants that are preferential to a specific substrate. The method provides for the identification of tyrosinase variants that are preferential to phenols or catechols present in a specific sequence, in a negatively charged or positively charged environment. The method generally involves: a) contacting a peptide with a test tyrosinase and thiol-modified biotin (biotin-thiol), wherein the peptide has a length of about 4 amino acids to about 25 amino acids (e.g., about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 to 20 or 20 to 25 amino acids), and wherein the peptide has a C-terminal Tyr residue; b) contacting the biotin-peptide conjugate with streptavidin-conjugated beads (e.g., streptavidin conjugated with magnetic beads) to generate a streptavidin-biotin-peptide complex; and c) determining the amino acid sequence of the peptide in the streptavidin-biotin-peptide complex. In some cases, the method further includes the step of washing the streptavidin-biotin-peptide complex to remove unbound peptides (peptides not conjugated with biotin). In some cases, the method further includes the step of releasing the peptide from the streptavidin-biotin-peptide complex before determining the amino acid sequence of the peptide. The peptide can be released (eluted) from the streptavidin-biotin-peptide complex by incubating the streptavidin-biotin-peptide complex in a mixture containing excess free biotin, acetonitrile, and formic acid (e.g., 80% acetonitrile, 5% formic acid, and 2 mM biotin). The amino acid sequence of the peptide present in the streptavidin-biotin-peptide complex (e.g., the eluted peptide) can be determined using any of a variety of well-known methods, including, for example, mass spectrometry (MS) (e.g., tandem MS). Peptide libraries can be used to determine whether the tyrosinase being tested has a preference for a specific amino acid sequence, a negatively charged environment, or a positively charged environment.
[0510] Examples of non-limiting aspects of this disclosure
[0511] Aspect A
[0512] The various aspects of the subject matter of the present invention described above, including the various embodiments, may be advantageous individually or in combination with one or more other aspects or embodiments. Without limiting the foregoing description, certain non-limiting aspects of this disclosure, numbered 1-39, are provided below. As will be apparent to those skilled in the art upon reading this disclosure, each of the individually numbered aspects may be used or combined with any of the preceding or subsequent individually numbered aspects. This is intended to support all such combinations of aspects and is not limited to the combinations of aspects explicitly provided below:
[0513] Aspect 1. A method for chemically selectively modifying a target molecule, the method comprising: contacting a target molecule comprising a thiol moiety with a biomolecule comprising a reactive moiety; the biomolecule comprising the reactive moiety being generated by reacting a biomolecule comprising a phenolic moiety or a catechol moiety with an enzyme capable of oxidizing the phenolic or catechol moiety; and wherein the contact is carried out under conditions sufficient to conjugate the target molecule with the biomolecule, thereby generating a modified target molecule.
[0514] Aspect 2. The method as described in aspect 1, wherein the target molecule is a polypeptide.
[0515] Aspect 3. The method as described in aspect 1 or 2, wherein the enzyme is a tyrosinase.
[0516] Aspect 4. The method of any one of Aspects 1 to 3, wherein the enzyme is bound to a solid carrier.
[0517] Aspect 5. The method of any one of Aspects 1 to 4, wherein the phenolic moiety is present in a tyrosine residue.
[0518] Aspect 6. The method of any one of Aspects 1 to 5, wherein the thiol moiety is present in a cysteine residue.
[0519] Aspect 7. The method of aspect 6, wherein the cysteine residue is a natural cysteine residue.
[0520] Aspect 8. The method of any one of Aspects 1 to 7, wherein the biomolecule comprises one or more portions selected from: fluorophores, active small molecules, affinity tags, and metal chelators.
[0521] Aspect 9. The method of any one of Aspects 1 to 8, wherein the reactive portion is an ortho-quinone or semi-quinone group, or a combination thereof.
[0522] Aspect 10. The method of any one of Aspects 1 to 9, wherein the biomolecule is a polypeptide.
[0523] Aspect 11. The method of aspect 10, wherein the biomolecule is a polypeptide selected from fluorescent proteins, antibodies, enzymes, receptor ligands, and receptors.
[0524] Aspect 12. The method of any one of Aspects 1 to 11, wherein the biomolecule comprising a phenolic or catechol moiety has formula (I), and the biomolecule comprising a reactive moiety has formula (II) or (IIA), or a combination thereof:
[0525]
[0526] in:
[0527] Y 1 It is a biomolecule, which optionally comprises one or more portions selected from: active small molecules, affinity tags, fluorophores, and metal chelators;
[0528] X 1 Selected from hydrogen and hydroxyl; and
[0529] L is an optional connector.
[0530] Aspect 13. The method of any one of Aspects 1 to 12, wherein the target molecule comprising the thiol moiety has formula (III), and wherein the modified target molecule has formula (IV) or (IVA), or a combination thereof:
[0531]
[0532] in:
[0533] Y 1 It is a biomolecule, which optionally comprises one or more portions selected from: active small molecules, affinity tags, fluorophores, and metal chelators;
[0534] Y 2 It is the second biomolecule;
[0535] L is an optional connector; and
[0536] n is an integer from 1 to 3.
[0537] Aspect 14. The method as described in aspect 13, wherein the modified target molecule of formula (IV) has any one of formulas (IV1)-(IV3):
[0538]
[0539] The target molecule modified according to formula (IVA) has any one of formulas (IVA1)-(IVA3):
[0540]
[0541] Aspect 15. The method as described in aspect 13, wherein the modified target molecule of formula (IV) has any one of formulas (IV5)-(IV6):
[0542]
[0543] The target molecule modified according to formula (IVA) has any one of formulas (IVA4)-(IVA5):
[0544]
[0545] Aspect 16. The method of any one of Aspects 1 to 15, wherein the biomolecule comprising a phenolic or catechol moiety is described by formula (IA):
[0546]
[0547] in:
[0548] Y 1 It is a biomolecule, which optionally includes one or more groups selected from the following: active small molecules, affinity tags, fluorophores and metal chelators;
[0549] Each R 1 Independently selected from hydrogen, acyl, substituted acyl, alkyl, and substituted alkyl;
[0550] X 1 Selected from hydrogen and hydroxyl; and
[0551] L 1 It is a linker selected from straight-chain or branched alkyl groups, straight-chain or branched substituted alkyl groups, polyethylene glycol (PEG), substituted PEG, and one or more peptides.
[0552] Aspect 17. The method of aspect 16, wherein the fluorophore is a rhodamine dye or a succinate dye.
[0553] Aspect 18. The method of any one of Aspects 1 to 17, wherein the modified target molecule is described by formula (IVB) or (IVC) or a combination thereof:
[0554]
[0555] in:
[0556] Y 1 It is a biomolecule, which optionally includes one or more groups selected from the following: active small molecules, affinity tags, fluorophores and metal chelators;
[0557] Each R 1 Independently selected from hydrogen, acyl, substituted acyl, alkyl, and substituted alkyl;
[0558] Y 2 It is the second biomolecule;
[0559] L 1 It is a linker selected from straight-chain or branched alkyl groups, straight-chain or branched substituted alkyl groups, polyethylene glycol (PEG), substituted PEG, and one or more peptides; and
[0560] n is an integer from 1 to 3.
[0561] Aspect 19. The method as described in aspect 18, wherein the target molecule modified by formula (IVB) has any one of formulas (IVB1)-(IVZB3):
[0562]
[0563] The target molecule modified according to formula (IVC) has any one of formulas (IVC1)-(IVC3):
[0564]
[0565] Aspect 20. The method as described in aspect 18, wherein the modified target molecule of formula (IVB) has any one of formulas (IVB5)-(IVB6):
[0566]
[0567] The target molecule modified according to formula (IVC) has any one of formulas (IVC4)-(IVC5):
[0568]
[0569] Aspect 21. The method of any one of aspects 1 to 20, wherein the method is carried out at a pH of 4 to 9.
[0570] Aspect 22. The method as described in aspect 21, wherein the method is carried out at a neutral pH.
[0571] Aspect 23. The method of any one of Aspects 1 to 22, wherein the target molecule comprising a thiol group is a CRISPR-Cas effector polypeptide.
[0572] Aspect 24. A composition comprising:
[0573] The target molecule containing thiols in formula (III):
[0574] (III); and
[0575] Biomolecules of formula (I) containing a phenolic or catechol moiety:
[0576]
[0577] in:
[0578] Y 1 It is a biomolecule, which optionally comprises one or more portions selected from: active small molecules, affinity tags, fluorophores, and metal chelators;
[0579] X 1 Selected from hydrogen and hydroxyl groups;
[0580] L is an optional connector; and
[0581] Y 2 It is the second biomolecule.
[0582] Aspect 25. The composition as described in aspect 24, wherein Y 2 It is a CRISPR-Cas effector polypeptide.
[0583] Aspect 26. The composition as described in aspect 24 or 25, wherein formula (I) is described by formula (IA):
[0584]
[0585] in:
[0586] Y 1 It is a biomolecule, which optionally includes one or more groups selected from the following: active small molecules, affinity tags, fluorophores and metal chelators;
[0587] Each R 1 Independently selected from hydrogen, acyl, substituted acyl, alkyl, and substituted alkyl;
[0588] X 1 Selected from hydrogen and hydroxyl; and
[0589] L 1 It is a linker selected from straight-chain or branched alkyl groups, straight-chain or branched substituted alkyl groups, polyethylene glycol (PEG), substituted PEG, and one or more peptides.
[0590] Aspect 27. A reagent kit comprising:
[0591] A first container, the first container comprising the composition as described in any one of aspects 24 to 26; and
[0592] A second container contains an enzyme capable of oxidizing the phenol or catechol moiety.
[0593] Aspect 28. The kit of claim 27, wherein the enzyme is a tyrosinase.
[0594] Section 29. A compound of formula (IV) or (IVA):
[0595]
[0596] in:
[0597] Y 1It is a biomolecule, which optionally comprises one or more portions selected from: active small molecules, affinity tags, fluorophores, and metal chelators;
[0598] L is an optional connector;
[0599] Y 2 It is the second biomolecule; and
[0600] n is an integer from 1 to 3.
[0601] Aspect 30. The compound as described in aspect 29, wherein the target molecule modified by formula (IV) has any one of formulas (IV1)-(IV5):
[0602]
[0603] Aspect 31. The compound as described in aspect 29, wherein the modified target molecule of formula (IVA) has any one of formulas (IVA1)-(IVA5):
[0604]
[0605] Aspect 32. The compound of any one of Aspects 29 to 31, wherein L is a pyrolytic linker.
[0606] Aspect 33. The compound as described in any one of Aspects 29 to 32, wherein Y 1 It is a polypeptide.
[0607] Aspect 34. The compound as described in aspect 33, wherein Y 1 Selected from fluorescent proteins, antibodies, and enzymes.
[0608] Aspect 35. The compound of any one of Aspects 29 to 34, wherein the compound is described by formula (IVB) or (IVC):
[0609]
[0610] in:
[0611] Y 1 It is a biomolecule, which optionally includes one or more groups selected from the following: active small molecules, affinity tags, fluorophores and metal chelators;
[0612] Each R 1 Independently selected from hydrogen, acyl, substituted acyl, alkyl, and substituted alkyl;
[0613] Y 2 It is the second biomolecule;
[0614] L1 It is a linker selected from straight-chain or branched alkyl groups, straight-chain or branched substituted alkyl groups, polyethylene glycol (PEG), substituted PEG, and one or more peptides; and
[0615] n is an integer from 1 to 3.
[0616] Aspect 36. The compound as described in aspect 35, wherein the target molecule modified by formula (IVB) has any one of formulas (IVB1)-(IVZB5):
[0617]
[0618] Aspect 37. The compound as described in aspect 35, wherein the target molecule modified by formula (IV) has any one of formulas (IVC1)-(IVC5):
[0619]
[0620] Aspect 38. The compound of any one of Aspects 29 to 37, wherein the compound is described by any one of the formulas (IVD)-(IVG):
[0621]
[0622] in:
[0623] R 2 Selected from alkyl and substituted alkyl groups;
[0624] R 3 Selected from hydrogen, alkyl, substituted alkyl, peptide and polypeptide; and
[0625] n is an integer from 1 to 3.
[0626] Aspect 39. The compound as described in any one of Aspects 29 to 38, wherein Y 2 It is a CRISPR-Cas effector polypeptide.
[0627] Aspect B
[0628] The various aspects of the subject matter of the present invention described above, including the various embodiments, may be advantageous individually or in combination with one or more other aspects or embodiments. Without limiting the foregoing description, certain non-limiting aspects of this disclosure, numbered 1-71, are provided below. As will be apparent to those skilled in the art upon reading this disclosure, each of the individually numbered aspects may be used or combined with any of the preceding or subsequent individually numbered aspects. This is intended to support all such combinations of aspects, and is not limited to the combinations of aspects explicitly provided below:
[0629] Aspect 1. A method for chemically selectively modifying a target molecule, the method comprising: contacting a target molecule comprising a thiol moiety with a biomolecule comprising a reactive moiety; the biomolecule comprising the reactive moiety being generated by reacting a biomolecule comprising a phenolic moiety or a catechol moiety with an enzyme capable of oxidizing the phenolic or catechol moiety; and wherein the contact is carried out under conditions sufficient to conjugate the target molecule with the biomolecule, thereby generating a modified target molecule.
[0630] Aspect 2. The method as described in aspect 1, wherein the target molecule is a polypeptide or a polynucleotide.
[0631] Aspect 3. The method as described in aspect 1 or aspect 2, wherein the enzyme is a tyrosinase polypeptide.
[0632] Aspect 4. The method of any one of Aspects 1 to 3, wherein the tyrosinase polypeptide is Agaricus bisporus tyrosinase (abTYR) polypeptide.
[0633] Aspect 5. The method of any one of Aspects 1 to 3, wherein the tyrosinase polypeptide comprises with Figure 8 or Figure 9 The amino acid sequence of abTYR described in the figure has at least 75% amino acid sequence identity.
[0634] Aspect 6. The method as described in aspect 4 or aspect 5, wherein the biomolecule comprising the phenolic portion or the catechol portion is neutral or positively charged within 50 angstroms (Å) of the phenolic or catechol portion.
[0635] Aspect 7. The method of any one of Aspects 1 to 3, wherein the tyrosinase polypeptide is a Bacillus megaterium tyrosinase (bmTYR) polypeptide.
[0636] Aspect 8. The method of any one of Aspects 1 to 3, wherein the tyrosinase polypeptide comprises with Figures 10A to 10Z and Figures 10AA to 10VV An amino acid sequence that describes any one of the amino acid sequences and has at least 75% amino acid sequence identity.
[0637] Aspect 9. The method as described in aspect 7 or aspect 8, wherein the biomolecule comprising the phenolic or catechol moiety is negatively charged within 50 Å of the phenolic or catechol moiety.
[0638] Aspect 10. The method of any one of Aspects 1 to 9, wherein the target molecule is a polynucleotide.
[0639] Aspect 11. The method of aspect 10, wherein the target molecule is a DNA molecule.
[0640] Aspect 12. The method of aspect 10, wherein the target molecule is an RNA molecule.
[0641] Aspect 13. The method of any one of Aspects 10 to 12, wherein the biomolecule is a polypeptide.
[0642] Aspect 14. The method of any one of Aspects 1 to 13, wherein the enzyme is bound to a solid carrier.
[0643] Aspect 15. The method of any one of Aspects 1 to 14, wherein the phenolic moiety is present in a tyrosine residue.
[0644] Aspect 16. The method of any one of Aspects 1 to 15, wherein the thiol moiety is present in a cysteine residue.
[0645] Aspect 17. The method of aspect 16, wherein the cysteine residue is a natural cysteine residue.
[0646] Aspect 18. The method of any one of Aspects 1 to 17, wherein the biomolecule comprises one or more portions selected from: fluorophores, active small molecules, affinity tags, and metal chelators.
[0647] Aspect 19. The method of any one of Aspects 1 to 18, wherein the reactive portion is an ortho-quinone or semi-quinone group, or a combination thereof.
[0648] Aspect 20. The method of any one of Aspects 1 to 19, wherein the biomolecule is a polypeptide.
[0649] Aspect 21. The method of aspect 20, wherein the biomolecule is a polypeptide selected from fluorescent proteins, antibodies, enzymes, receptor ligands, and receptors.
[0650] Aspect 22. The method of any one of Aspects 1 to 21, wherein the biomolecule comprising a phenolic or catechol moiety has formula (I), and the biomolecule comprising a reactive moiety has formula (II) or (IIA), or a combination thereof:
[0651]
[0652] in:
[0653] Y 1 It is a biomolecule, which optionally comprises one or more portions selected from: active small molecules, affinity tags, fluorophores, and metal chelators;
[0654] X 1 Selected from hydrogen and hydroxyl; and
[0655] L is an optional connector.
[0656] Aspect 23. The method of any one of Aspects 1 to 22, wherein the target molecule comprising the thiol moiety has formula (III), and wherein the modified target molecule has formula (IV) or (IVA), or a combination thereof:
[0657]
[0658] in:
[0659] Y 1 It is a biomolecule, which optionally comprises one or more portions selected from: active small molecules, affinity tags, fluorophores, and metal chelators;
[0660] Y 2 It is the second biomolecule;
[0661] L is an optional connector; and
[0662] n is an integer from 1 to 3.
[0663] Aspect 24. The method as described in aspect 23, wherein the target molecule modified by formula (IV) has any one of formulas (IV1)-(IV3):
[0664]
[0665] The target molecule modified according to formula (IVA) has any one of formulas (IVA1)-(IVA3):
[0666]
[0667] Aspect 25. The method as described in aspect 23, wherein the modified target molecule of formula (IV) has any one of formulas (IV5)-(IV6):
[0668]
[0669] The target molecule modified according to formula (IVA) has any one of formulas (IVA4)-(IVA5):
[0670]
[0671] Aspect 26. The method of any one of Aspects 1 to 25, wherein the biomolecule comprising a phenolic or catechol moiety is described by formula (IA):
[0672]
[0673] in:
[0674] Y 1 It is a biomolecule, which optionally includes one or more groups selected from the following: active small molecules, affinity tags, fluorophores and metal chelators;
[0675] Each R 1 Independently selected from hydrogen, acyl, substituted acyl, alkyl, and substituted alkyl;
[0676] X 1 Selected from hydrogen and hydroxyl; and
[0677] L 1 It is a linker selected from straight-chain or branched alkyl groups, straight-chain or branched substituted alkyl groups, polyethylene glycol (PEG), substituted PEG, and one or more peptides.
[0678] Aspect 27. The method of aspect 26, wherein the fluorophore is a rhodamine dye or a succinate dye.
[0679] Aspect 28. The method of any one of Aspects 1 to 27, wherein the modified target molecule is described by formula (IVB) or (IVC) or a combination thereof:
[0680]
[0681] in:
[0682] Y 1 It is a biomolecule, which optionally includes one or more groups selected from the following: active small molecules, affinity tags, fluorophores and metal chelators;
[0683] Each R 1 Independently selected from hydrogen, acyl, substituted acyl, alkyl, and substituted alkyl;
[0684] Y 2 It is the second biomolecule;
[0685] L 1 It is a linker selected from straight-chain or branched alkyl groups, straight-chain or branched substituted alkyl groups, polyethylene glycol (PEG), substituted PEG, and one or more peptides; and
[0686] n is an integer from 1 to 3.
[0687] Aspect 29. The method as described in aspect 28, wherein the target molecule modified by formula (IVB) has any one of formulas (IVB1)-(IVZB3):
[0688]
[0689] The target molecule modified according to formula (IVC) has any one of formulas (IVC1)-(IVC3):
[0690]
[0691] Aspect 30. The method as described in aspect 28, wherein the modified target molecule of formula (IVB) has any one of formulas (IVB5)-(IVB6):
[0692]
[0693] The target molecule modified according to formula (IVC) has any one of formulas (IVC4)-(IVC5):
[0694]
[0695] Aspect 31. The method of any one of aspects 1 to 30, wherein the method is carried out at a pH of 4 to 9.
[0696] Aspect 32. The method as described in aspect 31, wherein the method is carried out at a neutral pH.
[0697] Aspect 33. The method of any one of Aspects 1 to 32, wherein the target molecule comprising a thiol group is a CRISPR-Cas effector polypeptide.
[0698] Aspect 34. A composition comprising:
[0699] The target molecule containing thiols in formula (III):
[0700] (III); and
[0701] Biomolecules of formula (I) containing a phenolic or catechol moiety:
[0702]
[0703] in:
[0704] Y 1 It is a biomolecule, which optionally comprises one or more portions selected from: active small molecules, affinity tags, fluorophores, and metal chelators;
[0705] X 1 Selected from hydrogen and hydroxyl groups;
[0706] L is an optional connector; and
[0707] Y 2 It is the second biomolecule.
[0708] Aspect 35. The composition of aspect 34, wherein the biomolecule comprising the phenolic or catechol moiety is neutral or positively charged within 50 Å of the phenolic or catechol moiety.
[0709] Aspect 36. The composition of aspect 34, wherein the biomolecule comprising the phenolic or catechol moiety is negatively charged within 50 Å of the phenolic or catechol moiety.
[0710] Aspect 37. The composition as described in any one of Aspects 34 to 36, wherein Y 1 It is a polypeptide and in which Y 2 It is a polypeptide.
[0711] Aspect 38. The composition as described in any one of Aspects 34 to 36, wherein Y 1 It is a polynucleotide and in which Y 2 It is a polypeptide.
[0712] Aspect 39. The composition as described in any one of Aspects 34 to 38, wherein Y 2 It is a CRISPR-Cas effector polypeptide.
[0713] Aspect 40. The composition as described in any one of aspects 34 to 39, wherein formula (I) is described by formula (IA):
[0714]
[0715] in:
[0716] Y 1 It is a biomolecule, which optionally includes one or more groups selected from the following: active small molecules, affinity tags, fluorophores and metal chelators;
[0717] Each R 1 Independently selected from hydrogen, acyl, substituted acyl, alkyl, and substituted alkyl;
[0718] X 1 Selected from hydrogen and hydroxyl; and
[0719] L 1 It is a linker selected from straight-chain or branched alkyl groups, straight-chain or branched substituted alkyl groups, polyethylene glycol (PEG), substituted PEG, and one or more peptides.
[0720] Aspect 41. A reagent kit comprising:
[0721] A first container, the first container comprising the composition as described in any one of aspects 34 to 40; and
[0722] A second container contains an enzyme capable of oxidizing the phenol or catechol moiety.
[0723] Aspect 42. The kit as described in aspect 41, wherein the enzyme is a tyrosinase polypeptide.
[0724] Aspect 43. The kit as described in aspect 42, wherein the tyrosinase is Agaricus bisporus tyrosinase (abTYR).
[0725] Aspect 44. The kit as described in aspect 42, wherein the tyrosinase polypeptide comprises with Figure 8 or Figure 9 The amino acid sequence of abTYR described in the figure has at least 75% amino acid sequence identity.
[0726] Aspect 45. The kit as described in aspect 42, wherein the tyrosinase is Bacillus megaterium tyrosinase (bmTYR).
[0727] Aspect 46. The kit as described in aspect 42, wherein the tyrosinase polypeptide comprises with Figures 10A to 10Z and Figures 10AA to 10VV An amino acid sequence that describes any one of the amino acid sequences and has at least 75% amino acid sequence identity.
[0728] Section 47. A compound of formula (IV) or (IVA):
[0729]
[0730] in:
[0731] Y 1 It is a biomolecule, which optionally comprises one or more portions selected from: active small molecules, affinity tags, fluorophores, and metal chelators;
[0732] L is an optional connector;
[0733] Y 2 It is the second biomolecule; and
[0734] n is an integer from 1 to 3.
[0735] Aspect 48. The compound as described in aspect 47, wherein the target molecule modified by formula (IV) has any one of formulas (IV1)-(IV5):
[0736]
[0737] Aspect 49. The compound as described in aspect 47, wherein the modified target molecule of formula (IVA) has any one of formulas (IVA1)-(IVA5):
[0738]
[0739] Aspect 50. The compound of any one of Aspects 47 to 49, wherein L is a pyrolytic linker.
[0740] Aspect 51. The compound as described in any one of Aspects 47 to 50, wherein Y 1 It is a polypeptide.
[0741] Aspect 52. The compound as described in aspect 51, wherein Y 1 Selected from fluorescent proteins, antibodies, and enzymes.
[0742] Aspect 53. The compound of any one of Aspects 47 to 52, wherein the compound is described by formula (IVB) or (IVC):
[0743]
[0744] in:
[0745] Y 1 It is a biomolecule, which optionally includes one or more groups selected from the following: active small molecules, affinity tags, fluorophores and metal chelators;
[0746] Each R 1 Independently selected from hydrogen, acyl, substituted acyl, alkyl, and substituted alkyl;
[0747] Y 2 It is the second biomolecule;
[0748] L 1 It is a linker selected from straight-chain or branched alkyl groups, straight-chain or branched substituted alkyl groups, polyethylene glycol (PEG), substituted PEG, and one or more peptides; and
[0749] n is an integer from 1 to 3.
[0750] Aspect 54. The compound as described in aspect 53, wherein the target molecule modified by formula (IVB) has any one of formulas (IVB1)-(IVZB5):
[0751]
[0752] Aspect 55. The compound as described in aspect 53, wherein the target molecule modified by formula (IV) has any one of formulas (IVC1)-(IVC5):
[0753]
[0754] Aspect 56. The compound of any one of Aspects 47 to 55, wherein the compound is described by any one of the formulas (IVD)-(IVG):
[0755]
[0756] in:
[0757] R 2 Selected from alkyl and substituted alkyl groups;
[0758] R 3 Selected from hydrogen, alkyl, substituted alkyl, peptide and polypeptide; and
[0759] n is an integer from 1 to 3.
[0760] Aspect 57. The compound as described in any one of Aspects 47 to 56, wherein Y 2 It is a CRISPR-Cas effector polypeptide.
[0761] Aspect 58. A method for chemically selectively coupling a first polypeptide and a second polypeptide to a coupled polypeptide, the method comprising:
[0762] a) Contact the first polypeptide with the conjugated polypeptide to generate a first polypeptide-conjugated polypeptide conjugate.
[0763] The first polypeptide contains a thiol moiety.
[0764] The coupled polypeptide includes an N-terminal reactive portion that forms a covalent bond with the thiol moiety present in the first polypeptide.
[0765] The conjugated polypeptide, which includes the N-terminal reactive portion, is produced by reacting a polypeptide comprising an N-terminal phenolic or catechol portion and a C-terminal phenolic or catechol portion with a first enzyme capable of oxidizing the N-terminal phenolic or catechol portion but not the C-terminal phenolic or catechol portion to produce the N-terminal reactive portion.
[0766] The coupling polypeptide comprises two or more positively charged or neutral amino acids within ten amino acids of the N-terminal phenolic or catechol moiety and comprises two or more negatively charged amino acids within ten amino acids of the C-terminal phenolic or catechol moiety; and
[0767] b) Contact the second polypeptide with the first polypeptide-coupled polypeptide conjugate.
[0768] The second polypeptide contains a thiol moiety.
[0769] The first polypeptide-coupled polypeptide conjugate includes a C-terminal reactive moiety that forms a covalent bond with the thiol moiety present in the second polypeptide.
[0770] The first polypeptide-coupled polypeptide conjugate, which includes the C-terminal reactive portion, is produced by reacting the first polypeptide-coupled polypeptide conjugate with a second enzyme capable of oxidizing the C-terminal phenolic or catechol portion to generate the C-terminal reactive portion; and
[0771] The contact produces a first polypeptide-coupled polypeptide-second polypeptide conjugate.
[0772] Aspect 59. The method as described in aspect 58, wherein:
[0773] a) The first enzyme contains and Figure 8 or Figure 9 The abTYR amino acid sequence described herein is a tyrosinase polypeptide with at least 75% amino acid sequence identity; and
[0774] b) The second enzyme contains... Figures 10A to 10Z and Figures 10AA to 10VV A tyrosinase polypeptide whose amino acid sequence has at least 75% amino acid sequence identity for any one of the amino acid sequences described by either of the above.
[0775] Aspect 60. A method for chemically selectively coupling a first polypeptide and a second polypeptide to a coupled polypeptide, the method comprising:
[0776] a) Contact the first polypeptide with the conjugated polypeptide to generate a first polypeptide-conjugated polypeptide conjugate.
[0777] The first polypeptide contains a thiol moiety.
[0778] The coupled polypeptide includes an N-terminal reactive portion that forms a covalent bond with the thiol moiety present in the first polypeptide.
[0779] The conjugated polypeptide, which includes the N-terminal reactive portion, is produced by reacting a polypeptide comprising an N-terminal phenolic or catechol portion and a C-terminal phenolic or catechol portion with a first enzyme capable of oxidizing the N-terminal phenolic or catechol portion but not the C-terminal phenolic or catechol portion to produce the N-terminal reactive portion.
[0780] The coupling polypeptide comprises two or more negatively charged amino acids within the ten amino acids of the N-terminal phenolic or catechol moiety and comprises two or more positively charged or neutral amino acids within the ten amino acids of the C-terminal phenolic or catechol moiety; and
[0781] b) Contact the second polypeptide with the first polypeptide-coupled polypeptide conjugate.
[0782] The second polypeptide contains a thiol moiety.
[0783] The first polypeptide-coupled polypeptide conjugate includes a C-terminal reactive moiety that forms a covalent bond with the thiol moiety present in the second polypeptide.
[0784] The first polypeptide-coupled polypeptide conjugate, which includes the C-terminal reactive portion, is produced by reacting the first polypeptide-coupled polypeptide conjugate with a second enzyme capable of oxidizing the C-terminal phenolic or catechol portion to generate the C-terminal reactive portion; and
[0785] The contact produces a first polypeptide-coupled polypeptide-second polypeptide conjugate.
[0786] Aspect 61. The method as described in aspect 60, wherein:
[0787] a) The first enzyme contains and Figures 10A to 10Z and Figures 10AA to 10VV A tyrosinase polypeptide having at least 75% amino acid sequence identity for any one of the amino acid sequences described by either of the above; and
[0788] b) The second enzyme contains... Figure 8 or Figure 9 The abTYR amino acid sequence depicted is a tyrosinase polypeptide with at least 75% amino acid sequence identity.
[0789] Aspect 62. A method for covalently linking a first polypeptide to a second polypeptide, the method comprising:
[0790] a) Contact the first polypeptide with the immobilized reactive portion.
[0791] The fixed reactive portion is generated by reacting a fixed phenolic or catechol portion with a first enzyme, wherein the first enzyme is capable of oxidizing the fixed phenolic or catechol portion, thereby generating the fixed reactive portion.
[0792] The first polypeptide comprises: i) a thiol moiety; and ii) a phenolic or catechol moiety, wherein the first polypeptide comprises two or more negatively charged amino acids within ten amino acids of the phenolic or catechol moiety.
[0793] The fixed reactive portion forms a covalent bond with the thiol portion present in the first polypeptide, thereby producing a fixed first polypeptide;
[0794] b) Contacting the fixed first polypeptide with a second enzyme, wherein the second enzyme is capable of oxidizing the phenolic or catechol moiety present in the first polypeptide to produce a fixed first polypeptide comprising a reactive moiety; and
[0795] c) Contact the fixed first polypeptide containing the reactive portion with the second polypeptide.
[0796] The second polypeptide comprises: i) a thiol moiety; and ii) a phenolic or catechol moiety, wherein the second polypeptide comprises two or more neutral or positively charged amino acids within ten amino acids of the phenolic or catechol moiety.
[0797] The reactive portion present in the fixed first polypeptide forms a covalent bond with the thiol portion present in the second polypeptide, thereby producing a fixed conjugate comprising the first polypeptide covalently linked to the second polypeptide.
[0798] Aspect 63. The method as described in aspect 62, wherein the first enzyme comprises with Figure 8 , Figure 9 , Figures 10A to 10Z and Figures 10AA to 10VV A tyrosinase polypeptide whose amino acid sequence has at least 75% amino acid sequence identity for any one of the amino acid sequences described by either of the above.
[0799] Aspect 64. The method of aspect 62 or aspect 63, wherein the thiol moiety present in the first polypeptide is present in Cys, and wherein the phenolic moiety present in the first polypeptide is present in Tyr residues.
[0800] Aspect 65. The method of aspect 64, wherein the Tyr residue is present in an amino acid segment comprising EEEY (SEQ ID NO: 953), EEEY (SEQ ID NO: 955), DDDDY (SEQ ID NO: 965) or DDDDY (SEQ ID NO: 965).
[0801] Aspect 66. The method of any one of aspects 62 to 65, wherein the second enzyme comprises with Figures 10A to 10Z and Figures 10AA to 10VV A tyrosinase polypeptide whose amino acid sequence has at least 75% amino acid sequence identity for any one of the amino acid sequences described by either of the above.
[0802] Aspect 67. The method of any one of aspects 62 to 66, further comprising:
[0803] c) Contacting the fixed conjugate with a third enzyme, wherein the third enzyme is capable of oxidizing the phenolic or catechol moiety present in the second polypeptide to produce a fixed conjugate comprising a reactive moiety; and
[0804] c) Contact the fixed conjugate containing the reactive portion with the third polypeptide.
[0805] The third polypeptide comprises: i) a thiol moiety; and ii) a phenolic or catechol moiety, wherein the third polypeptide contains two or more negatively charged amino acids within ten amino acids of the phenolic or catechol moiety.
[0806] The reactive portion present in the fixed conjugate forms a covalent bond with the thiol portion present in the second polypeptide, thereby producing a fixed conjugate comprising the third polypeptide covalently linked to the second polypeptide.
[0807] Aspect 68. The method as described in aspect 67, wherein the third enzyme comprises with Figure 8 or Figure 9 The amino acid sequence described herein is a tyrosinase polypeptide with at least 75% amino acid sequence identity.
[0808] Aspect 69. The method as described in aspects 67 or 68, wherein between step (b) and step (c), the second enzyme is inactivated or removed.
[0809] Aspect 70. The method of any one of Aspects 67 to 69, wherein the thiol moiety present in the second polypeptide is present in Cys, and wherein the phenolic moiety present in the second polypeptide is present in Tyr residues.
[0810] Aspect 71. The method of aspect 70, wherein the Tyr residue is present in an amino acid segment comprising RRRY (SEQ ID NO: 949), RRRRY (SEQ ID NO: 951), KKKY (SEQ ID NO: 966), or KKKKY (SEQ ID NO: 967).
[0811] Example
[0812] The following examples are provided to offer a complete disclosure and description of how to prepare and use the invention to those skilled in the art, and are not intended to limit the scope of the invention as the inventors perceive it to be, nor are they intended to represent all or only the experiments performed. Efforts have been made to ensure accuracy regarding the numerical values used (e.g., quantities, temperatures, etc.), but some experimental errors and biases should be accounted for. Unless otherwise stated, parts are parts by weight, molecular weight is weight average molecular weight, temperature is degrees Celsius, and pressure is atmospheric pressure or close to atmospheric pressure. Standard abbreviations may be used, such as bp, base pair; kb, kilobase; pl, picoliter; s or sec, second; min, minute; h or hr, hour; aa, amino acid; kb, kilobase; bp, base pair; nt, nucleotide; im, intramuscular; ip, intraperitoneal; sc, subcutaneous; etc.
[0813] Exemplary phenolic and catechol-containing intermediates can be synthesized using any convenient method. Methods applicable to the preparation of exemplary phenolic and catechol-containing intermediates of this disclosure include those described by Maza et al. in “Enzymatic Modification of N-Terminal Proline Residues Using Phenol Derivatives”, J. Am. Chem. Soc. (2019), 141, 3885-3892, the disclosure of which is incorporated herein by reference in its entirety. Numerous generally known chemical synthetic schemes and conditions applicable to the synthesis of the disclosed compounds are also available (see, for example, Smith and March, March's Advanced Organic Chemistry: Reactions, Mechanisms, and Structure, 5th ed., Wiley-Interscience, 2001; or Vogel, A Textbook of Practical Organic Chemistry, Including Qualitative Organic Analysis, 4th ed., New York: Longman, 1978). The reaction can be monitored by thin-layer chromatography (TLC) and LC / MS, and the reaction products can be analyzed by LC / MS and... 1 H NMR characterization.
[0814] Example 1
[0815] Unless otherwise stated, all chemicals were purchased from Sigma Aldrich. Peptides were purchased from Genescript. Peptide sequences can be found at Sigma.
[0816] Tyrosinase Coupling Reaction
[0817] MS2 conjugation was performed on a cysteine mutant of the capsid, which replaced the asparagine residue at position 87 with cysteine (N87C). The conjugation conditions were 20 mM pH 6.5 phosphate, 10 μM N87C MS2, and tyrosinase (CAS No. 9002-10-2) purchased from Sigma-Aldrich, diluted in 50 mM phosphate at pH 6.5, added at a 1:10 ratio to a final concentration of 0.16 μM. Unless otherwise specified, the conjugant was added to a final concentration of 50 μM or 5x the MS2 monomer concentration. Ultrapure water (Milipore Sigma, 18 uohm resistor) was added to a final volume of 20 μL. The reaction was carried out at room temperature for 30 min, then quenched with 2 μL of 20 mM tyrosine and 20 mM TCEP to obtain a final concentration of 2 mM for each.
[0818] For stability studies, the loose tyrosinase was replaced with an enzyme coupled to the resin. The reaction was carried out as described above, with a filtration step through a 0.2 μm filter to remove excess tyrosinase, followed by quenching with 1 mM tolphenidone and TCEP.
[0819] Cas9 coupling occurred under the following conditions: 20 mM Tris HCl, 300 mM KCl, 50 mM trehalose, pH 7.0 (Buffer A), 4°C for 1 hr, and 10 μM Cas9. All samples were quenched using the above quenching solution, and then the solvent was exchanged three times in Buffer A using a 100,000 kDa MWCO spin concentrator. For peptide coupling, peptides were added at a 5x ratio to obtain a peptide concentration of 50 μM. In protein-protein coupling, a 1:1 Cas9:target ratio yielded near-quantitative conversion to monomodified Cas9 after filtration, while a 1:5 Cas9:target ratio produced a completely doubly modified product.
[0820] Figure 1 illustrates the reaction scheme at the protein scale.
[0821] Figure 2 Figure A illustrates how maleimide-terminated thiols on proteins block the addition reaction catalyzed by tyrosinase, and conversely, where tyrosinase occurs first, the reaction with maleimide is also blocked.
[0822] Figure 2 Figure B illustrates a series of stability studies, demonstrating that the bond is stable over time under various buffer conditions.
[0823] Figure 3 Different arrays of substrates already used in the subject approach are shown.
[0824] Figure 4 Mass spectrometry data supporting the addition of multiple peptides using a subject-based approach are shown.
[0825] Figure 5 Figure A demonstrates that the topic approach can be used to modify Cas9.
[0826] Figure 5 Figure B demonstrates that the modified Cas9 remains active even when reacting on the apo protein, i.e., Cas9 does not have its guide RNA.
[0827] Figure 5 Figure C demonstrates that Cas9 can be modified with another protein, in this case GFP with an N-terminal tyrosine residue.
[0828] Figure 5 The D diagram demonstrates that the GFP-Cas9 conjugate retains activity.
[0829] Figure 6 It was demonstrated that Cas9 modified with a peptide, the 2NLS sequence Ac-YGPKKKRKVGGSPKKKRKV (SEQ ID NO:943), showed a 20-fold improvement in editing within neural progenitor cells.
[0830] Fi...
Claims
1. A method for chemically selectively modifying a target molecule, the method comprising: This allows target molecules containing thiol moieties to come into contact with biomolecules containing reactive moieties. The biomolecule containing the reactive portion is produced by reacting a biomolecule containing a phenolic or catecholic portion with an enzyme capable of oxidizing the phenolic or catecholic portion; and The contact is carried out under conditions sufficient to conjugate the target molecule with the biomolecule, thereby producing a modified target molecule.
2. The method of claim 1, wherein the target molecule is a polypeptide or a polynucleotide.
3. The method of claim 1 or claim 2, wherein the enzyme is a tyrosinase polypeptide.
4. The method according to any one of claims 1-3, wherein the tyrosinase polypeptide is Agaricus bisporus tyrosinase (abTYR) polypeptide.
5. The method of any one of claims 1-3, wherein the tyrosinase polypeptide comprises an amino acid sequence having at least 75% amino acid sequence identity with the abTYR amino acid sequence depicted in Figure 8 or Figure 9.
6. The method of claim 4 or claim 5, wherein the biomolecule comprising the phenolic moiety or the catechol moiety is neutral or positively charged within 50 Å of the phenolic or catechol moiety.
7. The method according to any one of claims 1-3, wherein the tyrosinase polypeptide is Bacillus megaterium tyrosinase (bmTYR) polypeptide.
8. The method of any one of claims 1-3, wherein the tyrosinase polypeptide comprises an amino acid sequence having at least 75% amino acid sequence identity with any one of the amino acid sequences depicted in any one of Figures 10A-10Z and Figures 10AA-10VV.
9. The method of claim 7 or claim 8, wherein the biomolecule comprising the phenolic or catechol moiety is negatively charged within 50 Å of the phenolic or catechol moiety.
10. The method of any one of claims 1-9, wherein the target molecule is a polynucleotide.
11. The method of claim 10, wherein the target molecule is a DNA molecule.
12. The method of claim 10, wherein the target molecule is an RNA molecule.
13. The method of any one of claims 10-12, wherein the biomolecule is a polypeptide.
14. The method of any one of claims 1 to 13, wherein the enzyme is bound to a solid carrier.
15. The method of any one of claims 1 to 14, wherein the phenolic moiety is present in a tyrosine residue.
16. The method of any one of claims 1 to 15, wherein the thiol moiety is present in a cysteine residue.
17. The method of claim 16, wherein the cysteine residue is a natural cysteine residue.
18. The method of any one of claims 1 to 17, wherein the biomolecule comprises one or more portions selected from: fluorophores, active small molecules, affinity tags, and metal chelators.
19. The method of any one of claims 1 to 18, wherein the reactive portion is an ortho-quinone or semi-quinone group, or a combination thereof.
20. The method of any one of claims 1 to 19, wherein the biomolecule is a polypeptide.
21. The method of claim 20, wherein the biomolecule is a polypeptide selected from fluorescent proteins, antibodies, enzymes, receptor ligands, and receptors.
22. The method of any one of claims 1 to 21, wherein the biomolecule comprising a phenolic or catechol moiety has formula (I), and the biomolecule comprising a reactive moiety has formula (II) or (IIA), or a combination thereof: in: Y 1 It is a biomolecule, which optionally comprises one or more portions selected from: active small molecules, affinity tags, fluorophores, and metal chelators; X 1 Selected from hydrogen and hydroxyl; and L is an optional connector.
23. The method of any one of claims 1 to 22, wherein the target molecule comprising the thiol moiety has formula (III), and wherein the modified target molecule has formula (IV) or (IVA), or a combination thereof: in: Y 1 It is a biomolecule, which optionally comprises one or more portions selected from: active small molecules, affinity tags, fluorophores, and metal chelators; Y 2 It is the second biomolecule; L is an optional connector; and n is an integer from 1 to 3.
24. The method of claim 23, wherein the modified target molecule of formula (IV) has any one of formulas (IV1)-(IV3): The target molecule modified according to formula (IVA) has any one of formulas (IVA1)-(IVA3): 。 25. The method of claim 23, wherein the modified target molecule of formula (IV) has any one of formulas (IV5)-(IV6): The target molecule modified according to formula (IVA) has any one of formulas (IVA4)-(IVA5): 。 26. The method of any one of claims 1 to 25, wherein the biomolecule comprising a phenolic or catechol moiety is described by formula (IA): in: Y 1 It is a biomolecule, which optionally includes one or more groups selected from the following: active small molecules, affinity tags, fluorophores and metal chelators; Each R 1 Independently selected from hydrogen, acyl, substituted acyl, alkyl, and substituted alkyl; X 1 Selected from hydrogen and hydroxyl; and L 1 It is a linker selected from straight-chain or branched alkyl groups, straight-chain or branched substituted alkyl groups, polyethylene glycol (PEG), substituted PEG, and one or more peptides.
27. The method of claim 26, wherein the fluorophore is a rhodamine dye or a xaton dye.
28. The method of any one of claims 1 to 27, wherein the modified target molecule is described by formula (IVB) or (IVC) or a combination thereof: in: Y 1 It is a biomolecule, which optionally includes one or more groups selected from the following: active small molecules, affinity tags, fluorophores and metal chelators; Each R 1 Independently selected from hydrogen, acyl, substituted acyl, alkyl, and substituted alkyl; Y 2 It is the second biomolecule; L 1 It is a linker selected from straight-chain or branched alkyl groups, straight-chain or branched substituted alkyl groups, polyethylene glycol (PEG), substituted PEG, and one or more peptides; and n is an integer from 1 to 3.
29. The method of claim 28, wherein the modified target molecule of formula (IVB) has any one of formulas (IVB1)-(IVZB3): The target molecule modified according to formula (IVC) has any one of formulas (IVC1)-(IVC3): 。 30. The method of claim 28, wherein the modified target molecule of formula (IVB) has any one of formulas (IVB5)-(IVB6): The target molecule modified according to formula (IVC) has any one of formulas (IVC4)-(IVC5): 。 31. The method of any one of claims 1 to 30, wherein the method is carried out at a pH of 4 to 9.
32. The method of claim 31, wherein the method is performed at a neutral pH.
33. The method of any one of claims 1 to 32, wherein the target molecule comprising a thiol group is a CRISPR-Cas effector polypeptide.
34. A composition comprising: The target molecule containing thiols in formula (III): Biomolecules of formula (I) containing a phenolic or catechol moiety: in: Y 1 It is a biomolecule, which optionally comprises one or more portions selected from: active small molecules, affinity tags, fluorophores, and metal chelators; X 1 Selected from hydrogen and hydroxyl groups; L is an optional connector; and Y 2 It is the second biomolecule.
35. The composition of claim 34, wherein the biomolecule comprising the phenolic or catechol moiety is neutral or positively charged within 50 Å of the phenolic or catechol moiety.
36. The composition of claim 34, wherein the biomolecule comprising the phenolic or catechol moiety is negatively charged within 50 Å of the phenolic or catechol moiety.
37. The composition of any one of claims 34-36, wherein the Y 1 It is a polypeptide and in which Y 2 It is a polypeptide.
38. The composition according to any one of claims 34-36, wherein Y 1 It is a polynucleotide and in which Y 2 It is a polypeptide.
39. The composition of any one of claims 34 to 38, wherein Y 2 It is a CRISPR-Cas effector polypeptide.
40. The composition according to any one of claims 34 to 39, wherein formula (I) is described by formula (IA): in: Y 1 It is a biomolecule, which optionally includes one or more groups selected from the following: active small molecules, affinity tags, fluorophores and metal chelators; Each R 1 Independently selected from hydrogen, acyl, substituted acyl, alkyl, and substituted alkyl; X 1 Selected from hydrogen and hydroxyl; and L 1 It is a linker selected from straight-chain or branched alkyl groups, straight-chain or branched substituted alkyl groups, polyethylene glycol (PEG), substituted PEG, and one or more peptides.
41. A reagent kit comprising: A first container, the first container comprising the composition as described in any one of claims 34 to 40; as well as A second container contains an enzyme capable of oxidizing the phenol or catechol moiety.
42. The kit of claim 41, wherein the enzyme is a tyrosinase polypeptide.
43. The kit of claim 42, wherein the tyrosinase is Agaricus bisporus tyrosinase (abTYR).
44. The kit of claim 42, wherein the tyrosinase polypeptide comprises an amino acid sequence having at least 75% amino acid sequence identity with the abTYR amino acid sequence depicted in Figure 8 or Figure 9.
45. The kit of claim 42, wherein the tyrosinase is Bacillus megaterium tyrosinase (bmTYR).
46. The kit of claim 42, wherein the tyrosinase polypeptide comprises an amino acid sequence having at least 75% amino acid sequence identity with any one of the amino acid sequences depicted in any of Figures 10A-10Z and Figures 10AA-10VV.
47. A compound of formula (IV) or (IVA): in: Y 1 It is a biomolecule, which optionally comprises one or more portions selected from: active small molecules, affinity tags, fluorophores, and metal chelators; L is an optional connector; Y 2 It is the second biomolecule; and n is an integer from 1 to 3.
48. The compound of claim 47, wherein the modified target molecule of formula (IV) has any one of formulas (IV1)-(IV5): 。 49. The compound of claim 47, wherein the modified target molecule of formula (IVA) has any one of formulas (IVA1)-(IVA5): 。 50. The compound of any one of claims 47 to 49, wherein L is a pyrolytic connector.
51. The compound according to any one of claims 47 to 50, wherein Y 1 It is a polypeptide.
52. The compound of claim 51, wherein Y 1 Selected from fluorescent proteins, antibodies, and enzymes.
53. The compound according to any one of claims 47 to 52, described by formula (IVB) or (IVC): in: Y 1 It is a biomolecule, which optionally includes one or more groups selected from the following: active small molecules, affinity tags, fluorophores and metal chelators; Each R 1 Independently selected from hydrogen, acyl, substituted acyl, alkyl, and substituted alkyl; Y 2 It is the second biomolecule; L 1 It is a linker selected from straight-chain or branched alkyl groups, straight-chain or branched substituted alkyl groups, polyethylene glycol (PEG), substituted PEG, and one or more peptides; and n is an integer from 1 to 3.
54. The compound of claim 53, wherein the modified target molecule of formula (IVB) has any one of formulas (IVB1)-(IVZB5): 。 55. The compound of claim 53, wherein the modified target molecule of formula (IV) has any one of formulas (IVC1)-(IVC5): 。 56. The compound according to any one of claims 47 to 55, wherein it is described by any one of formulas (IVD)-(IVG): in: R 2 Selected from alkyl and substituted alkyl groups; R 3 Selected from hydrogen, alkyl, substituted alkyl, peptide and polypeptide; and n is an integer from 1 to 3.
57. The compound according to any one of claims 47 to 56, wherein Y 2 It is a CRISPR-Cas effector polypeptide.
58. A method for chemically selectively coupling a first polypeptide and a second polypeptide to a coupled polypeptide, the method comprising: a) Contact the first polypeptide with the coupled polypeptide to generate a first polypeptide-coupled polypeptide conjugate. The first polypeptide contains a thiol moiety. The coupled polypeptide includes an N-terminal reactive portion that forms a covalent bond with the thiol moiety present in the first polypeptide. The conjugated polypeptide, which includes the N-terminal reactive portion, is produced by reacting a polypeptide comprising an N-terminal phenolic or catechol portion and a C-terminal phenolic or catechol portion with a first enzyme capable of oxidizing the N-terminal phenolic or catechol portion but not the C-terminal phenolic or catechol portion to produce the N-terminal reactive portion. The coupling polypeptide comprises two or more positively charged or neutral amino acids within ten amino acids of the N-terminal phenolic or catechol moiety and comprises two or more negatively charged amino acids within ten amino acids of the C-terminal phenolic or catechol moiety; and b) Contact the second polypeptide with the first polypeptide-coupled polypeptide conjugate. The second polypeptide contains a thiol moiety. The first polypeptide-coupled polypeptide conjugate includes a C-terminal reactive moiety that forms a covalent bond with the thiol moiety present in the second polypeptide. The first polypeptide-coupled polypeptide conjugate, which includes the C-terminal reactive portion, is produced by reacting the first polypeptide-coupled polypeptide conjugate with a second enzyme capable of oxidizing the C-terminal phenolic or catechol portion to generate the C-terminal reactive portion; and The contact produces a first polypeptide-coupled polypeptide-second polypeptide conjugate.
59. The method of claim 58, wherein: a) The first enzyme is a tyrosinase polypeptide containing an amino acid sequence having at least 75% amino acid sequence identity with the abTYR amino acid sequence depicted in Figure 8 or Figure 9; and b) The second enzyme is a tyrosinase polypeptide containing an amino acid sequence that has at least 75% amino acid sequence identity with any one of the amino acid sequences depicted in either of Figures 10A-10Z and Figures 10AA-10VV.
60. A method for chemically selectively coupling a first polypeptide and a second polypeptide to a coupled polypeptide, the method comprising: a) Contact the first polypeptide with the coupled polypeptide to generate a first polypeptide-coupled polypeptide conjugate. The first polypeptide contains a thiol moiety. The coupled polypeptide includes an N-terminal reactive portion that forms a covalent bond with the thiol moiety present in the first polypeptide. The conjugated polypeptide, which includes the N-terminal reactive portion, is produced by reacting a polypeptide comprising an N-terminal phenolic or catechol portion and a C-terminal phenolic or catechol portion with a first enzyme capable of oxidizing the N-terminal phenolic or catechol portion but not the C-terminal phenolic or catechol portion to produce the N-terminal reactive portion. The coupling polypeptide comprises two or more negatively charged amino acids within ten amino acids of the N-terminal phenolic or catechol moiety and comprises two or more positively charged or neutral amino acids within ten amino acids of the C-terminal phenolic or catechol moiety; and b) Contact the second polypeptide with the first polypeptide-coupled polypeptide conjugate. The second polypeptide contains a thiol moiety. The first polypeptide-coupled polypeptide conjugate includes a C-terminal reactive moiety that forms a covalent bond with the thiol moiety present in the second polypeptide. The first polypeptide-coupled polypeptide conjugate, which includes the C-terminal reactive portion, is produced by reacting the first polypeptide-coupled polypeptide conjugate with a second enzyme capable of oxidizing the C-terminal phenolic or catechol portion to generate the C-terminal reactive portion; and The contact produces a first polypeptide-coupled polypeptide-second polypeptide conjugate.
61. The method of claim 60, wherein: a) The first enzyme is a tyrosinase polypeptide comprising an amino acid sequence having at least 75% amino acid sequence identity with any one of the amino acid sequences depicted in any of Figures 10A-10Z and 10AA-10VV; and b) The second enzyme is a tyrosinase polypeptide containing an amino acid sequence that has at least 75% amino acid sequence identity with the abTYR amino acid sequence depicted in Figure 8 or Figure 9.
62. A method for covalently linking a first polypeptide to a second polypeptide, the method comprising: a) Contact the first polypeptide with the immobilized reactive portion. The fixed reactive portion is generated by reacting a fixed phenolic or catechol portion with a first enzyme, wherein the first enzyme is capable of oxidizing the fixed phenolic or catechol portion, thereby generating the fixed reactive portion. The first polypeptide comprises: i) a thiol moiety; and ii) a phenolic or catechol moiety, wherein the first polypeptide comprises two or more negatively charged amino acids within ten amino acids of the phenolic or catechol moiety. The fixed reactive portion forms a covalent bond with the thiol portion present in the first polypeptide, thereby producing a fixed first polypeptide; b) Contacting the fixed first polypeptide with a second enzyme, wherein the second enzyme is capable of oxidizing the phenolic or catechol moiety present in the first polypeptide to produce a fixed first polypeptide comprising a reactive moiety; and c) Contact the fixed first polypeptide containing the reactive portion with the second polypeptide. The second polypeptide comprises: i) a thiol moiety; and ii) a phenolic or catechol moiety, wherein the second polypeptide comprises two or more neutral or positively charged amino acids within ten amino acids of the phenolic or catechol moiety. The reactive portion present in the fixed first polypeptide forms a covalent bond with the thiol portion present in the second polypeptide, thereby producing a fixed conjugate comprising the first polypeptide covalently linked to the second polypeptide.
63. The method of claim 62, wherein the first enzyme is a tyrosinase polypeptide comprising an amino acid sequence having at least 75% amino acid sequence identity with any one of the amino acid sequences depicted in any of Figures 8, 9, 10A-10Z and 10AA-10VV.
64. The method of claim 62 or claim 63, wherein the thiol moiety present in the first polypeptide is present in Cys, and wherein the phenolic moiety present in the first polypeptide is present in Tyr residues.
65. The method of claim 64, wherein the Tyr residue is present in an amino acid segment comprising EEEY (SEQ ID NO: 953), EEEEY (SEQ ID NO: 955), DDDDY (SEQ ID NO: 965), or DDDDY (SEQ ID NO: 965).
66. The method of any one of claims 62-65, wherein the second enzyme is a tyrosinase polypeptide comprising an amino acid sequence having at least 75% amino acid sequence identity with any one of the amino acid sequences depicted in any one of Figures 10A-10Z and Figures 10AA-10VV.
67. The method of any one of claims 62-66, further comprising: c) Contacting the fixed conjugate with a third enzyme, wherein the third enzyme is capable of oxidizing the phenolic or catechol moiety present in the second polypeptide to produce a fixed conjugate comprising a reactive moiety; and c) Contact the fixed conjugate containing the reactive portion with the third polypeptide. The third polypeptide comprises: i) a thiol moiety; and ii) a phenolic or catechol moiety, wherein the third polypeptide contains two or more negatively charged amino acids within ten amino acids of the phenolic or catechol moiety. The reactive portion present in the fixed conjugate forms a covalent bond with the thiol portion present in the second polypeptide, thereby producing a fixed conjugate comprising the third polypeptide covalently linked to the second polypeptide.
68. The method of claim 67, wherein the third enzyme is a tyrosinase polypeptide comprising an amino acid sequence having at least 75% amino acid sequence identity with the amino acid sequence depicted in FIG8 or FIG9.
69. The method of claim 67 or 68, wherein between step (b) and step (c), the second enzyme is inactivated or removed.
70. The method of any one of claims 67-69, wherein the thiol moiety present in the second polypeptide is present in Cys, and wherein the phenolic moiety present in the second polypeptide is present in Tyr residues.
71. The method of claim 70, wherein the Tyr residue is present in an amino acid segment comprising RRRY (SEQ ID NO: 949), RRRRY (SEQ ID NO: 951), KKKY (SEQ ID NO: 966), or KKKKY (SEQ ID NO: 967).
Citation Information
Patent Citations
Processes for the production of multichain polypeptides or proteins
EP0120694B1
Recombinant immunoglobulin preparations, methods for their preparation, DNA sequences, expression vectors and recombinant host cells therefor
EP0125023B1
Production of chimeric antibodies
EP0194276B1
Recombinant antibodies and methods for their production
EP0239400B1
Bispecific and oligospecific, mono- and oligovalent receptors, production and applications thereof
EP0404097A2
Cited By
Tyrosinase mutant, immobilized enzyme and enzymatic synthesis method of 5, 6-dihydroxyindole
CN122128257A