Sequence-based framework design peptide-oriented degradation agent

The SaLT&PepPr module generates peptide-oriented degrading agents without structural information, solving the programmability and widespread targets of targeted protein degradation in the prior art, and achieving efficient intracellular degradation of a variety of pathogenic targets, and applying them to tumor treatment.

CN120476446APending Publication Date: 2025-08-12UBIQUITX INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380090579.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-07
Filing Date
2023-11-07
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing targeted protein degradation technology relies on small molecule screening and structural information, making it difficult to effectively degrade unstructured or unstable target proteins, and lacks programmability, so it cannot be widely used in a variety of pathogenic targets.

Method used

Using a pre-trained language model and protein interaction database, a peptide-oriented degrading agent without structural information was generated through the SaLT&PepPr module. A prediction model and a peptide synthesis device were used to design a peptide sequence with high binding affinity with the target protein, and targeted degradation was carried out by combining E3 ubiquitin ligase.

Benefits of technology

Peptide-oriented degradation that efficiently recognizes multiple pathogenic targets without structural information is achieved, and can reliably induce target protein degradation in cells, and is applied to diagnosis, analysis and treatment, especially degradation of endogenous β-catenin to regulate Wnt signaling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120476446A_ABST
    Figure CN120476446A_ABST
Patent Text Reader

Abstract

A method of generating a peptide sequence binding to a target sequence, the method comprising: receiving a data object corresponding to a target protein using a processor configured by code executed therein; retrieving at least one chaperone protein of the target protein in a protein interaction database using the data object; at least one chaperone protein identifying the target protein; providing the at least one chaperone protein to a computational model configured to output a predicted protein sequence that predicts interaction with the target sequence; and identifying at least one sub-sequence of the predicted protein sequence that meets a predetermined interaction threshold.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. patent application No. 63 / 423,320, filed on November 7, 2022, the entire contents of which are incorporated herein by reference. Technical Field

[0002] The present disclosure relates to systems and methods that provide a unified, sequence-based framework for designing peptide-directed degraders without the need for structural information, wherein such peptide-directed degraders are useful in diagnostic, analytical, and therapeutic applications, and compositions related thereto. Background Art

[0003] Curing malignant tumors is one of the greatest challenges facing future human health, and targeted therapies have emerged as an effective solution to this problem. Small molecule inhibitors, in particular, have achieved remarkable success in the clinic, but their therapeutic potential remains limited. Most notably, they are “occupancy-driven” and therefore rely on high doses and must bind to active sites that are either absent or difficult to access on traditional “undruggable” target proteins. To overcome these limitations, targeted protein degradation (TPD) technology offers a unique opportunity to transiently bind to intracellular proteins and induce their degradation by hijacking the cell’s native ubiquitin-proteasome pathway (UPP). 1 For example, proteolysis-targeting chimeras (PROTACs) and molecular glues utilize small molecules that can both bind to target proteins and recruit endogenous E3 ubiquitin ligases, thereby enabling the transfer of ubiquitin to the target protein for subsequent proteasomal degradation. However, small-molecule-based degraders lack programmability: they require extensive small-molecule screening and design at the target end, currently utilize only a few of the approximately 600 E3 ubiquitin ligases, and are unable to degrade proteins without accessible binding sites. As a more modular strategy, we and others have recently fused compact protein binders (including FN3s, DARPins, nanobodies, and peptides) to various E3 ubiquitin ligase domains to achieve binding, selective ubiquitination, and intracellular degradation of a wide range of pathogenic targets. Compared with the more restrictive small-molecule-based approaches, constructing a programmable system for designing these genetically encoded “ubiquitin antibodies” (uAbs) would represent a more efficient TPD approach.

[0004] Previous approaches to binder design have included high-throughput screening and structure-based rational design. Recent computational protein design tools include interface predictors, docking software, and repair models that leverage advances in protein structure prediction (e.g., AlphaFold) to infer novel sequences from user-specified structures. These algorithms (e.g., ProteinMPNN) rely heavily on the existence of cocrystal complexes or accurate structure predictions of the target protein, thereby excluding disordered or unstable proteins (e.g., transcription factors) that have important impacts on disease and are difficult to resolve through experimental or computational protein structure determination methods. Recently, language models have been pre-trained on millions of natural protein sequences to generate latent embeddings that capture relevant physicochemical, functional, and, most importantly, tertiary structural information. Transfer learning using these models allows the prediction of peptide binding sites based on sequence and structure in a rotationally and translationally invariant manner. More interestingly, early results suggest that sequence-based protein transformers can generate novel protein sequences with functional capabilities.

[0005] In: Békés, M., Langley, DR&Crews, CMPROTAC Nat.Rev.Drug Discov.21,181–200(2022).Schreiber,SLThe RiseofMolecular Glues.Cell 184,3–9(2021).Gao,H.,Sun,X.&Rao,Y.PROTAC Technology:Opportunities and Challenges.ACS Med.Chem.Lett.11,237–240(2020).Portnoff,AD,Stephens,EA,Varner,JD&DeLisa,MPUbiquibodies,synthetic E3 ubiquitinligases endowedwith unnatural substrate specificity for targetedproteinsilencing.J.Biol.Chem.289,7844–7855(2014).Chatterjee,P.et al.Targetedintracellular degradation of SARS-CoV-2via computationally optimized peptidefusions.Communications Biology 3,1–8(2020).Stephens,EAet al.EngineeringSingle Pan-Specific BioPROTACs as Versatile Modulators of Intracellular Therapeutic Targetsincluding Proliferating Cell Nuclear Antigen(PCNA).ACS Synth.Biol.10,2396–2408(2021).Lim,S.etal117,5791–5800(2020).Sheehan,J.&Marasco,W.A.Phageand Yeast Display.Microbiol Spectr 3,AID–0028–2014(2015).Barlow,K.A.etal.Flex ddG:Rosetta Ensemble-Based Estimation ofChanges in Protein-ProteinBinding Affinity upon Mutation.J.Phys.Chem.B 122,5389–5399(2018).Chevalier,A.et al.Massively parallel de novo protein design for targetedtherapeutics.Nature 550,74–79(2017).Pagadala,N.S.,Syed,K.&Tuszynski,J.Software for molecular docking:a review.Biophys.Rev.9,91–102(2017).Abdin,O.,Nim,S.,Wen,H.&Kim,P.M.PepNN:a deep attention model for the identificationofpeptide binding sites.Communications Biology 5,1–10(2022).Cao,L.etal.Design ofprotein-binding proteins from the target structure alone.Nature605,551–560(2022).Anishchenko,I.et al.De novo protein design by deep networkhallucination.Nature 600,547–552(2021).Jumper,J.et al.Highly accurate proteinstructure prediction with AlphaFold.Nature 596,583–589(2021).Dauparas,J.etal.Robust deep learning-basedprotein sequence design usingProteinMPNN.Science eadd2187(2022).Varadi,M.et al.AlphaFold Protein StructureDatabase:massively expanding the structural coverage ofprotein-sequence spacewith high-accuracy models.Nucleic Acids Res.50,D439–D444(2022).Rives,A.etal.Biological structure and function emerge from scaling unsupervisedlearning to 250million protein sequences.Proc.Natl.Acad.Sci.U.S.A.118,(2021).Madani,A.et al.ProGen:Language Modeling for Protein Generation.bioRxiv2020.03.07.982272(2020)doi:10.1101 / 2020.03.07.982272.Elnaggar,A.etal.ProtTrans:Toward Understanding the Language ofLife Through Self-SupervisedLearning.IEEE Trans.Pattern Anal.Mach.Intell.44,7112–7127(2022).Madani,A.etal.Deep neural language modeling enables functional protein generation acrossfamilies.bioRxiv 2021.07.18.452833(2021)doi:10.1101 / 2021.07.18.452833.Bian,J.,Dannappel,M.,Wan,C.&Firestein,R.Transcriptional Regulation of Wnt / β-Catenin Pathway in Colorectal Cancer.Cells 9,(2020).Zhao,H.et al.Wntsignaling in colorectal cancer:pathogenic role and therapeutictarget.Mol.Cancer 21,1–34(2022).Khalaf,A.M.et al.Role of Wnt / β-cateninsignaling in hepatocellular carcinoma,pathogenesis,and clinicalsignificance.J Hepatocell Carcinoma 5,61–73(2018).Sedan,Y.,Marcu,O.,Lyskov,S.&Schueler-Furman,O.Peptiderive server:derive peptide inhibitors fromprotein-protein interactions.Nucleic Acids Res.44,W536–41(2016).Radford,A.etal.Learning Transferable Visual Models From Natural Language Supervision.(2021)doi:10.48550 / arXiv.2103.00020.Rao,R.et al.MSA Transformer.bioRxiv2021.02.12.430858(2021)doi:10.1101 / 2021.02.12.430858.Porras,P.et al.Towards aunified open access dataset of molecular interactions.Nat.Commun.11,1–12(2020).Oughtred,R.et al.The BioGRID interaction database:2019update.NucleicAcids Res.47,D529–D541(2018).Johnson,K.L.et al.Revealing protein-proteininteractions at the transcriptome scale by sequencing.Mol.Cell 81,4091–4103.e9(2021).Palepu,K.et al.Design of Peptide-Based Protein Degraders viaContrastive Deep Learning.bioRxiv 2022.05.23.493169(2022)doi:10.1101 / 2022.05.23.493169.Lin,Z.et al.Language models ofprotein sequences at thescale of evolution enable accurate structure prediction.bioRxiv2022.07.20.500902(2022)doi:10.1101 / 2022.07.20.500902.Mirdita,M.etal.ColabFold:making protein folding accessible to all.Nat.Methods 19,679–682(2022).Evans,R.et al.Protein complex prediction with AlphaFold-Multimer.bioRxiv 2021.10.04.463034(2022)doi:10.1101 / 2021.10.04.463034.Huang,L.,Guo,Z.,Wang,F.&Fu,L.KRAS mutation:from undruggable to druggable incancer.Signal Transduction and Targeted Therapy 6,1–20(2021).Graham,R.P.etal.DNAJB1-PRKACA is specific for fibrolamellar carcinoma.Mod.Pathol.28,822–829(2015).Dong,X.C.PNPLA3—A Potential Therapeutic Target for PersonalizedTreatment ofChronic Liver Disease.Front.Med.0,(2019).Hou,X.,Zaks,T.,Langer,R.&Dong,Y.Lipid nanoparticles for mRNA delivery.Nature Reviews Materials 6,1078–1094(2021).Tran,T.H.et al.KRAS interaction with RAF1 RAS-binding domain and cysteine-rich domain provides insights into RAS-mediated RAFactivation. Nat. Commun. 12, 1–16 (2021). Zhan, T., Rindtorff, N. & Boutros, M. Wntsignaling in cancer. Oncogene 36,1461–1473(2016).Nusse,R.&Clevers,H.Wnt / β-Catenin Signaling,Disease,and Emerging Therapeutic Modalities.Cell 169,985–999(2017).Korinek,V.et al.Constitutive transcriptional activation by a beta-catenin-Tcfcomplex in APC- / -colon carcinoma.Science 275,1784–1787(1997).Ludwicki,MBet al.Broad-Spectrum Proteome Editing with an EngineeredBacterial Ubiquitin Ligase Mimic.ACS Cent Sci 5,852–866(2019).Cong,F.,Zhang,J.,Pao,W.,Zhou,P.&Varmus,HA protein knockdown strategy to study the function ofbeta-catenin in tumorigenesis.BMC Mol.Biol.4,10(2003). Summary of the Invention

[0006] In one or more embodiments of the subject matter described herein, systems and methods are provided for implementing a unified, sequence-based framework for designing peptide-directed degraders in the absence of structural information, wherein such peptide-directed degraders are useful in diagnostic, analytical, and therapeutic applications, and compositions related thereto.

[0007] In a specific embodiment, the inventors discovered that by using a pre-trained language model and a protein interaction database, a unified, sequence-based framework can be generated to design peptide-directed degraders without structural information. This sequence-based system is referred to throughout this article as the Structure-agnostic Language Transformer & Peptide Prioritization (SaLT&PepPr) module. The SaLT&PepPr module can effectively down-select peptides from known binding protein sequences for downstream screening.

[0008] In one or more embodiments, a system for generating a binding protein sequence configured to bind to a target protein sequence is provided. In certain configurations, the system includes a pre-trained prediction model, wherein the pre-trained prediction model includes a protein language model, and the protein language model is configured to output position data to a multilayer perceptron classification neural network, wherein the perceptron classification neural network is configured to output a value corresponding to the probability of each amino acid sequence binding to each position of the target sequence. The system is further configured to generate a binding sequence that binds to the target sequence based on the output probabilities.

[0009] The system may further include one or more peptide synthesis devices configured to synthesize one or more peptide sequences generated by the pre-trained prediction model.

[0010] In a further embodiment, a peptide generation system is provided. The peptide generation system includes: a search engine configured to search an interactome database; a prediction engine configured to generate proposed binding sequence peptides based on the search results of the interactome database; and an output module configured to extract subsequences from the proposed binding sequences that are predicted to have a binding affinity with the target above a predetermined threshold.

[0011] A method for predicting the similarity between target molecules and peptide molecules is also provided, the method comprising: generating feature-rich embeddings for multiple target molecules and peptide molecules using a pre-trained ESM-2 model; forming a matrix whose rows correspond to target molecules and columns correspond to peptide molecules; using the trained model to predict the cosine similarity between each pair of target molecules and peptide molecules in the matrix; calculating the average of the cross-entropy losses of the rows and columns of the matrix; and outputting the predicted cosine similarities and the average cross-entropy losses as a measure of the performance of the trained model.

[0012] In one embodiment, a method for generating a peptide sequence that binds to a target sequence is provided, the method comprising: receiving a data object corresponding to a target protein using a processor configured by code executed therein; using the data object to retrieve at least one partner protein of the target protein in a protein interaction database; identifying at least one partner protein of the target protein; providing the at least one partner protein to a computational model, the computational model configured to output a predicted protein sequence predicted to interact with the target sequence; and identifying at least one subsequence in the predicted protein sequence that meets a predetermined interaction threshold.

[0013] In one or more further embodiments, a chimeric molecule is provided that comprises one or more peptides generated using a SaLT&PepPr-derived sequence. In one approach, such a chimeric molecule is used to post-translationally modify a target of interest. For example, the SaLT&PepPr-derived sequence is configured to link the target to one or more E3 ubiquitin ligase domains. Such a chimeric molecule can be used for targeted degradation of a biological target of a specific interest.

[0014] In further embodiments, the SaLT&PepPr module can reliably identify candidates that exhibit potent intracellular degradation of multiple pathogenic targets in human cells, including those with minimal structural information.

[0015] In another embodiment, a peptide-directed degrader is provided, wherein the peptide is generated using the SaLT&PepPr module and has negligible off-target effects as determined by whole-cell proteomics. Such a peptide-directed degrader is capable of degrading endogenous β-catenin in a colorectal cancer cell model and subsequently downregulating Wnt signaling, thus potentially being used as a drug for the treatment of a variety of diseases. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In the accompanying drawings, which are not necessarily drawn to scale, like numbers may represent similar components in different views. Like numbers with different letter suffixes may represent different instances of similar components. The accompanying drawings generally illustrate various embodiments discussed herein by way of example and not limitation.

[0017] Figure 1 is a block diagram detailing the arrangement of components of the system described herein according to one embodiment of the present invention.

[0018] Figure 2 is a schematic diagram showing the relationship between the various modules of the system.

[0019] Figure 3 is a flowchart detailing the process of generating new peptide sequences.

[0020] Figure 4 is a graph detailing the percentage of the proteome with documented protein-protein interactions.

[0021] Figure 5 is a flowchart detailing the data flow in one embodiment of the model.

[0022] Figure 6 is a table detailing the relevant performance of this method.

[0023] Figure 7 is a table that benchmarks the model against different protein prediction methods.

[0024] Figure 8 A comparison of the output of the model in this application with the output of other methods is shown.

[0025] Figure 9 A comparison of the output of the model in this application with the output of other methods is shown.

[0026] Figure 10 Flowchart detailing the steps for generating new peptide sequences based on target input.

[0027] Figures 11A-11D The characterization of derived uAbs for target modulation is described in detail.

[0028] Figures 12A-12D The characterization of derived uAbs for target modulation is described in detail. DETAILED DESCRIPTION

[0029] As an overview and introduction, the subject matter of this application relates to systems and methods for generating peptides that can bind to identified targets. Specifically, this application provides a system, method, and approach for generating peptides based solely on sequence data without using structural information.

[0030] Inspired by the programmability of RNA-guided genome editing, the systems, methods, and computer-implemented methods utilize one or more pre-trained protein language models to design short "guide" peptide sequences that can bind to user-identified or computationally identified target proteins.

[0031] By training the computer-predicted binding affinities on the protein interface, peptide fragments that significantly contribute to the overall binding energy of the target protein can be isolated. In addition, by using predicted or experimentally verified binding proteins and specific target proteins as starting scaffolds for splicing, the structure-independent language converter and peptide prioritization (SaLT&PepPr) module can be used to generate peptides with a high probability of binding to the target protein.

[0032] In one or more specific embodiments, these designed peptide sequences are used as components of chimeric molecules that are designed to induce post-translational modifications. For example, but not limited to, when these designed peptide sequences are fused to modular E3 ubiquitin ligase domains, they are configured to induce target degradation. However, other post-translational modifications can also be used with the computationally designed peptide sequences.

[0033] For example, in one or more embodiments, a system for generating a binding protein sequence is provided, wherein the binding protein sequence is configured to bind to a target protein sequence. In a specific configuration, the system includes a pre-trained prediction model, wherein the pre-trained prediction model includes a protein language model, and the protein language model is configured to output position data to a multi-layer perceptron classification neural network, wherein the perceptron classification neural network is configured to output a value corresponding to the probability of each amino acid sequence binding to each position of the target sequence. The system is also configured to generate a binding sequence that binds to the target sequence based on the output probability.

[0034] The sequence-based binding prediction system SaLT&PepPr can generate peptides reliably and efficiently. When these peptides are incorporated into uAb constructs, the resulting constructs can induce robust post-translational modifications. For example, such methods can be used to induce the degradation of a variety of pathogenic targets in human cells, but are not limited thereto. For example, the method can be used to develop degraders for β-catenin, a cytoplasmic variant of which can lead to abnormal Wnt signaling, thereby inducing a variety of cancers, including colorectal cancer and hepatocellular carcinoma. In particular, the uAbs designed by SaLT&PepPr can bind with high affinity, induce endogenous β-catenin degradation, and subsequently downregulate Wnt signaling in colorectal cancer cell models, thereby promoting the clinical translation of such uAb platforms.

[0035] Now let's see Figure 1 , the system includes a user interface device 101 , an interactome database 106 , a peptide generation system 108 and one or more output devices 114 .

[0036] like Figure 1 and Figure 10 As shown, the protein target selections made by the user through the user interface device 101 are provided to the processor 108. In a specific embodiment, the user interface device 101 is a standard computing device, such as a desktop or portable computing device. However, in certain embodiments, the user interface device 101 is a custom computing platform specifically designed to perform the tasks described herein.

[0037] like Figure 1As shown, user interface device 101 is configured to transmit one or more protein selections to a peptide processing platform, such as processing platform 108. In one or more configurations, user interface device 101 is equipped or configured with a network interface or protocol that can be used to communicate over a network (e.g., the Internet). In such configurations, the user's selections can be from a stored local protein collection, a public protein database, or other protein data source.

[0038] Optionally, the user interface device 101 is connected to one or more computers or processors (eg, processing platform 108) using a standard interface (eg, USB, FIREWIRE, Wi-Fi, Bluetooth, and other wired or wireless communication technologies suitable for transmitting protein data).

[0039] The protein data generated or transmitted to one or more processing platforms 108 is then evaluated as a function of one or more hardware or software modules. As used herein, the term "module" generally refers to one or more discrete components that contribute to the effectiveness of the presently described systems, methods, and protocols. A module may include software elements, including but not limited to functions, algorithms, classes, etc. In one embodiment, a software module is stored as software in the memory 205 of the processing platform 108, such as Figure 2 shown.

[0040] In some embodiments, the modules may include discrete or specific hardware components. In one embodiment, the user interface device 101 and the processing platform 108 are located in the same device. For example, each of the user interface 101 and the processing platform 108 is a software module executed by one or more processors of a server, computing cluster, or other data processing and execution platform. However, in another embodiment, the processing platform 108 and the user interface device 101 may be remote or separate and communicate via one or more communication links.

[0041] In one configuration, the processing platform 108 is configured via one or more software modules to generate, calculate, process, output, or otherwise manipulate data provided by the user interface device 101 .

[0042] In one embodiment, processing platform 108 is a commercially available computing device. For example, processing platform 108 may be a collection of computers, servers, processors, cloud computing units, micro-computing units, computers on a chip, home entertainment consoles, media players, set-top boxes, prototyping equipment, or "hobby" computing units.

[0043] Furthermore, depending on the specific embodiment, processing platform 108 may include a single processor, multiple discrete processors, a multi-core processor, or other types of processors known to those skilled in the art. In a specific example, processing platform 108 executes software code on hardware such as a custom or commercially available cell phone, smartphone, laptop, workstation, or desktop computer that is configured to receive data or measurements captured by sample color sensor 106 directly or via a communication link.

[0044] The processing platform 108 is configured to execute a commercially available or customized operating system, such as an operating system based on Microsoft WINDOWS, Apple OSX, UNIX, or Linux, to execute instructions or code.

[0045] In one or more embodiments, processing platform 108 is further configured to access various peripheral devices and network interfaces. For example, processing platform 108 is configured to communicate with one or more remote servers, computers, peripheral devices, or other hardware via the Internet using standard or custom communication protocols and settings (e.g., TCP / IP, etc.).

[0046] Processing platform 108 may include one or more memories. The memories may be persistent or non-persistent storage devices (e.g., IC memory elements) that may be used to store an operating system and one or more software modules. According to one or more embodiments, the memories include one or more volatile and non-volatile memories, such as read-only memory ("ROM"), random access memory ("RAM"), electrically erasable programmable read-only memory ("EEPROM"), phase change memory ("PCM"), single inline memory ("SIMM"), dual inline memory ("DIMM"), or other types of memory. Such memories may be fixed or removable, as known to those skilled in the art, such as by using removable media cards or modules. In one or more embodiments, the memories of processing platform 108 are used to store application programs and data files. The one or more memories provide program code that processing platform 108 reads and executes upon receiving a startup or initialization signal.

[0047] The computer memory may also include secondary computer memory, such as a magnetic disk drive, an optical disk drive, or flash memory, which provides long-term storage of data in a manner similar to persistent memory devices. In one or more embodiments, the memory of processing platform 108 may store application programs and data files as needed.

[0048] The processing platform 108 is configured to store data locally in one or more storage devices. Alternatively, the processing platform 108 is configured to store data (e.g., data or processing results) in a local or remotely accessible database 112. The physical structure of the database 112 can be embodied as a solid-state memory (e.g., ROM), a hard drive system, a RAID, a disk array, a storage area network ("SAN"), a network attached storage ("NAS"), and / or any other suitable system for storing computer data. In addition, the database 112 may also include a cache, including a database cache and / or a web cache. From a programming perspective, the database 112 can include a flat file data store, a relational database, an object-oriented database, a relational-object hybrid database, a key-value data store (e.g., HADOOP or MONGODB), and other systems for constructing and retrieving data that are well known to those skilled in the art. The database 112 includes the necessary hardware and software to enable the processing platform 108 to retrieve and store data within the database 112.

[0049] In one embodiment, Figure 1 Each element provided in the system is configured to communicate with each other via one or more direct connections (e.g., via a common bus). Alternatively, each element is configured to communicate with other elements via a network connection or interface (e.g., a local area network (LAN) or a data cable connection). In another embodiment, the user interface device 101, the processing platform 108, and the database 112 can all be connected to a network 110 (e.g., the Internet) and configured to communicate and exchange data using well-known and well-understood communication protocols.

[0050] In a particular embodiment, processing platform 108 is a computer, workstation, thin client, or portable computing device, such as an Apple or device or other commercially available mobile electronic devices configured to receive data from database 112 or 106 or output data to database 112 or 106 .

[0051] In one configuration, processing platform 108 communicates with output device 114 to transmit, generate, display, or exchange data. In one configuration, output device 114 and processing platform 108 are integrated into a single device form factor, such as a computing device connected to one or more protein synthesis devices or systems. For example, such a device or system can be a workbench or desktop protein synthesis device. In this context, such an integrated system includes one or more computers or other data processing devices, tools, apparatus, or reaction or synthesis elements required to synthesize a peptide sequence.

[0052] Those skilled in the art will recognize that additional functional features, such as power supplies, power supplies, power management circuits, control interfaces, relays, adapters and / or other components for powering and interconnecting electronic components and controlling activation, are considered to be included in the present invention and may be understood to be employed.

[0053] like Figure 1 、 2 As shown in Figure 10, the user submits a target protein to the system. The target protein is used to retrieve an interactome database (e.g., database 106). For example, the processor of the user interface device 101 is configured with one or more connections to the interactome database 106. In this way, the user can retrieve the target protein of interest from the content of the interactome database 106. After selecting one or more target targets, the processor of the user interface device 101 is configured to send the user selection to the processing platform 108. Here, the processing platform 108 is configured by the user input module 202 to receive the user's protein (target) selection. For example, the user input module 202 configures one or more processors of the processing platform 108 to receive a data file, object, stream, or link generated by the user input device 101.

[0054] like Figure 3 As shown in step 302, the user's selection is passed to the processing platform 108 for further processing. In one embodiment, the processing platform 108 transmits the user's selection directly to the interactome database 106 to identify one or more partner proteins. In a specific embodiment, the processing platform 108 changes or otherwise processes the user's selection before providing the data to the interactome database 106.

[0055] As described and referenced Figure 4 There are multiple protein interaction databases. However, not all protein databases are suitable for the systems and methods described herein.

[0056] To create a comprehensive dataset of computationally derived putative peptides, one approach is to apply PeptiDerive to all high-resolution The PDB co-crystal complexes of the peptides were used to generate a database (e.g., the interactome database) with a total of approximately 100 million peptide-target pairs with associated binding interface scores. This database represents a comprehensive collection of peptide-protein pairs and can therefore be used as a standardized training set for interface modeling.

[0057] It will be further appreciated that, in one approach, the percentage of the human proteome with at least one binding partner can be estimated by screening three databases: IMEx (https: / / www.imexconsortium.org / ), BioGRID (https: / / thebiogrid.org / ), or PROPER (https: / / genemo.ucsd.edu / proper / ) 6–8 .

[0058] In one embodiment, only databases that clearly provide experimental evidence of physical binding are used to create the interactome database 106. For example, in one embodiment, StringDB, another protein interaction database that does not guarantee physical binding, is excluded. However, in one or more further embodiments of the described systems, methods, and approaches, any source of protein interaction data can be included.

[0059] In one particular approach, the gene symbols corresponding to each human protein in the dataset were obtained. For example, gene symbols were downloaded from UniProt (a total of 20,601). For each database, symbols were scanned using pandas and a list of proteins involved in at least one PPI was compiled. Screening was performed separately for heterologous interactions and homologous (self-binding) interactions. To account for different curation criteria, in one approach, the entire process was repeated twice using different filters. The most stringent or least inclusive filter included PROPER entries with p < 0.01, all IMEx entries, and BioGRID entries supported by low-throughput (LTP) physical evidence. The least stringent or most inclusive filter included PROPER entries with p < 0.05, all IMEx entries, and BioGRID entries supported by LTP or high-throughput (HTP) physical evidence.

[0060] To quantify the availability of structural data for PPIs, we scanned the Protein Data Bank (PDB) for co-crystal complexes of two human proteins. Complexes are divided into two categories: heteromeric and homomeric. The PDB provides an entry ID for each co-crystal and a FASTA sequence for its two component proteins.

[0061] Because species designations and constituent entry IDs are not directly available, determining the composition of co-crystals can involve a multi-step process: (i) mapping co-crystal entry IDs to organisms and filtering only human-human interactions (reference: source.idx from the PDB archive, https: / / ftp.wwpdb.org / pub / pdb / derived_data / index / ) (ii) mapping the constituent proteins in each co-crystal to entry IDs based on FASTA sequences (reference: pdb_seqres.txt from the PDB archive, https: / / ftp.wwpdb.org / pub / pdb / derived_data / ) (iii) mapping entry IDs to UniProtKB identifiers and UniProt gene symbols (reference: SIFTS database pdb_chain_uniprot.csv, https: / / www.ebi.ac.uk / pdbe / docs / sifts / quick.html, UniProtRetrieve / ID Mapping tool) (iv) comparing PDB-derived UniProt gene symbols to the complete human genome. The final result represents the total number of human proteins involved in at least one co-crystal in the PDB.

[0062] For example, in one embodiment, a PDB-derived dataset is generated by mining verified high-resolution PPI structures in the RCSB PDB. For example, the process of extracting useful data first retrieves every interaction of every assembly of every co-crystal from the PDB. Then, the interactions are filtered to ensure their uniqueness (a unique interaction is one with a unique pair of interaction partners, or for the same pair of interaction partners, the buried surface area is significantly different). ). In one specific example, this filtering step generated 420,000 PPIs. Next, all interaction structures with amino acid sequence lengths greater than 50 and less than 1023 (to increase computational training speed) were processed using Rosetta PeptiDerive. However, it will be appreciated that amino acid sequences with lengths between 11 and 35,000 (TAL to Titin) can also be used when creating the necessary dataset. Next, a list of derived peptides and their corresponding Rosetta energy scores (REUs) are extracted. In this embodiment, the lower the score, the higher the predicted stability. Next, entries with an REU lower than -1000 are filtered out. The REU scores of the 10 peptide segments at each position are then averaged to estimate the energy score for each amino acid position. Next, the energy scores for each position in the matched derived protein sequence are averaged so that the dataset does not contain redundant entries. A threshold for the energy score is then set, for example -1. A binary classification task is established, in which amino acids with energies less than -1 are protein-binding amino acids and amino acids with energies greater than -1 are non-binding amino acids.

[0063] Transfer back to Figure 3 , as further shown in step 304, the interactome database 106 is searched using the user input generated in step 302. In one approach, the processing platform 108 is configured by the peptide search module 204 to access the interactome database 106 and search the interactome database using one or more search criteria according to the user's selection, as shown in step 304. For example, the peptide search module 204 causes the processing platform 108 to access an interactome database, such as the database 106. Here, the peptide search module 204 further configures the processing platform to submit the target protein obtained in step 302 to a search program or algorithm.

[0064] As shown in step 306, the search results performed in step 304 are fed back to the processing platform 108, as shown in step 306. For example, the search results module 206 configures the processing platform 108 to receive the search results from the interactome database 106. In one embodiment, the search results module 206 configures one or more processors of the processing platform 108 to generate a file or data object containing the search results of the interactome database. In another embodiment, the search results module 206 configures one or more processors of the processing platform 108 to convert the returned mate sequences into a FASTA format, as shown in step 306.

[0065] Next, as shown in step 308, the search results are provided to the prediction model. In one embodiment, the processing platform 108 evaluates the FASTA-converted form of the search results. However, in another embodiment, the processing platform 108 directly evaluates the direct results of the search in step 306. For example, the peptide generation module causes the processing platform 108 to access a computational model that is configured to receive input data in the form of a FASTA sequence and generate output values in the form of a nucleotide sequence, an amino acid sequence, or a sequence.

[0066] In one or more embodiments, the developed computational model is configured to generate putative partner proteins for a given target. The model utilizes data from the curated PPI database constructed as described herein.

[0067] For example, in one or more embodiments, the computational model implemented or accessed by the peptide generation module 208 is a machine learning model that uses data obtained from a custom PPI database to generate binding partners for a target sequence. In one configuration, the machine learning model is a large language model. In another configuration, the machine learning model is a neural network. For example, a pre-trained neural network is used to evaluate the data in the PPI database and generate binding partners for the target protein. In another configuration, the machine learning model is a protein language model.

[0068] In one embodiment, the machine learning model is a protein language model with millions of parameters. For example, the machine learning model is the ESM-2 model with 650 million parameters. It is understood that the ESM-2 model is a protein language model from Meta AI. However, other protein language models are also applicable, which should be understood and recognized. For example, any protein language model that can characterize protein interactions without generating multiple sequence alignments (MSAs) can be used. Although the lack of MSA features (which is a costly derivation procedure) represents an advancement in the field of machine learning models, it should be understood that any machine learning model, including models using MSA features, can be used in the above-mentioned systems, methods, and approaches described herein.

[0069] Through a more detailed overview, those skilled in the art will understand that ESM-2 is a Transformer-based model that uses a self-attention mechanism to learn long-range dependencies between amino acids in protein sequences. The self-attention mechanism enables the model to learn how different parts of a protein sequence interact with each other, which is crucial for understanding protein structure and function. Therefore, any model that incorporates this or similar features is suitable for the methods described herein.

[0070] Furthermore, ESM-2 consists of a stack of encoder layers, each containing a self-attention layer and a feed-forward layer. The encoder layer is responsible for extracting latent features from the input protein sequence. The encoder output is fed into a decoder layer, which is responsible for generating an embedding for the input protein sequence. In one embodiment, the decoder layer is a transformer layer with a different architecture from the encoder layer. For example, ESM-2 is trained using a masked language modeling objective. Therefore, the ESM-2 model is trained to predict masked amino acids in the input protein sequence. This objective enables the model to learn the relationships between amino acids in a protein sequence and encode various biological information in its embeddings.

[0071] In one or more embodiments, the provided machine learning model can be fine-tuned on a dataset of labeled examples. For example, to fine-tune ESM-2 for protein structure prediction, the model can be trained on a protein sequence dataset stored in the Intergroup Database or other protein databases.

[0072] Here, our method further improves the ESM-2 model by fine-tuning its final three layers. These final layers are fine-tuned using data from protein interaction (PPI) datasets, interactome databases, or other protein databases.

[0073] The ESM-2 model is then paired with a classification head. In one specific configuration, the classification head is a neural network that receives the output of the fine-tuned ESM-2 model. In another configuration, the neural network is a perceptron classification head. For example, the classification head is a multilayer perceptron (MLP) classification head paired with the fine-tuned ESM-2 model. Here, the MLP is a neural network configured to perform classification tasks. The MLP consists of at least three layers of fully connected neurons with nonlinear activation functions. Here, the MLP takes the embeddings from the ESM-2 model as input and outputs a predicted class for each amino acid in the sequence. For example, the MLP is trained to classify the interaction position of each amino acid.

[0074] As a more detailed example, the last three layers of the ESM-2650M are fine-tuned together with a four-layer fully connected neural network classification head that processes each position output of ESM-2 to predict the probability of each position. Here, the model configuration is as follows: each protein is passed to the model with the binary class of each amino acid as the target of the cross-entropy loss: -(ylog(p)+(1-y)log(1-p)), where y is the binary class label (0 = non-binding, 1 = binding) and p is the predicted probability of the amino acid belonging to the binding site.

[0075] More specifically, the computational model is configured to learn how to predict the similarity between two different sequences. To do this, it is trained on a dataset of sequences that have been labeled as similar or dissimilar. The model first generates a representation for each sequence, called an embedding. An embedding is a numerical vector that represents information about a sequence. The model then calculates the cosine similarity between each pair of sequence embeddings. Cosine similarity measures the similarity between two vectors. The higher the cosine similarity, the more similar the two sequences are. The model is then trained to predict the cosine similarity between each pair of sequences in the dataset. The model's predictions are compared with the actual cosine similarities, and the model is updated to improve its predictions. The average of the cross-entropy loss across the rows and columns of the matrix is used to measure the model's performance. The cross-entropy loss measures the degree of difference between two probability distributions. The smaller the cross-entropy loss, the better the model's performance.

[0076] In a specific embodiment, the ESM-2 model and its MLP head were implemented using PyTorch and trained until the validation loss began to increase. When using the PPBS dataset, to be consistent with ScanNet, the weighted approach used in the paper by Tubiana et al. was adopted, that is, the loss was multiplied by a specified weight.

[0077] It is understood that in one or more embodiments, the machine learning module accessed by the peptide generation module 208 is a pre-trained model. However, in one or more embodiments, the peptide generation module can train the computational model through one or more submodules before outputting the peptide sequence.

[0078] In a specific scheme, training, validation, and test sets were created with 26,423 training sequences, 3487 validation sequences, and 3817 test sequences, without entries belonging to the same cluster in different sets. This approach ensured that validation and test metrics did not reflect memorization of properties of homologous protein sequences. Proteins that were clustered by MMseqs as partner proteins selected in in vitro testing were also moved to the test set. For benchmarking, the Dockground-based PPBS dataset used in ScanNet was used.

[0079] The developed model, referred to herein as the Structure-Independent Language Transformer and Peptide Prioritizer (SaLT&PepPr) module, is configured to utilize the complete partner sequence. Consequently, the SaLT&PepPr model described herein outperforms the Cut&CLIP model in ranking peptides from the same protein on the test set (see Figure 6 ), its final Spearman correlation coefficient was 0.4, while the correlation coefficient of the Cut&CLIP model was less than 0.05. Therefore, the model represents an improvement in the field of protein-protein interaction modeling technology and other computational methods for evaluating proteins.

[0080] It should be understood and appreciated that the SaLT&PepPr inference model and module described herein are integrated in one or more embodiments with existing gold standard, experimentally validated PPI datasets including IMEx, BioGRID, and PROPER. Such datasets cover over 75% of all target proteins in the human proteome (compared to less than 25% covered by existing co-crystal datasets). The SaLT&PepPr module can be further extended in one or more embodiments by utilizing multimeric target proteins as its own peptide-derived scaffolds, e.g., see Figure 6 . It will be understood by those skilled in the art that a multimeric target protein is a protein composed of two or more subunits. These subunits can be the same or different. A peptide scaffold is a protein that can be used to display peptides on its surface. Therefore, the SaLT&PepPr module can be used to design peptides that can bind to the subunits of a multimeric target protein. This will make the binding affinity of the peptide to the target stronger than when it binds to only one subunit. In addition, the SaLT&PepPr module is configured to derive peptides from the amino acid sequence of the multimeric target. This approach can design peptides that specifically bind to multimeric targets.

[0081] As shown in step 310, the prediction model evaluates the search result data. For example, the processing platform 108 is configured by the peptide generation module 208, which provides an example of the prediction model described herein. Here, the processing platform 108 is configured by the peptide generation module 308 to generate peptide sequences that are predicted to bind to the target sequence identified by the user in step 202.

[0082] Once the peptide sequence is generated, it can be output as shown in step 312. For example, the output module 310 is configured to process one or more processors of the processing platform 108 to evaluate the generated sequence and select subsequences for further use or analysis. For example, the output module 310 is configured to evaluate the output of the prediction model and provide a probability of interaction with the target sequence for each amino acid.

[0083] In one embodiment, the output module 310 is configured to identify sub-portions of the generated sequence based on a likelihood threshold. For example, when the predicted interaction of a group of amino acids in the generated sequence is higher than a predetermined threshold (e.g., an interaction probability higher than 20, 30, 40, 50, 60, 70, 80, 90, 95), the output module selects these groups and outputs these groups or sub-sequences for further use and evaluation. For example, in one or more configurations, the target information and the predicted interacting sequences are stored in the results database 112 for further use and access. In addition, the computational model (SaLT&PepPr) is configured to generate peptides that can be sampled in the partner sequence to maximize the selection range and integrate prior knowledge of known binding domains. Specifically, the SaLT&PepPr module predicts the probability of each amino acid in the partner protein sequence being an interaction site.

[0084] As part of this prediction process, contiguous peptides are “cut” from the complete chaperone sequence to select the fragment with the highest average prediction score. In one approach, to sample different regions of the chaperone protein, a greedy sampling approach is used to obtain peptides of a user-specified length with the highest predicted binding probability, e.g. Figure 10 shown.

[0085] Understandably, because the SaLT&PepPr computational model is a purely sequence-based model, the total inference time for a single target protein is approximately one minute on a standard machine equipped with two CPU cores, 8GB of RAM, and no GPU. Therefore, the SaLT&PepPr module represents an unconventional and non-traditional advancement in the field of protein interaction modeling. For example, compared to SaLT&PepPr, a functionally equivalent approach using the publicly accessible ColabFold software coupled with the PeptiDerive step for modeling interacting sequence pairs requires over an hour of computational time and cannot be reliably applied to large multimeric complexes due to hardware limitations.

[0086] For example, Figure 8 Figure 3. Detailed comparison of the predicted SaLT&PepPr scores with experimentally annotated PPBS binding sites on different protein structures in the PPBS dataset. Red indicates amino acids with high binding probability, blue indicates amino acids with low binding probability, and the data are normalized for each protein chain.

[0087] same, Figure 9A representative example of the energy landscapes derived from the model and calculated by PeptiDerive for a specific PDB cocrystal entry is shown. Red indicates high binding probability, white indicates low binding probability, blue indicates low binding probability, and gray indicates amino acids that were discarded due to ineffectiveness for PeptiDerive. Note that the visualized PeptiDerive scores only reflect the binding sites captured in the specific PDB entry.

[0088] Overall, the computational model described in this paper exhibits robust model performance across a variety of target proteins, especially proteins with known binding domains. Figure 6 As shown, when trained and tested on a PDB-derived dataset, SaLT&PepPr achieved a test set area under the receiver operating characteristic (AUROC) curve of 0.77. Alternatively, when the ESM-2 weights were kept constant, the test set AUROC was 0.7, demonstrating the benefit of fine-tuning the final layers of the original model. This approach, utilizing the sequence of the binding partner, achieved a Spearman correlation of 0.4 with the PeptiDerive energy score on a test set with sequence identity <25%.

[0089] The computational model was also trained and tested on the ScanNet PPBS dataset to compare the described computational model with baseline and state-of-the-art models that require tertiary structure and / or multiple sequence alignments (MSAs) to identify protein-protein interacting residues. Figure 7 As shown in Figure 3, despite not using structure as input, the performance of the computational model is still competitive with the structure-based baseline, while the performance is slightly reduced compared to ScanNet. Specifically, in the "test none" partition reflecting the most distant proteins, SaLT&PepPr outperforms the baseline methods based on structural homology and manual feature selection, indicating that it has strong generalization ability to non-homologous proteins from different families. Finally, we visualized the SaLT&PepPr predictions for paired proteins with crystal structures in the PDB database, both in the independent structures of the single binding protein (such as Figure 8 ), or in a eutectic (as Figure 9 ), highlighting the model's ability to identify concise interaction interfaces. Experimental testing

[0090] Recently, building on the pioneering work of Portnoff et al., our group reprogrammed the specificity of a modular human E3 ubiquitin ligase called CHIP (carboxyl terminus of Hsc70-interacting protein) by replacing its natural substrate-binding domain, TPR, with designed "guide" peptides to generate minimal and programmable uAb architectures. To demonstrate that guide peptides from known, selected interacting partners could function as powerful guide peptides, we first focused on designing uAbs against β-catenin, as aberrant Wnt / β-catenin signaling has been widely implicated in multiple cancers, including colorectal, hepatocellular, lung, and pancreatic cancers.

[0091] Specifically, mutant β-catenin accumulates in the cytoplasm of diseased cells, while wild-type β-catenin binds to the transmembrane protein E-cadherin 20. Therefore, to target endogenous cytoplasmic β-catenin for degradation, we screened target peptides from the E-cadherin / β-catenin binding interface for subsequent uAb generation, leveraging its known sequence interaction with E-cadherin. These peptides were then scored using SaLT and PepPr. We then transfected the uAb constructs (SnP_1 to SnP_8) into DLD1 colon cancer cells, which express unusually high levels of wild-type β-catenin. Immunoblotting of cytoplasmic fractions revealed that all but one uAb significantly promoted β-catenin degradation compared to untransfected DLD1 control cells, with several constructs (SnP_3, SnP_5, and SnP_8) degrading greater than 60% of cytoplasmic β-catenin (Figure 11a). For example, a. Endogenous β-catenin degradation in the cytoplasmic fraction of DLD1 cells was assessed by immunoblotting using anti-β-catenin and anti-β-tubulin antibodies. Blots represent independent replicates (n = 3). Relative degradation activity was determined by densitometric analysis of anti-β-catenin immunoblots. b. β-catenin / TCF transcriptional activity was assessed using a TOPFlash luciferase reporter. The FOPFlash reporter served as a negative control. c. β-catenin binding activity was assessed by ELISA using immobilized β-catenin (β-cat). Binding to bovine serum albumin (BSA) served as a negative control. d. Total protein harvested from HEK293T cells cotransfected with plasmids encoding SnP_8uAb and β-catenin-sfGFP was analyzed by nano-liquid chromatography tandem mass spectrometry (Nano LC-MS / MS). Data were log2-normalized and processed for fold change and p-value (unpaired two-tailed t-test) to generate a volcano plot of differentially abundant proteins. STUB1 represents the CHIPΔTPR domain of the overexpressed SnP_8uAb. Data in panels (a-c) are mean ± SD of independent replicate transfections (n = 3). Statistical significance was determined by two-tailed Student's t-test for individual samples.

[0092] Using TOPFlash22 (a luciferase reporter that reliably reads β-catenin-dependent transcriptional activity), we observed that the potent SnP_8 degrader significantly reduced the β-catenin transcriptional response compared to empty vector control cells (Figure 11b). In contrast, the SnP_7 degrader exhibited less inhibitory effect on β-catenin signaling, consistent with its moderate degradation activity. We confirmed by quantitative ELISA (Figure 11c) that peptide-directed uAbs can promote target degradation through specific peptide-mediated β-catenin binding. Specifically, purified SnP_7 and SnP_8 uAbs exhibited strong affinity for immobilized β-catenin and little binding to the immobilized bovine serum albumin (BSA) control. The strong binding of SnP_7 and SnP_8 to β-catenin should be attributed to the SaLT&PepPr peptide, as evidenced by the lack of binding to the CHIPΔTPR ubiquitination domain alone. We note that the high binding activity of these uAbs to β-catenin is consistent with the binding affinities measured for other uAbs.1,23 Given the similar binding activity but varying degrees of silencing of β-catenin, other factors, such as proximity / orientation of binding, must also influence the efficacy of peptide-directed uAbs. Finally, to test the off-target propensity of peptide-directed uAbs, we performed one-dimensional liquid chromatography-tandem mass spectrometry (1D-LC-MS / MS) analysis of total protein harvested from cells overexpressing β-catenin, with or without treatment with the uAb candidates, quantifying approximately 6,700 proteins. Our analysis revealed the expected increase in the abundance of uAb-associated proteins, including a tryptic peptide assigned to the human CHIP protein (STUB1), with a corresponding decrease in β-catenin abundance between control and treated samples for both tested uAbs (Figure 11d). In contrast, the abundance of other proteins did not change significantly with uAb expression, confirming the absence of statistically significant off-target effects of uAb expression or degradation.

[0093] Experimental validation of SaLT&PepPr interface predictions for endogenous target degradation. Having identified interacting proteins as effective scaffolds for generating targeting peptides, we sought to test SaLT&PepPr's ability to screen for highly effective targeting peptides in a data-driven manner. To this end, we first selected eukaryotic initiation factor 4E binding protein 2 (4E-BP2), a relatively small and disordered protein involved in the initiation of eukaryotic translation and associated with cancer. 24,25 4E-BP2 has a known specific interactor: eukaryotic initiation factor 4E (eIF4E). 26 Using eIF4E as input for SaLT&PepPr, we obtained six high-scoring peptides from its sequence. These peptides were cloned into our uAb plasmid and transfected into A673 Ewing sarcoma cells, which overexpress 4E-BP2. Following Western blot analysis, we successfully identified two degraders: 4E-BP2_SnP_3 and 4E-BP2_SnP_6, which degraded endogenous 4E-BP2 by more than 50% compared to the non-targeted control plasmid (Figure 12a-b), validating the utility of our algorithm. Next, we turned our attention to TRIM8, which is itself an E3 ubiquitin ligase that regulates the levels of EWS-FLI1, the core fusion oncoprotein that drives Ewing sarcoma.11 Loss of TRIM8 induces EWS-FLI1-mediated overexpression in Ewing sarcoma cells, leading to upregulation of apoptosis. We used TRIM8 as input to our carefully curated PPI database to identify multiple interacting partners ( Figure 12A ) and used SaLT&PepPr to screen the top six peptides from different partners and incorporated them into our uAb architecture. Next, we transfected these uAbs into A673 Ewing's sarcoma cells and successfully identified two candidates, TRIM8_SnP_5 and TRIM8_SnP_6, which were able to degrade endogenous TRIM8 ( Figure 12C , Figure 12D ), which was statistically significant. We then co-transfected these six uAbs with a GFP-based apoptosis fluorescent caspase reporter (called ZipGFP27) in A673 cells and observed that the most potent degrader induced upregulation of apoptosis, as expected from previous studies ( Figure 12E ).

[0094] For further data, see "SaLT & PepPr is an interface-predicting language model for designing peptide-guided protein degraders," which is available at https: / / doi.org / 10.1038 / s42003-023-05464-z and is incorporated herein by reference as if presented in their respective complete forms. In addition, U.S. Provisional Application No. 63 / 344,820, entitled "Contrastive Learning for Peptide Based Degrader Design and Uses Thereof," is also incorporated herein by reference as if presented in its entirety. Treatment method

[0095] The present disclosure provides methods and compositions for creating engineered chimeras between synthetic binding proteins (e.g., antibodies, DARPins, FN3s, monomers, nanobodies, etc.) and post-translational modification (PTMs) domains that have extended half-lives within cells.

[0096] The present disclosure also provides a chimeric molecule, wherein the targeting domain is computationally designed.

[0097] The present disclosure also provides a chimeric molecule wherein the targeting domain is computationally designed and is relatively non-homologous (eg, non-natural sequence) to a wild-type binder to the target.

[0098] The present disclosure also provides a chimeric molecule wherein the PTM domain is designed by computer (eg, a computer-designed enzyme).

[0099] As used herein, the terms "chimeric molecule" or "ubiquitosome" are used interchangeably to refer to a molecule having a degradation domain and a targeting domain connected by a linker region, as defined herein.

[0100] As used herein, the term "deubiquitinating enzymes" or "DUBs" refers to enzymes that remove ubiquitin molecules from proteins in a process called deubiquitination. Ubiquitin is a small protein that is added to other proteins as a post-translational modification, a modification that affects protein function, localization, and stability. DUBs play an important role in regulating the ubiquitin system by reversing the effects of ubiquitination. There are many different types of DUBs, each with unique properties and functions. Some DUBs remove ubiquitin from single ubiquitination sites on proteins, while others cleave entire ubiquitin chains. DUBs are involved in a variety of cellular processes, including DNA repair, protein degradation, and immune responses. Dysregulation of DUB activity is associated with a variety of diseases, including cancer, neurodegenerative disorders, and inflammatory disorders. Humans have nearly 100 DUB genes, which can be divided into two major categories: cysteine proteases and metalloproteases. Cysteine proteases include ubiquitin-specific proteases (USPs), ubiquitin C-terminal hydrolases (UCHs), Machado-Josephin domain proteases (MJDs), and ovarian tumor proteases (OTUs). The metalloprotease group only includes Jab1 / Mov34 / Mpr1Pad1 N-terminal+ (MPN+) (JAMM) domain proteases.

[0101] Therefore, one aspect of the present disclosure relates to a chimeric molecule comprising: (i) a PTMs domain comprising a degradation domain composed of a deubiquitinating enzyme; and (ii) a targeting domain comprising a substrate binding motif heterologous to the deubiquitinating enzyme. A linker connects the PTMs domain to the targeting domain.

[0102] In some embodiments of the compositions and methods of the present disclosure, the chimeric molecule (or test agent) is an isolated chimeric molecule (or isolated test agent). As used herein, the term "isolated" or "purified" polypeptide, peptide, molecule or chimeric molecule is substantially free of cellular material or other contaminating polypeptides from a cell or tissue source, or substantially free of chemical precursors or other chemicals when chemically synthesized. For example, the chimeric molecule does not contain substances that interfere with the intended function, diagnosis or therapeutic use of the molecule. Such interfering substances may include proteins or fragments, enzymes, hormones, and other proteins and non-protein solutes other than those contained in the chimeric molecule.

[0103] In some embodiments of the compositions and methods of the present disclosure, the linker is heterologous to the PTMs domain and the targeting domain. According to such embodiments, the linker is heterologous to both the PTMs motif of the degradation domain and the substrate binding motif of the targeting domain.

[0104] As described herein, the substrate binding motif of the targeting domain is heterologous to the PTM domain. Thus, the PTM domain may be heterologous to the targeting domain. Likewise, in some embodiments, the PTM domain does not comprise a substrate binding motif.

[0105] In one or more embodiments, a peptide-based therapeutic is provided, wherein the therapeutic comprises any peptide-E3 ubiquitin ligase or other polynucleotide described herein, or a sequence having 80% homology thereto. In another embodiment, the peptide therapeutic comprises any of the aforementioned polynucleotides and is coupled to a delivery vehicle, wherein the delivery vehicle can be a virus or a micelle. The peptide-based therapeutic comprises a fusion of any of the aforementioned polynucleotides, wherein the peptide fusion is further fused to a cell penetrating motif or a cell surface receptor binding motif. In certain embodiments, the compositions and methods of the present disclosure can be used to prevent and / or treat symptoms of cancer and metastasis. In certain embodiments, the compositions and methods of the present disclosure can be used to prevent and / or treat cancer and metastasis.

[0106] In one embodiment, the subject has cancer and metastases. In some embodiments, the cancer or metastases are selected from the group consisting of basal cell carcinoma (BCC), head and neck squamous cell carcinoma (HNSCC), prostate cancer (CaP), pilomatricoma (PTR), and medulloblastoma (MDB).

[0107] Plasmid Construction All uAb plasmids were constructed from a standard pcDNA3 vector containing a cytomegalovirus (CMV) promoter and a C-terminal IRES-mCherry expression cassette. Target coding sequences (CDS) were synthesized as gBlocks by Integrated DNA Technologies (IDT). Sequences were amplified with overhangs for insertion into the pcDNA3-SARS-CoV-2-S-RBD-sfGFP backbone vector (Addgene #141184) via Gibson Assembly, which was double-digested with NheI and BamHI. Following PCR amplification using mutagenic primers (Genewiz), an Esp3I restriction site was introduced immediately upstream of the CHIPΔTPR CDS and the flexible GSGSG linker using KLD enzyme mix (NEB). For uAb assembly, candidate peptide oligonucleotides were annealed and ligated to the Esp3I-digested uAb backbone using T4 DNA ligase.

[0108] The assembled constructs were transformed into 50 μL of NEB Turbo Competent Escherichia coli cells and plated on LB agar supplemented with the corresponding antibiotics for subsequent colony sequence verification and plasmid purification (Genewiz). To purify the proteins, the gene encoding each uAb construct was PCR amplified from a pcDNA3-based plasmid using primers that introduced HindIII and XhoI overhangs. The resulting PCR amplicon was ligated into an empty pET28a vector that had been double-digested with HindIII and XhoI. This process generated plasmids encoding each selected peptide (with CHIPΔTPR and a 6×His tag at the C-terminus). All plasmids were verified by DNA sequencing at Genewiz or the Genomics Facility of the Biotechnology Resource Center (BRC) at Cornell Biotechnology and purified. Cell Culture: The DLD1 cell line was generously donated by Dr. Pengbo Zhou. DLD1 cells (ATCC CCL-221), HEK293T cells (ATCC CRL-3216), and A673 cells (ATCC CRL-1598) were cultured in DMEM supplemented with 100 units / ml penicillin, 100 mg / ml streptomycin, and 10% fetal bovine serum. Unless otherwise specified, 0.3 × 10⁶ cells were seeded per well of a 6-well plate the day before transfection. Plasmids expressing uAbs were prepared using the PureYield miniprep kit to remove endotoxins. On the day of transfection, plasmids were transfected using Lipofectamine 3000. After incubation for 3 days after transfection, cell lysates were collected for immunoblot analysis. Cell fractionation and immunoblotting for probing Figure 2To detect β-catenin in the culture medium, on the day of harvest, 0.05% trypsin-EDTA was added to detach the cells, and the cell pellets were washed twice with ice-cold 1× PBS. The cells were then lysed according to the manufacturer's instructions, and a subcellular protein fractionation kit (ThermoFisher) was used to lyse the cells and separate subcellular components from the lysate. Specifically, ice-cold cytoplasmic extraction buffer was added to the cell pellet, the mixture was placed at 4°C for 10 minutes with gentle shaking, and then centrifuged at 500×g for 10 minutes at 4°C. The supernatant was immediately collected into a pre-chilled PCR tube, placed on ice, and then subjected to immunoblotting or stored at -20°C for later use. Ice-cold membrane extraction buffer was then added to the pellet. The mixture was incubated at 4°C for 10 minutes and then centrifuged at 3000×g for 5 minutes. The supernatant was immediately transferred to a pre-chilled tube. Protein concentration was quantified using a Pierce BCA protein kit (ThermoFisher). Equal amounts of total protein were loaded onto Precise Tris-HEPES 4-20% sodium dodecyl sulfate (SDS)-polyacrylamide gels (ThermoFisher) and separated by electrophoresis. Immunoblotting analysis was performed according to standard protocols.

[0109] Briefly, proteins were transferred to poly(vinylidene fluoride) (PVDF) membranes (Millipore), blocked with 5% nonfat milk (Carnation) in 1× tris-buffered saline (TBS) containing 0.05% Tween 20 (TBST) for 1 hour at room temperature, washed three times with TBST for 10 minutes each, and probed with rabbit anti-β-catenin antibody (Cell Signaling, Cat#8480S; diluted 1:1000) or rabbit anti-β-tubulin (Cell Signaling Cat#2146; diluted 1:1000). The blots were washed again with TBST three times for 5 minutes each and then probed with donkey anti-rabbit horseradish peroxidase (HRP) secondary antibody (Abcam, Cat#7083; diluted 1:2500) for 1 hour at room temperature. Blots were detected by chemiluminescence using a ChemiDoc MP imager (Bio-Rad). Densitometry analysis of protein bands in immunoblots was performed using ImageJ software, see: https: / / imagej.nih.gov / ij / docs / examples / dot-blot / . Briefly, bands in each lane were grouped into a row or horizontal "lane" and quantified using the gel analysis function of ImageJ. uAb band intensity data were normalized to the empty vector control group from six independent experiments. Figure 3TRIM8 and 4E-BP2 in the culture medium were isolated on the day of harvest by adding 0.05% trypsin-EDTA to detach the cells and wash the cell pellets twice with ice-cold 1× PBS. The cells were then lysed and subcellular components were isolated from the lysate using a protease inhibitor cocktail (Millipore Sigma) diluted 1:100 in Pierce RIPA buffer (ThermoFisher). Specifically, protease inhibitor cocktail-RIPA buffer was added to the cell pellet, the mixture was placed at 4°C for 30 minutes, and then centrifuged at 15,000 rpm for 10 minutes at 4°C. The supernatant was immediately collected into a pre-cooled PCR tube and 4×Bolt 500 containing 5% β-mercaptoethanol was added at a ratio of 3:1. TM After adding LDS loading buffer (ThermoFisher), the mixture was incubated at 95°C for 10 minutes and then subjected to immunoblotting analysis. Immunoblotting analysis was performed according to the standard protocol. Briefly, equal volumes of sample were added to Bolt TM Bis-Tris Plus Mini Protein Gel (ThermoFisher) was then electrophoresed. TM 2Transfer Stacks (Invitrogen) were used for membrane blotting and SuperBlock TMAfter incubation in blocking buffer (ThermoFisher) for 1 hour at room temperature, proteins were probed with rabbit anti-TRIM8 antibody (Cell Signaling, Cat#4936, 1:500 dilution), rabbit anti-4EBP2 antibody (Cell Signaling, Cat#2845T, 1:500 dilution), rabbit anti-Vinculin antibody (ThermoFisher, Cat#700062, 1:500 dilution), or mouse anti-GAPDH (Santa Cruz Biotechnology, Cat#sc-47724; 1:500 dilution) and incubated overnight at 4°C. The blot was washed three times with 1× TBST for 5 minutes each and then incubated with secondary antibodies, goat anti-rabbit IgG (H+L), horseradish peroxidase (HRP) (see Poly-HRP COMMUNICATIONS BIOLOGY | https: / / doi.org / 10.1038 / s42003-023-05464-z ARTICLE COMMUNICATIONS BIOLOGY | (2023) 6:1081 | https: / / doi.org / 10.1038 / s42003-023-05464-z | www.nature.com / commsbio 7) (ThermoFisher, Cat# 32230, 1:2000 dilution) at room temperature for 1-2 hours. After washing three times with 1× TBST for 5 minutes each, the blot was detected by chemiluminescence using the iBright 1500 Imaging System (ThermoFisher). Densitometry analysis of protein bands in immunoblots was performed using FIJI software (see: https: / / imagej.nih.gov / ij / docs / examples / dotblot / ). Briefly, bands in each lane were grouped into a row or horizontal "lane" and quantified using the gel analysis function of FIJI. The intensity data of the uAb bands were first normalized to the band intensity of GAPDH (for TRIM8) or vinculin (for 4E-BP2) in each lane and then normalized to the average band intensity of the empty uAb vector control group in replicate experiments. TOPFlash Assay: 20-24 hours before transfection, 1×10⁴ DLD1 cells were seeded in a white-bottomed 96-well plate. On the day of transfection, the following plasmids were added to each well: M50 Super 8×TOPFlash plasmid (Addgene #12456) or M51 Super 8×FOPFlash (TOPFlash mutant; Addgene #12457), pCMV-Renilla29, and pcDNA3-SnP_7 or pcDNA3-SnP_8. 100 ng of plasmid DNA was mixed with Lipofectamine 3000 reagent in serum-free Opti-MEM medium at a ratio of 1:0.1:3 for TOPFlash / FOPFlash:Renilla:SnP_7 / SnP_8 uAb. After incubation at room temperature for 15 minutes, the cells were added dropwise to each well. After 48 hours of incubation, the cells were lysed, and the luminescence signals of firefly and Renilla were measured using a dual-luciferase reporter kit (Promega). The plates were read using a microplate reader (Tecan). Luciferase activity was measured and normalized to that of the control Renilla. Protein expression and purification All purified uAb constructs and unfused CHIPΔTPR were obtained from BL21(DE3) E. coli cultures carrying the pET28a plasmid encoding the corresponding SnP_7 and SnP_8 uAbs or CHIPΔTPR3.

[0110] Cells were cultured in Luria-Bertani (LB) medium according to the previously described protocol 3. Briefly, when the culture density (calculated as optical density at 600 nm (OD600)) reached 0.5-0.7, protein expression was induced with 1 M isopropyl β-D-1-thiogalactopyranoside (IPTG) and cultured at 37°C for 12-16 hours. After expression, cells were harvested by centrifugation at 10,000 × g for 10 minutes at 4°C. The resulting pellet was resuspended in 10 mL of phosphate-buffered saline (PBS) and lysed using an EmulsiFlex-C5 high-pressure homogenizer (Avestin). Insoluble material was removed from the lysate by centrifugation at 10,000 × g for 10 minutes at 4°C. The clarified lysate containing 6xHis-tagged protein was subjected to gravity flow Ni2+ affinity purification using HisPur-Ni-NTA resin (ThermoFisher) according to the manufacturer's protocol. The purified protein can be stored at 4°C for 2 weeks. The final purity of all proteins was confirmed by Coomassie blue staining of SDS-PAGE gels. ELISA ELISA assays were performed according to previously published protocols. Briefly, 50 μL of 1 μg / mL β-catenin (Biomatik, Cat# RPU40704) diluted in PBS (pH 7.4) was added to each well of a 96-well plate (MaxiSorp; Nunc Nalgene) and incubated overnight at 4°C. The plate was incubated overnight at 4°C with 200 μL of blocking buffer (5% (w / v) skim milk (Carnation) in PBS) and then washed three times with 200 μL of PBS-T (0.1% (v / v) Tween 20 in PBS) per well. The plate was then incubated with EZ-Link ELISA kit according to the manufacturer's instructions. TM Purified uAb constructs were biotinylated with NHS-Biotin (ThermoFisher, Cat#20217). The biotinylated uAb constructs were serially diluted three times in PBS and added to the ELISA plate. The plate was incubated at 37°C for 1 hour. The plate was washed three times with PBS-T and then incubated in the presence of HRP-conjugated streptavidin (ThermoFisher, Cat#N100; diluted 1:20,000) at room temperature for 1 hour with shaking at 450 rpm. After washing three more times with PBS-T, 100 μL of 3,3'-5,5'-tetramethylbenzidine substrate (1-Step UltraTMB-ELISA; ThermoFisher) was added to each well and incubated at room temperature in the dark. The reaction was terminated by adding 100 μL of 2 M H2SO4, and the absorbance was measured at 450 nm using a FilterMaxF5 microplate spectrophotometer (Agilent). Proteomics HEK293T cells were cultured in DMEM supplemented with 100 U / mL penicillin, 100 mg / mL streptomycin, and 10% FBS. 1 μg Target-sfGFP plasmid and 1 μg Target-sfGFP+1 μg pcDNA-uAb plasmid were transfected into the cells in triplicate (8×104 cells per well / 6-well plate) using Lipofectamine 3000 (Invitrogen) in Opti-MEM (Gibco). Three days after transfection, cells were harvested and washed four times with 500 μL 1X cold PBS. The cell pellet was resuspended in 200 μL Pierce RIPA buffer (VWR) and incubated on ice for 15 minutes. The homogenate was treated with triethylammonium bicarbonate buffer containing 20% (w / v) SDS at pH 8.5, then sonicated with a probe and heated at 80°C for 5 minutes. After centrifugation, the supernatant was collected and concentrations were determined using a detergent-compatible Bradford assay. A 20 μg aliquot of each sample was reduced and alkylated, followed by trypsin digestion using an S-trap microdevice. The peptide eluate was freeze-dried, reconstituted, and equal volumes of each sample were combined to create an SPQC pool. Approximately 1 μg of each sample, as well as triplicate samples of the SPQC pool, were analyzed by 1D-LCMS / MS. Samples were analyzed using a Nanospray Flex ion source coupled to an M-class UPLC system (Waters) connected to an Exploris 480 high-resolution, precise tandem mass spectrometer (ThermoFisher) and processed using Spectronaut 16. P values were calculated by Student's t-test on log2FC values. Log2FC values were calculated as the difference in mean protein abundance in the presence and absence of the uAb. Functional assay To perform apoptosis experiments, 3×105 A673 cells / well were seeded in 24-well plates 20-24 hours before transfection. On the day of transfection, the following plasmids were added to each well: ZipGFP-Casp3 plasmid (Addgene#81241) and pcDNA3-SnP_TRIM8_#. 500ng of plasmid DNA was mixed with Lipofectamine 2000 reagent in serum-free Opti-MEM medium at a ratio of ZipGFP-Casp3:pcDNA3-SnP_TRIM8_#=1:1 and incubated at room temperature for 20 minutes before adding dropwise to each well. After incubation for 60 hours, cells were collected and analyzed using a method similar to uAb screening. Cells expressing mCherry were gated and analyzed using FlowJo software ( https: / / flowjo.com / ) Normalized EGFP cell fluorescence was calculated and compared with samples transfected with non-targeting uAb.

[0111] To ensure reliable reproducibility of all results, experiments were performed with at least three biological replicates and three technical measurements. The sample size was not predetermined based on statistical methods but was selected based on field standards (at least three independent biological replicates per condition) and was sufficient to detect the target effect size. All data are presented as mean ± standard deviation (SD). For individual samples, statistical significance was determined by paired Student's t-test unless otherwise stated (*p < 0.05, **p < 0.01; ***p < 0.001; ****p < 0.0001). All figures and tables were generated using Prism 9 for macOS, version 9.2.0. No data were excluded from the analyses. The experiments were not randomized. The researchers were aware of the group assignments during the implementation and evaluation of the experiments. Pharmaceutical composition

[0112] Therefore, the present disclosure provides pharmaceutical compositions comprising peptide-post-translationally modified fusion compounds and pharmaceutically acceptable carriers. The compounds of the present disclosure can be formulated into pharmaceutical compositions and administered to mammalian hosts, such as human patients, in a variety of forms suitable for the selected route of administration.

[0113] Routes of administration include, but are not limited to, oral, topical, mucosal, nasal, parenteral, gastrointestinal, intraspinal, intraperitoneal, intramuscular, intravenous, intrauterine, intraocular, intradermal, intracranial, intratracheal, intravaginal, intracerebroventricular, intracerebral, subcutaneous, ocular, transdermal, rectal, buccal, epidural, and sublingual.

[0114] As used herein, the term "administration" refers broadly to any and all ways of introducing a compound described herein into a host subject. The compounds described herein can be administered in unit dosage form and / or in the form of compositions comprising one or more pharmaceutically acceptable carriers, adjuvants, diluents, excipients and / or vehicles, and combinations thereof.

[0115] As used herein, the term "composition" generally refers to any product comprising more than one ingredient, including the compounds described herein. It should be understood that the compositions described herein can be prepared from the compounds described herein or salts, solutions, hydrates, solvates, and other forms of the compounds described herein. It should be understood that these compositions can be prepared from various amorphous, non-amorphous, partially crystalline, crystalline, and / or other forms of the compounds described herein, and that these compositions can be prepared from various hydrates and / or solvates of the compounds described herein. Therefore, such pharmaceutical compositions reciting the compounds described herein include each, any combination, or individual forms of the various forms and / or solvates or hydrates of the compounds described herein.

[0116] In some embodiments, the peptide-post-translational modification fusion-based therapeutic can be systemically (e.g., orally) administered in combination with a pharmaceutically acceptable carrier (e.g., an inert diluent or an assimilable edible carrier). For oral therapeutic administration, the active compound can be combined with one or more excipients and used in the form of swallowable tablets, buccal tablets, sublingual tablets, lozenges, capsules, elixirs, suspensions, syrups, wafers, etc. The percentage of active ingredient and excipients in the compositions and formulations can vary, and the percentage of active ingredient and excipients can be from about 1% to about 99%. Excipients include, but are not limited to, binders, fillers, diluents, disintegrants, lubricants, surfactants, sweeteners, flavorings, colorants, buffers, antioxidants, preservatives, chelating agents (e.g., ethylenediaminetetraacetic acid) and agents for regulating rehydration (e.g., sodium chloride), ethylenediaminetetraacetic acid) and agents for regulating rehydration, such as sodium chloride.

[0117] Suitable binders include, but are not limited to, polyvinylpyrrolidone, copovidone, hydroxypropylmethylcellulose, starch, and gelatin.

[0118] Suitable fillers include, but are not limited to, sugars (such as lactose, sucrose, mannitol, or sorbitol) and their derivatives (such as amino sugars), ethylcellulose, microcrystalline cellulose, and silicified microcrystalline cellulose.

[0119] Suitable diluents include, but are not limited to, dicalcium phosphate dihydrate, sugar, lactose, calcium phosphate, cellulose, kaolin, mannitol, sodium chloride, and dry starch.

[0120] Suitable disintegrants include, but are not limited to, pregelatinized starch, crospovidone, croscarmellose sodium, and combinations thereof.

[0121] Suitable lubricants include, but are not limited to, sodium stearyl fumarate, stearic acid, polyethylene glycol, or stearates such as magnesium stearate.

[0122] Suitable surfactants or emulsifiers include, but are not limited to, polyvinyl alcohol (PVA), polysorbates, polyethylene glycol, polyoxyethylene-polyoxypropylene block copolymers known as "poloxamers," polyglycerol fatty acid esters (such as decaglyceryl monolaurate and decaglyceryl monomyristate), sorbitan fatty acid esters (such as sorbitan monostearate), polyoxyethylene sorbitan fatty acid esters (such as polyoxyethylene sorbitan monooleate (Tween)), polyethylene glycol fatty acid esters (such as polyoxyethylene monostearate), polyoxyethylene alkyl ethers (such as polyoxyethylene lauryl ether), polyoxyethylene castor oil, and hardened castor oil (such as polyoxyethylene hardened castor oil).

[0123] Suitable flavoring agents and sweeteners include, but are not limited to, sweeteners (e.g., sucralose), synthetic flavor oils and flavoring fragrances, natural oils, plant extracts (e.g., leaf, flower, and fruit extracts), and combinations thereof. Exemplary flavoring agents include cinnamon oil, wintergreen oil, peppermint oil, clover oil, hay oil, anise oil, eucalyptus oil, vanilla oil, citrus oils (e.g., lemon oil, orange oil), grape oil, and grapefruit oil, as well as fruit essences (e.g., apple, peach, pear, strawberry, raspberry, cherry, plum, pineapple, and apricot).

[0124] Suitable coloring agents include, but are not limited to, alumina (dried aluminum hydroxide), annatto extract, calcium carbonate, canthaxanthin, caramel, beta-carotene, cochineal extract, carmine, potassium sodium copper chlorophyllin (chlorophyll-copper complex), dihydroxyacetone, bismuth oxychloride, synthetic iron oxide, ammonium ferric ferrocyanide, ferric ferrocyanide, chromium hydroxide green, chromium oxide green, guanine, mica-based pearlescent pigments, pyrophyllite, mica, tooth powder, talc, titanium dioxide, aluminum powder, bronze powder, copper powder, and zinc oxide.

[0125] Suitable buffers or pH adjusters include, but are not limited to, acidic buffers (e.g., short-chain fatty acids, citric acid, acetic acid, hydrochloric acid, sulfuric acid, and fumaric acid); and alkaline buffers (e.g., tris, sodium carbonate, sodium bicarbonate, sodium hydroxide, potassium hydroxide, and magnesium hydroxide).

[0126] Suitable osmotic pressure enhancing agents include, but are not limited to, ionic and non-ionic agents such as alkali or alkaline earth metal halides, urea, glycerol, sorbitol, mannitol, propylene glycol, and glucose.

[0127] Suitable wetting agents include, but are not limited to, glycerin, cetyl alcohol, and glyceryl monostearate.

[0128] Suitable preservatives include, but are not limited to, benzalkonium chloride, benzophenone chloride, thimerosal, phenylmercuric nitrate, phenylmercuric acetate, phenylmercuric borate, methylparaben, propylparaben, chlorobutanol, benzyl alcohol, benzyl alcohol, chlorhexidine, and polyhexamethylene biguanide.

[0129] Suitable antioxidants include, but are not limited to, sorbic acid, ascorbic acid, ascorbate, glycine, alpha-tocopherol, butylated hydroxyanisole (BHA), and butylated hydroxytoluene (BHT).

[0130] The peptide-post-translational modification fusion therapy of the present disclosure can also be administered via infusion or injection (e.g., using a needle (including microneedle) syringe and / or a needle-free syringe). The solution of the active composition can be aqueous, optionally mixed with a non-toxic surfactant, and / or can contain a carrier or excipient, such as a salt, a carbohydrate, and a buffer (preferably a pH of 3 to 9). For certain applications, they may be more suitable for formulation into a sterile non-aqueous solution or dry form for use with a suitable carrier (e.g., sterile, pyrogen-free water or phosphate-buffered saline). For example, dispersions can be prepared in glycerol, liquid polyethylene glycol, triacetin, and mixtures thereof, as well as oils. The formulation may further comprise a preservative to prevent the growth of microorganisms.

[0131] Pharmaceutical compositions can be formulated for parenteral administration (e.g., subcutaneous injection, intravenous injection, arterial injection, transdermal injection, intraperitoneal injection, or intramuscular injection) and can include aqueous and non-aqueous isotonic sterile injection solutions, which may contain antioxidants, buffers, antibacterial agents, and solutes that make the formulation isotonic with the blood of the intended recipient, as well as aqueous and non-aqueous sterile suspensions, which may contain suspending agents, solubilizers, thickeners, stabilizers, and preservatives. When the pharmaceutical composition is administered intravenously, water is a preferred carrier. Saline solutions and aqueous solutions of glucose and glycerol can also be used as liquid carriers, particularly for injections. Oils (e.g., petroleum, animal, vegetable, or synthetic oils) and soaps (e.g., fatty alkali metal salts, ammonium salts, and triethanolamine salts) and suitable detergents can also be used for parenteral administration. In addition, the composition may contain one or more nonionic surfactants. Suitable surfactants include polyethylene sorbitol fatty acid esters, such as sorbitol monooleate, and high molecular weight adducts of ethylene oxide and a hydrophobic group (formed by the condensation of propylene oxide and propylene glycol). Suitable preservatives include sodium benzoate, benzoic acid and sorbic acid. Suitable antioxidants include sulfites, ascorbic acid and alpha-tocopherol.

[0132] Parenteral compounds / compositions may readily be prepared under sterile conditions using standard pharmaceutical techniques well known to those skilled in the art, for example by lyophilization.

[0133] Compositions for inhalation or isolation include solutions and suspensions in pharmaceutically acceptable aqueous or organic solvents or mixtures thereof, and powders. Liquid or solid compositions may contain suitable pharmaceutically acceptable excipients as described above. In one embodiment, the composition is administered via oral or nasal respiratory route to achieve local or systemic effects. Compositions in pharmaceutically acceptable solvents can be atomized using an inert gas. The atomized solution can be inhaled directly from the atomizing device, or the atomizing device can be connected to a mask or intermittent positive pressure breathing machine. Solutions, suspensions or powder compositions can be administered orally or nasally by a device that delivers the preparation in an appropriate manner.

[0134] In another embodiment, the composition is prepared for topical administration, for example, as an ointment, gel, drops, or cream. For topical application to the body surface using, for example, a cream, gel, drops, ointment, etc., the compounds of the present disclosure can be prepared and applied in a physiologically acceptable diluent, with or without a pharmaceutical carrier. Excipients for topical administration or gel matrices can include, for example, sodium carboxymethylcellulose, polyacrylates, polyoxyethylene-polyoxypropylene block polymers, polyethylene glycol, and wood wax alcohols.

[0135] Alternative formulations include nasal sprays, liposomal formulations, sustained release formulations, pumps (including mechanical or osmotic pumps) to deliver the drug into the body, controlled release formulations, and the like, as are known in the art. dose

[0136] As used herein, the term "therapeutically effective dose" refers, unless expressly stated otherwise, to that amount of a compound that, when administered at one time or over the course of a treatment cycle, affects the health, well-being, or mortality of a subject.

[0137] The peptide-post-translational modification fusion-based therapeutics described herein can be present in the composition in the following amounts: about 0.001 mg, about 0.005 mg, about 0.01 mg, about 0.02 mg, about 0.03 mg, about 0.04 mg, about 0.05 mg, about 0.06 mg, about 0.07 mg, about 0.08 mg, about 0.09 mg, about 0.1 mg, about 0.2 mg, about 0.3 mg, about 0.4 mg, about 0.5 mg, about 0.6 mg, about 0.7 mg, about 0.8 mg, about 0.9 mg, about 1 mg, about 1.5 mg, about 2 mg, about 2.5 mg, about 3 mg, about 3.5 mg, about 4 mg, about 4.5 mg, about 5 mg, about 5.5 mg, about 6 mg, about 6.5 mg, about 7 mg, about 8 mg, about 9 mg, about 10 mg, about 11 mg, about 12 mg, about 13 mg, about 14 mg, about 15 mg, about 16 mg, about 17 mg, about 18 mg, about 19 mg, about 20 mg, about 21 mg, about 22 mg, about 23 mg, about 24 mg, about 25 mg, about 26 mg, about 27 mg, about 28 mg, about 29 mg, about 30 mg .5 mg, about 8 mg, about 8.5 mg, about 9 mg, about 0.5 mg, about 10 mg, about 10.5 mg, about 11 mg, about 12 mg, about 12.5 mg, about 13 mg, about 13.5 mg, about 14 mg, about 14.5 mg, about 15 mg, about 15.5 mg, about 16 mg, about 16.5 mg, about 17 mg, about 17.5 mg, about 18 mg, about 18.5 mg, about 19 mg, about 19.5 mg, about 20 mg, about 25 mg, about 30 mg, about 35 mg, about 40 mg, about 45 mg, about 50 mg, about 55 mg, about 60 mg, about 65 mg, about 70 mg, about 75 mg, about 80 mg, about 85 mg, about 90 mg, about 95 mg, about 100 mg.

[0138] The peptide-post-translational modification fusion-based therapeutics described herein may be present in the composition in the following ranges: about 0.1 mg to about 100 mg; 0.1 mg to about 75 mg; about 0.1 mg to about 50 mg; about 0.1 mg to about 25 mg; about 0.1 mg to about 10 mg; 0.1 mg to about 7.5 mg, 0.1 mg to about 5 mg; 0.1 mg to about 2.5 mg; about 0.1 mg to about 1 mg; about 0.5 mg to about 100 mg; about 0.5 mg to about 75 mg; about 0.5 mg to about 50 mg; about 0.5 mg to about 25 mg; about 0.5 mg to about 10 mg; about 0.5 mg to about 5 mg, about 0.5 mg to about 2.5 mg; about 0.5 mg to about 1 mg; about 1 mg to about 100 mg; about 1 mg to about 75 mg; about 0.1 mg to about 50 mg; about 0.1 mg to about 25 mg; about 0.1 mg to about 10 mg; about 0.1 mg to about 5 mg; about 0.1 mg to about 2.5 mg; about 0.1 mg to about 1 mg. Dosage regimen

[0139] The compounds described herein can be administered by any dosing regimen or dosing schedule depending on the patient and / or condition being treated. Administration can be once a day (qd), twice a day (bid), three times a day (tid), once a week, twice a week, three times a week, once every two weeks, once every three weeks, or twice a month, etc.

[0140] In some embodiments, the treatment based on the peptide-post-translational modification fusion is administered for at least one day. In other embodiments, the treatment based on the peptide-post-translational modification fusion is administered for at least two days. In other embodiments, the treatment based on the peptide-post-translational modification fusion is administered for at least three days. In other embodiments, the treatment based on the peptide-post-translational modification fusion is administered for at least four days. In other embodiments, the treatment based on the peptide-post-translational modification fusion is administered for at least five days. In other embodiments, the treatment based on the peptide-post-translational modification fusion is administered for at least six days. In other embodiments, the treatment based on the peptide-post-translational modification fusion is administered for at least seven days. In other embodiments, the treatment based on the peptide-post-translational modification fusion is administered for at least ten days. In other embodiments, the treatment based on the peptide-post-translational modification fusion is administered for at least 14 days. In other embodiments, the treatment based on the peptide-post-translational modification fusion is administered for at least one month. In some embodiments, the treatment based on the peptide-post-translational modification fusion is administered long-term, until treatment is needed.

[0141] The subject matter described herein will be more particularly illustrated by the following non-limiting examples. It should be understood that changes and modifications may be made thereto without departing from the scope and spirit of the present disclosure. It should also be understood that various theories as to why the present disclosure is effective are not intended to limit the present disclosure.

[0142] The description of the foregoing specific embodiments will fully reveal the general nature of the present disclosure so that others can use the knowledge within the scope of the relevant art technology (including the literature content cited and incorporated by reference herein) to easily modify and / or adjust these specific embodiments for various applications without departing from the overall concept of the present disclosure without excessive experimentation. Therefore, such adjustments and modifications are intended to be based on the teachings and guidance provided herein, within the meaning and scope of the equivalents of the disclosed embodiments. It should be understood that the wording or terminology herein is for descriptive purposes only and not for limiting purposes, and therefore, those skilled in the art should interpret the terms or wording in this specification according to the teachings and guidance provided herein, in combination with the knowledge of those skilled in the relevant art.

[0143] The terms used herein are used only to describe specific embodiments and are not intended to limit the present invention. As used herein, the singular forms "a", "an" and "the" also include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "include" and / or "comprise" used in this specification specify the presence of the features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.

[0144] Although various embodiments of the present disclosure have been described above, it should be understood that these embodiments are intended to be illustrative and non-limiting. Those skilled in the art will appreciate that various modifications in form and detail may be made to the present disclosure without departing from the spirit and scope of the present disclosure. Therefore, the present disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents.

[0145] For example, it can be understood as the following implementation.

[0146] Example 1. A method for generating a peptide sequence that binds to a target sequence, the method comprising: a. receiving a data object corresponding to a target protein using a processor configured by a code executed therein; b. using the data object to search for at least one partner protein of the target protein in a protein interaction database; c. recognizing at least one partner protein of the target protein; d. providing at least one chaperone protein to a computational model configured to output a predicted protein sequence predicted to interact with the target sequence; and

[0147] Identifying at least one subsequence in the predicted protein sequence that satisfies a predetermined interaction threshold. A method as described in any of the above, further comprising: a. Converting the identified at least one chaperone protein into a FASTA data format before providing the at least one chaperone protein to the computational model.

[0148] The method of any preceding implementation, wherein the computational model is a protein language model.

[0149] The method of any preceding implementation, wherein the protein language model is configured to generate a predicted protein sequence without using structural data associated with the target protein.

[0150] The method of any preceding implementation, wherein the computational model is a protein language model of at least 500 million parameters.

[0151] The method of any preceding implementation, wherein the computational model further comprises a multilayer perceptron classification head configured to receive an output of the protein language model.

[0152] The method of any preceding implementation, wherein identifying at least one subsequence comprises receiving a user-specified length of the target binding generated sequence and generating subsequences of the predicted protein sequence based on a highest predicted probability of binding to the target using greedy sampling.

[0153] An implementation wherein a pre-trained protein language model, when executed by a processor, configures the processor to generate one or more guide peptide sequences without using structural data.

[0154] An implementation of a method for training a machine learning model to predict cosine similarity between two protein sequences, comprising the following steps: generating a matrix of all possible target sequence and peptide sequence pairs; generating an embedding for each sequence in the matrix using a protein language model; calculating the cosine similarity between each pair of embeddings in the matrix; calculating the cross-entropy loss between the predicted cosine similarity and the actual cosine similarity; calculating the average cross-entropy loss of the matrix; and updating the model parameters to minimize the average cross-entropy loss.

[0155] The method of any preceding implementation, further comprising generating a training dataset consisting of experimentally validated protein-protein interactions and training the machine learning model.

Claims

1. A method for generating a peptide sequence that binds to a target sequence, the method comprising: a. receiving a data object corresponding to a target protein using a processor configured by the code executed therein; b. using the data object to search for at least one partner protein of the target protein in a protein interaction database; c. recognizing at least one partner protein of the target protein; d. providing the at least one chaperone protein to a computational model configured to output a predicted protein sequence predicted to interact with the target sequence; and e. Identifying at least one subsequence in the predicted protein sequence that meets a predetermined interaction threshold.

2. The method according to claim 1, further comprising: a. Converting the identified at least one chaperone protein into a FASTA data format before providing the at least one chaperone protein to the computational model. The method according to claim 1 , wherein the computational model is a protein language model.

4. The method of claim 3, wherein the protein language model is configured to generate a predicted protein sequence without using structural data associated with the target protein.

5. The method of claim 3, wherein the computational model is a protein language model with at least 500 million parameters.

6. The method of claim 3, wherein the computational model further comprises a multilayer perceptron classification head configured to receive an output of the protein language model.

7. The method of claim 1 , wherein identifying at least one subsequence comprises receiving a user-specified length of the target binding generated sequence and using greedy sampling to generate subsequences of the predicted protein sequence based on the highest predicted probability of binding to the target.

8. A pre-trained prediction model configured to be executed as code by at least one processor, wherein the prediction model is configured to receive one or more sequence input data and generate one or more guide peptide sequences without using structural data, wherein the prediction model includes a pre-trained protein language, and the pre-trained protein language is configured to output data to a four-layer fully connected neural network classification head.

9. The pre-trained prediction model of claim 8, wherein the neural network classification head is configured to receive output data from a protein language model and generate a predicted interaction likelihood for each amino acid.

10. A method for training a machine learning model to predict cosine similarity between two protein sequences, comprising the following steps: Generate a matrix of all possible pairings of target and peptide sequences; Generate an embedding for each sequence in the matrix using a protein language model; Compute the cosine similarity between each pair of embeddings in the matrix; compute the cross entropy loss between the predicted cosine similarities and the actual cosine similarities; compute the average cross entropy loss over the matrix; Update the model parameters to minimize the average cross entropy loss.

11. The method of claim 1 , further comprising generating a training dataset consisting of experimentally validated protein-protein interactions and training a machine learning model.