Methods and materials for detecting protein-protein interactions
SpARC-map leverages neoR fragments and Extreme Value Theory to rapidly map protein-protein interfaces, overcoming the limitations of traditional methods by using routine microbiological techniques and next-generation sequencing, effectively identifying interaction interfaces in both known and weakly bound complexes.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-10-02
- Publication Date
- 2026-04-09
AI Technical Summary
Existing methods for mapping protein-protein interactions, such as x-ray crystallization and cryo-EM, are complex and costly, and many protein complexes cannot be reconstituted in vitro, limiting the identification of interaction interfaces, especially for weakly bound complexes.
The SpARC-map method uses aminoglycoside phosphotransferase (neoR) fragments to reconstitute functional neoR polypeptides in host cells based on bait-prey peptide interactions, enabling rapid mapping of protein-protein interfaces through routine microbiological methods and Extreme Value Theory analysis of next-generation sequencing data.
SpARC-map efficiently identifies bona fide protein-protein interaction interfaces, including those from weakly bound complexes, without requiring complex instrumentation, and can be multiplexed to map multiple interactions simultaneously, providing cost-effective insights into protein complex structures and functions.
Smart Images

Figure IMGF000013_0001 
Figure IMGF000021_0001 
Figure IMGF000022_0001
Abstract
Description
[0001] Attorney Docket No.14017-0126WO1 / 2024-141 METHODS AND MATERIALS FOR DETECTING PROTEIN-PROTEIN INTERACTIONS CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Patent Application Serial No.63 / 703,738, filed on October 4, 2024. The disclosure of the prior application is considered part of, and is incorporated by reference in, the disclosure of this application. STATEMENT REGARDING FEDERAL FUNDING This invention was made with government support under GM024129 awarded by the National Institutes of Health. The government has certain rights in the invention. SEQUENCELISTINGThis application contains a Sequence Listing that has been submitted electronically as an XML file named “14017-0126WO1_SL_ST26.XML.” The XML file, created on September 23, 2025, is 3,034 bytes in size. The material in the XML file is hereby incorporated by reference in its entirety. TECHNICALFIELDThis document relates to identifying protein-protein interactions. For example, this document provides fragments of an aminoglycoside phosphotransferase (neoR) polypeptide as well as methods for using such fragments to map protein-protein interfaces. In some cases, an N-terminal fragment of neoR (neoRN) polypeptide can be fused to a bait polypeptide and a C-terminal fragment of neoR (neoRC) polypeptide can be used fused to a prey polypeptide, such that an interaction between the bait polypeptide and a particular prey polypeptide can lead to reconstitution of a functional neoRpolypeptide, thereby restoring kanamycin resistance to a host cell. The presence of a functional neoRpolypeptide (e.g., the restoration of kanamycin resistance) can be used to detect a protein-protein interaction. SUMMARYWe describe SpARC-map, a method for rapidly mapping the probable interaction interface between two interacting proteins. SpARC-map does not require complex instrumentation or reconstituting protein complexes in vitro, utilizing only routine Attorney Docket No.14017-0126WO1 / 2024-141 microbiological methods and a straightforward application of Extreme Value Theory. Using SpARC-map, we recover the known protein-protein interaction interface of the PCNA-p21 complex. We also use SpARC-map to map hereto unknown protein-protein interaction interfaces in the purinosome, the weakly bound complex responsible for de novo purine biosynthesis. There, we identify putative interaction interfaces between sequential purinosome enzymes that satisfy structural requirements for substrate channeling; we also identify multispecific protein surfaces that participate in multiple interactions, which we validate using site-specific photocrosslinking in live human cells. Our results show SpARC-map can identify bona fide protein-protein interaction interfaces, including those from weakly bound protein complexes inaccessible to existing approaches. This platform can rapidly map protein-protein interfaces. Compared to existing approaches, X-ray crystallization and cryo-EM, the required material and equipment are easily accessible and cost-friendly. It can be used to investigate the protein complex interactions and also applied to invent high affinity peptide towards the protein of interest, which could be potential antibody alternatives. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Although methods and materials similar or equivalent to those described herein can be used to practice the invention, suitable methods and materials are described below. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting. The details of one or more embodiments of the invention are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the invention will be apparent from the description and drawings, and from the claims. DESCRIPTION OF THE DRAWINGS Figure 1. Schematic of the SpARC-map workflow. Pair of interacting bait and prey proteins; prey residues on the bait-prey PPI interface are highlighted (101). dsDNA Attorney Docket No.14017-0126WO1 / 2024-141 containing the prey CDS is sheared into random fragments (102) and ligated into thebicistronic SpARC vector (103), in upstream fusion with the neoRC. The intact baitprotein is fused to neoRN; the fusion terminus depends on the bait. Expression of the baitand prey peptide are under the control of ITPG and aTc inducible promoters. The SpARClibrary is transformed into E. coli host and selected in the presence of kanamycin (104).SpARC vectors that express bait-binding prey peptides can result in kanamycin resistance through the reconstitution of neoRN and neoRC fragments into a functional neoR. SpARC vectors from kanamycin-resistant colonies are sequenced, mapped, and analyzed to extract the PPI interface (105). Figures 2A-2C. Extreme Value Theory analysis of NGS species abundance.Figure 2A) In addition to in-sense prey CDS fragments (orientation = +, length mod 3 = 0, frameshift = 0), some nonsense fragments, when ligated into the SpARC vectorupstream of neoRC, can translate into a random peptide fused in frame with neoRCfragment. All other fragments (neglecting the possibility of alternative start codons) will disrupt neoRC fusion and cannot result in kanamycin resistance. Figure 2B) From a representative SpARC-map dataset, shown are the measured abundance of NGS species encoding random peptides and the resulting GPD fit (dashed line), compared against NGS species that encode true prey peptides . Figure 2C) From the GPD fitted to the randompeptide abundance, a p-value can be calculated for any bona fide prey peptide given itsabundance. Prey peptides with small p are unlikely to be random binders of the bait.Figures 3A-3D. Mapping the p21-PCNA interface with SpARC-map. Figure 3A) Map of atomic contacts between PCNA and p21 residues from the co-crystal structure of PCNA with p21139-160; each symbol represents an atomic contact between the indicated residues. Figure 3B) Heatmap of PCNA binding sites on p21; solid gray is heatmap compiled from all prey peptides, the outline is heatmap compiled only using prey peptides whose abundances significantly exceeded the random peptide background. Figure 3C) Heatmap of p21 binding sites on PCNA. Figure 3D) Co-crystal structure (PDB ID: 1AXC) of PCNA homo-trimer in complex with p21139-160peptide (light gray). PCNA is colored by its affinity heatmap for p21 (dark gray / white = most / least probable). Pseudo- bonds denoting atomic contacts between PCNA and p21139-160are shown. Figure 4. High confidence (p<0.003) PPI interfaces between sequential enzymesin the purinosome. Spherical atoms highlight enzyme active sites and ligand binding sites. Attorney Docket No.14017-0126WO1 / 2024-141 The gray arrows indicate the direction of substrate transfer from upstream to downstream enzyme (domain). Heatmap plots show probable GART (monomer) binding sites on PPAT (homo-tetramer) (401); PPAT binding sites on GART(402); PFAS (monomer) binding sites on GART(403); PAICS (homo-octamer) binding sites on GART(404); PAICS binding sites on ADSL (homo-tetramer) (405); and ADSL binding sites on ATIC (homo-dimer) (406). Figures 5A-5D. Site-specific photocrosslinking with ADSLY466azFin live cells. Figure 5A) Photoactive p-azido-L-phenylalanine (azF) is inserted onto a PPI interface of a bait protein. Upon UV-irradiation, the azF can form a covalent crosslink between the bait and any prey that interacts with the bait via the specific interface. The resulting adduct can then be pulled down and analyzed. Figure 5B) Western blot image showing anti-StrepTag pull-down from HEK293T transfected with 2×Strep-ADSLY466azF. High molecular weight crosslinked adducts were apparent following UV-irradiation. Figure 5C) Mass proteomic analysis of protein abundance in anti-StrepTag pull-down with andwithout UV irradiation. The six purinosome enzymes are highlighted. Figure 5D) p-values in ascending rank, indicating PFAS, PAICS, and ATIC were significantly enriched in the UV-irradiated vs the non-irradiated samples. Figure 6. Structure of SpARC vectors (pSEP4N and pSEP4C) used in this study are shown. The two vectors differ in whether neoRN fragment is fused to the N- or C- terminus of the bait protein. Figure 7. Sketch of affinity selection in SpARC-map. Changing the selectionconditions (kanamycin and inducer concentrations) shifted the survival probability psurv(KD|KD, avg*) (top row) to sample different portions of the library affinity distribution Plib(KD) (bottom row, showing locations of strong, moderate, and weak / nonspecific binders). Figures 8A-8B. C-terminal tagging of ADSL compromised its ability to rescue ADSL-deficient crADSL cells. Figure 8A) Proliferation of crADSL was rescued by transfection with 2×StrepTag-ADSL or ADSL-2×StrepTag. Dotted line indicates initial crADSL cell seeding density at the start of the rescue assay. Rescue with ADSL- 2×StrepTag led to only half as many surviving cells as rescue with ADSL-2×StrepTag. Figure 8B) Western blot analysis of reintegrated ADSL expression in crADSL stably rescued with either N-terminally tagged 2×StrepTag-ADSL or C-terminally tagged Attorney Docket No.14017-0126WO1 / 2024-141 ADSL-2×StrepTag, showing that ADSL hyper-expression is crADSL::ADSL- 2×StrepTag. Figures 9A-9B. AF3 models of the PPAT-GART complex. Figure 9A) Map ofatomic contacts between PPAT and GART monomers across 289 AF3 models is shown. Each data point is a PPAT-GART atomic contact found in at least one model; density contours are quadratic interpolations. Also shown are heatmaps indicating the fraction of model that indicated a specific residue on PPAT (right sub-panel) or GART (bottom sub-panel) to be a part of the PPAT-GART interface. Figure 9B) Example of an AF3 modelof PPAT-GART complex that is compatible with SpARC-map results is shown; PPAT (homo-tetramer) is show in light gray, GART (homodimer) is in dark gray. Figures 10A-10B. Peptide affinity optimization through templated mutagenesis. Figure 10A) Starting with the template, a library of mutants (mutated residues in black) was generated and screened. The mutants depleted by affinity selection will reveal the scaffold of residues (dark gray) that anchors the interaction. Figure 10B) Further optimization can be performed by mutational scanning of residues outside the invariant scaffold to obtain a tight-binding peptide. Figure 11. Synthetic evolution of high affinity peptides. Starting from a random peptide pool, repeated rounds of selection and random mutation, with selection stringency increasing with each round, leads to the evolution of peptides with higher bait-binding affinities. DETAILEDDESCRIPTIONThis document provides methods and materials for identifying protein-protein interactions. For example, this document provides fragments of a neoRpolypeptide where each fragment has disrupted neoRfunction and cannot provide a host cell (e.g., a bacterial host cell) with resistance to aminoglycoside antibiotics (e.g., kanamycin and neomycin). In some cases, a fragment of a neoRpolypeptide can be a N-terminal fragment of a neoR polypeptide (a neoRN polypeptide). In some cases, a fragment of a neoRpolypeptide can be a C-terminal fragment of a neoR polypeptide (a neoRC polypeptide). When a neoRN polypeptide provided herein and a neoRC polypeptide provided herein are reconstituted, a functional neoRcan be generated. For example, when a neoRN polypeptide provided herein and a neoRC polypeptide provided herein are reconstituted within a host cell, a Attorney Docket No.14017-0126WO1 / 2024-141 functional neoRcan be generated such that the host cell exhibits resistance to aminoglycoside antibiotics (e.g., such that the host cell exhibits kanamycin resistance). In some cases, a neoRN polypeptide that can be used in the methods and materials provided herein can comprise, consist essentially of, or consist of the following amino acid sequence: MIEQDGLHAGSPAAWVERLFGYDWAQQTIGCSDAAVFRLSAQGRPVLFVKTDLS GALNE (SEQ ID NO:1). An amino acid sequence that consists essentially of SEQ ID NO:1 can be can include the amino acid sequence set forth in SEQ ID NO:1 with zero, one, or two amino acid substitutions within the articulated sequence of SEQ ID NO:1, with zero, one, two, three, four, or five amino acid residues preceding the articulated sequence of SEQ ID NO:1, and / or with zero, one, two, three, four, or five amino acid residues following the articulated sequence of SEQ ID NO:1, provided that the neoRN polypeptide retains at least some activity exhibited by a neoRN polypeptide that consists of the amino acid sequence set forth in SEQ ID NO:1 (e.g., lacks the ability to provide resistance to aminoglycoside antibiotics alone, but can provide resistance to aminoglycoside antibiotics when reconstituted with a neoRC polypeptide provided herein). In some cases, a neoRC polypeptide that can be used in the methods and materials provided herein can comprise, consist essentially of, or consist of the following amino acid sequence: ELQDEAARLSWLATTGVPCAAVLDVVTEAGRDWLLLGEVPGQDLLSSHLAPAEK VSIMADAMRRLHTLDPATCPFDHQAKHRIERARTRMEAGLVDQDDLDEEHQGLA PAELFARLKARMPDGEDLVVTHGDACLPNIMVENGRFSGFIDCGRLGVADRYQDI ALATRDIAEELGGEWADRFLVLYGIAAPDSQRIAFYRLLDEFF (SEQ ID NO:2). An amino acid sequence that consists essentially of SEQ ID NO:2 can be can include the amino acid sequence set forth in SEQ ID NO:2 with zero, one, or two amino acid substitutions within the articulated sequence of SEQ ID NO:2, with zero, one, two, three, four, or five amino acid residues preceding the articulated sequence of SEQ ID Attorney Docket No.14017-0126WO1 / 2024-141 NO:2, and / or with zero, one, two, three, four, or five amino acid residues following the articulated sequence of SEQ ID NO:2, provided that the neoRC polypeptide retains at least some activity exhibited by a neoRC polypeptide that consists of the amino acid sequence set forth in SEQ ID NO:2 (e.g., lacks the ability to provide resistance to aminoglycoside antibiotics alone, but can provide resistance to aminoglycoside antibiotics when reconstituted with a neoRR polypeptide provided herein). Also provided are methods for using neoRN polypeptides provided herein and neoRC polypeptides provided herein. In some cases, neoRN polypeptides provided herein and neoRC polypeptides provided herein can be used to detect protein-protein interactions. For example, a neoRN polypeptide provided herein can be fused to a bait polypeptide and a neoRC polypeptide provided herein can be used fused to a prey polypeptide, such that an interaction between the bait polypeptide and a the prey polypeptide can lead to reconstitution of a functional neoRpolypeptide, thereby restoring kanamycin resistance to a host cell, and the presence of a functional neoRpolypeptide (e.g., the restoration of kanamycin resistance) can be used to detect a protein-protein interaction. In some cases, the methods for using neoRN polypeptides provided herein and neoRC polypeptides provided herein to detect protein-protein interactions can be multiplexed. For example, methods provided herein can be used to simultaneously detect protein-protein interactions from multiple protein pairs. Also provided herein are nucleic acid constructs encoding neoRN polypeptides provided herein and nucleic acid constructs encoding neoRC polypeptides provided herein. For example, a nucleic acid construct can include a promoter sequence operably linked to a nucleotide sequence that can encode a neoRN polypeptide provided herein, such that that the nucleic acid construct can be used to express the neoRN polypeptide. In some cases, a nucleic acid construct that includes a promoter sequence operably linked to a nucleotide sequence that can encode a neoRN polypeptide provided herein can be introduced into a host cell such that the host cell expresses the neoRN polypeptide. For example, a nucleic acid construct can include a promoter sequence operably linked to a nucleotide sequence that can encode a neoRC polypeptide provided herein, such that that the nucleic acid construct can be used to express the neoRC polypeptide. In some cases, a nucleic acid construct that includes a promoter sequence operably linked to a nucleotide Attorney Docket No.14017-0126WO1 / 2024-141 sequence that can encode a neoRC polypeptide provided herein can be introduced into a host cell such that the host cell expresses the neoRC polypeptide. When a nucleic acid construct containing a nucleotide sequence that can encode a neoRN polypeptide provided herein is used in a method for detecting a protein-protein interaction as described herein, the nucleic acid construct also can include a nucleotide sequence that can encode a bait polypeptide such that that the nucleic acid construct can be used to express a fusion polypeptide that includes the neoRN polypeptide and the bait polypeptide. Such nucleic acid constructs can be introduced into a host cell such that the host cell expresses a fusion polypeptide that includes the neoRN polypeptide and the bait polypeptide. When a nucleic acid construct containing a nucleotide sequence that can encode a neoRC polypeptide provided herein is used in a method for detecting a protein-protein interaction as described herein, the nucleic acid construct also can include a nucleotide sequence that can encode a prey polypeptide such that that the nucleic acid construct can be used to express a fusion polypeptide that includes the neoRC polypeptide and the prey polypeptide. Such nucleic acid constructs can be introduced into a host cell such that the host cell expresses a fusion polypeptide that includes the neoRC polypeptide and the prey polypeptide. The invention will be further described in the following examples, which do not limit the scope of the invention described in the claims. EXAMPLES Example 1: Multiplex mapping of protein-protein interaction interfaces Protein-protein interactions (PPIs) are vital to biological function, and elucidating the structure-function of protein complexes is a core objective of molecular and cellular biology. Here, the PPI interface is an essential piece of structural information: it can illuminate functional consequences of an interaction, and mechanisms by which it isregulated by the cell (Wang et al., Cur Opi In Stru Biol, 74:102352 (2022); Pawson et la.,Science, 300:445-452 (2003)). Although disrupting disease-critical PPIs is a promising approach for precision medicine, for most PPIs, the interaction interface remains unknown due in part to the significant technical demands of direct structural probes such Attorney Docket No.14017-0126WO1 / 2024-141 as x-ray crystallography and cryo-electron microscopy, and in part to the fact that manyprotein complexes cannot be readily reconstituted in vitro.This example describes peptide-mapping via split-antibiotic resistance complementation (SpARC-map), an indirect method for identifying the probable PPI interface between two interacting proteins. SpARC-map requires only routine microbiological methods and straightforward statistical analysis based upon Extreme Value Theory, without reliance on complex instrumentation. Furthermore, SpARC-mapdoes not require reconstituting protein complexes in vitro, and can therefore interrogateprotein complexes where in vitro reconstitution is not possible. Finally, SpARC-map canbe multiplexed to rapidly map PPI interfaces from multiple protein pairs simultaneously. As a proof-of-concept, we show SpARC-map recovers the known PPI interface of the p21-PCNA complex. As a further challenge, we use SpARC-map to identify probable PPI interfaces in the purinosome, the complex responsible for de novo purine biosynthesis(Pareek et al., Annu Rev Biochem, 91:89-106(2022)). This weakly-bound complex of sixenzymes has never been reconstituted in vitro, and no PPI interfaces are known. UsingSpARC-map, we identify putative PPI interfaces between sequential enzymes that satisfystructural demands for substrate channeling in this pathway (Pareek et al., Science,368:283-290(2020)). We also discover multispecific protein surfaces patches thatparticipate in multiple PPIs, which we validate through site-specific photocrosslinking inlive cells. These results show that SpARC-map can reveal true PPI interfaces of even weakly bound protein complexes, thereby shining new light into their structure, function, and regulation. RESULTS Peptide mapping through split-antibiotic resistance complementation SpARC-map was motivated by the observation that many PPIs were mediated by small structural motifs, or were anchored by a small number of interfacial hot spots (Roeyet al., Chem Rev, 114:6733-6778(2014); Rajamani et al., Proc. Ntl. Acad. Sci USA,101:11287-11292(2004); Gould et al., Nucleic Acids Res, 38:D167-180(2010). The aminoacid sequence of a protein (prey) was scanned, and regions that showed affinity for an interaction partner (bait) was identified — these regions contained portions of the bait-prey interface (Figure 1, 101). Attorney Docket No.14017-0126WO1 / 2024-141 A library of bacterial expression vectors (SpARC vectors) was constructed. These vector contained a bicistronic cassette expressing both the intact bait protein and arandom prey peptide (Figure 1, 103). The prey peptides were encoded by short (100-300bp) double stranded (dsDNA) inserts, obtained by shearing dsDNA containing the preycoding sequence (CDS) (Figure 1, 102). Library selection was performed via antibioticresistance that was conditional upon bait-prey peptide binding. Aminoglycoside phosphotransferase (neoR), which inactivates kanamycin, was split into two fragments: neoRN (AAs 1-59) and neoRC (AAs 59-264). Neither fragment alone can inactivate kanamycin; however, when fused to interacting proteins, neoRN and neoRC fragments canreconstitute a functional neoR (Paschon et al., J Mol Biol, 353:26-37 (2005)). By fusingthe bait and the prey peptide respectively to neoRN and neoRC, SpARC vectors that express bait-binding prey peptides can endow their bacterial host with kanamycinresistance (Figure 1, 104), provided the bait-prey peptide affinity is sufficiently strong.Selection threshold in SpARC-map is tunable. In the SpARC vector, bait and preypeptide expressions were controlled by orthogonal inducible promoters: isopropyl- -D-thiogalactopyranoside (IPTG) inducible for bait, and anhydrotetracycline (aTc) induciblefor prey peptide (Figure 1, 103). High bait and / or prey peptide expression favorscomplexation, enabling even weak interactions to reconstitute sufficient neoRactivity to result in kanamycin resistance. Conversely, low bait and / or prey peptide expression means only strong bait-prey peptide interactions can lead to kanamycin resistance. By adjusting the concentrations of kanamycin and inducers, we estimate we can tuneSpARC-map to discriminate bait-prey peptide affinities from KD = 10 nM-1 mM (Figure7). E. coli transformed with the SpARC library were plated onto LB / agar plates containing kanamycin and inducers . Kanamycin-resistant colonies were collected and pooled. The hosted SpARC vectors were extracted, PCR amplified, and sequenced using Illumina’s MiSeq next-generation sequencing (NGS) platform. Finally, the NGS results were analyzed to identify likely bait-binding regions along the prey amino acid sequence(Figure 1, 105).Extreme Value Theory analysis of sequence abundances Attorney Docket No.14017-0126WO1 / 2024-141 The premise of SpARC-map analysis is as follows: for a prey peptide to be a specific binder of the bait peptide, its bait-binding affinity must be significantly higher than that of random peptides (on average). In the NGS data, sequences encoding bait- binding prey peptides are more abundant than (hypothetical) random DNA sequences encoding random peptides. The process of constructing the SpARC library automatically supplies a pool of random sequences, encoding random peptides, from which the bait- random peptide background is estimated. Prey peptide coding sequences were generated by shearing dsDNA containing the full prey CDS. Only 1 / 18 of the resulting dsDNA fragments ligate in-sense and translate into a prey peptide fused in frame to neoRC; the remainder are inverted, and / or frameshifted, and / or contain extra bases. Most such nonsense inserts disrupt the downstream framing of the neoRC fusion, or contain stop codons. SpARC vectors with such inserts cannot result in kanamycin resistance (Figure 2A). However, some nonsense inserts translate through to effectively random peptides fused in frame with neoRC. These interact with the bait and potentially reconstitute neoR. These random peptides arerecognized in the NGS results, and for each species i, quantify its sequence abundance ni.The distribution PNS(ni) is the background due to bait-random peptide binding.For a random peptide to endow its host with sufficient kanamycin resistance to survive selection, its affinity for the bait must exceed the threshold determined by the selection condition: greater the exceedance, greater the survival . The Pickands– Balkema–De Haan Theorem, a central result of the Extreme Value Theory, states that the distribution of threshold exceedances, under generous conditions, will be a GeneralizedPareto Distribution (GPD) (Coles et al., Springer, (2001)). The bait-random peptidebackground is PNS(ni) = GPD(ni; , , ), with fitting parameters , , and are modeled(Figure 2B). The fitted PNS(ni) is used to evaluate whether a prey peptide species issignificantly more abundant than the bait-random peptide background (Figure 2C). If it is, the prey peptide is likely a specific binder of the bait that contains portions of the bait- prey PPI interface; if it is not, then it is likely a random binder that should be excluded from analysis. Mapping the PPI interface between PCNA and p21 Attorney Docket No.14017-0126WO1 / 2024-141 As a proof-of-concept, SpARC-map was used to probe the known PPI interface between PCNA and p21. PCNA, as a homo-trimer, forms a ring-shaped DNA clamp that is essential for DNA replication and repair. p21, better known as an inhibitor of cell-cycle dependent kinases, can bind PCNA and inhibit its normal functions (Waga et al., Nature,369:574-578(1994); Prives et al., Cell Cycle, 7:3840-3846(2008)). The PCNA-bindingregion of p21 is its C-terminal PIP-box motif (AAs 140-164), first identified through analyses of p21 truncations and fragments (Chen et al., Nucl Acids Res 24:1727-1733(1996); Warbrick et al., Curr Biol, 5:275-282(1995)). Subsequently, the complete PPIinterface was revealed by the x-ray crystal structure of PCNA in complex with p21139-160(Gulbis et la., Cell, 87:297-306(1996)).Two SpARC libraries were constructed: one with PCNA as bait and p21 CDS fragments as prey, another with p21 as bait and PCNA CDS fragments as prey. Figure 3 shows the PCNA-p21 atomic contact map (Figure 3A) juxtaposed with affinity heatmaps generated from these libraries post-selection (Figures 3B-3C). For heatmaps, the x-axis indicates prey residue number, and the y-axis indicates the probability that the prey residue is tiled by a mapped NGS read. For PCNA binding on p21, including all p21 peptides in the analysis resulted in a heatmap that fully covered the p21 amino acid sequence. Its main feature is the large peak spanning AAs 139-160, corresponding to the true p21 interface with PCNA (Figure 3B). Restricting analysis to species that exceeded the random peptide background suppressed heatmap signal outside the true interface region. From the tiling of the interface by distinct p21 fragments, it is inferred that while p21139-164had the highest PCNA affinity (the most abundant species), truncation to p21143- PCNA-binding but at a lower affinity (lower abundance). This is consistent with known PCNA binding affinities of p21 peptides: KD = 6.4±2.8 nM for p21140-163, and67±9 nM for p21143-157 (Prestel et la., Cell Mol Life Sci, 76:4923-4943(2019)).For the reciprocal PCNA interface that binds, including all species resulted in a heatmap containing two comparably sized features, AAs 104-139 and 152-207, along the PCNA peptide sequence. However, once the analysis was restricted to those species that significantly exceeded the random peptide background, the feature corresponding to AAs 152-207 vanished, leaving only the hotspot on AAs 104-139 (the most abundant species corresponding to PCNA115-133) (Figure 3C). This result is consistent with the crystal Attorney Docket No.14017-0126WO1 / 2024-141 structure of PCNA in complex with p21139-160, where the main p21 contact on PCNA, accounting for ~50% of atomic contacts, was on PCNA AAs 118-131 (Figure 3D). Mapping the PPI interface in the purinosome Compared to the p21-PCNA complex, the purinosome posed more challenging test. The purinosome is a complex of at least six enzymes — PPAT, GART, PFAS, PAICS, ADSL, and ATIC — responsible for de novo purine biosynthesis (DNPB) (Zhanget al., Cell Mol Life Sci, 65;3699-3724(2008)). Imaging studies have shown that theseenzymes can colocalize into punctate bodies in purine-starved cells, while proximity, co- fractionation, and co-immunoprecipitation assays have all supplied molecular evidencefor PPIs (Wan et al., Nature, 525:339-344(2015); Sha et al., J Biol Chem, 107620(2024)).However, it is unknown if any of the reported interactions are direct. All evidence indicates that the purinosome is a weakly-bound complex; it has never been reconstituted in vitro, wholly or partially, and no PPI interfaces are known. To map probable PPI interfaces in the purinosome, six SpARC libraries were constructed and screened, one per each purinosome enzyme as bait; and all of them incorporated peptide encoding sequences generated by shearing an equimolar mix of CDSs (13 kbps) from all six enzymes (4335 AAs total). First, the putative PPI interfaces between sequential enzymes were examined. Given the clear metabolomic signatures ofsubstrate channeling in the purinosome (Warbrick et al., Curr Biol, 5:275-282(1995)), atleast some enzyme pairs that catalyzed sequential reactions were expected to physically interact in a way that can facilitate substrate transfer. PPAT-GART. PRA, product of PPAT in reaction 1 of DNPB, was the substrate of GART’s GAR synthase domain for reaction 2. The instability of PRA (half-life <1 minute) was early motivation for suggesting PPAT and GART must physically interact totransfer such a labile intermediate (Schendel et al., Biochemistry, 27:2614-2623(1988);Rudolph et al., Biochemistry, 34:2241-2250(1995)). On GART, probable PPAT bindingsites were found on AAs 14-55 (p<0.001) and 159-232 (p<0.003); these flanked the GARsynthase active site that received PRA from PPAT (Figure 4, 401), as expected from thestructural requirements substrate channeling. On PPAT, the putative GART-binding sitewas on its C-terminus, AAs 468-504 (p<0.001) (Figure 4, 402). Attorney Docket No.14017-0126WO1 / 2024-141 GART-PFAS. GART contains three catalytic domains that catalyze reactions 2, 3, 5 of DNPB. PFAS catalyzes reaction 4, receiving its substrate FGAR from GART’s GAR transformylase domain, and sending its product FGAM forward to GART’s AIR synthase domain. On GART, three probable PFAS-binding sites (p<0.003) were found: AAs 24- 48, AAs 773-801 on the AIR synthase domain that received FGAM from PFAS, and AAs 1001-1010 on the GAR transformylase domain that supplied FGAR to PFAS (Figure 4, 403). The latter two were structurally consistent with substrate channeling from GART to PFAS and back. The reciprocal GART-binding interface on PFAS was not identified. GART-PAICS. AIR, the product of GART in reaction 5, was the first substrate of PAICS, which catalyzed reactions 6 and 7. On GART, a probable PAICS binding site, AAs 618-669 (p<10-3), was found on the AIR synthase domain that supplied AIR toPAICS (Figure 4, 404) and was structurally consistent with substrate transfer from GARTto PAICS. The reciprocal GART-binding interface on PAICS was not identified. PAICS-ADSL. SAICAR, the product of PAICS in reaction 7, was the substrate for ADSL for reaction 8. On ADSL, a probable PAICS binding site was found on AAs 451-484 (p<0.001) (Figure 4, 405). On PAICS, three potential ADSL-binding interfaces wereidentified: AAs 251-270, 281-348, and 413-425, with low confidence (p<0.004 – 0.01); two of these, AAs 251-270 and 413-425, were adjacent to PAICS’s SAICAR synthase domain that supplied ADSL. ADSL-ATIC. ADSL supplied its product AICAR as the initial substrate for ATIC, which catalyzed reactions 9 and 10. On ATIC, putative ADSL binding site was found on AAs 174-214 (p<0.002) and 489-529 (p<0.003); the latter lined the active site entrance on ATIC’s AICAR transformylase domain that received AICAR from ADSL (Figure 4, 406), and was structurally consistent with substrate channeling from ADSL to GART. Two low confidence (p<0.008) ATIC-binding sites were identified on ADSL, AAs 73- 130 and 447-484; the latter coincided with the probable PAICS-binding site on ADSL. In addition to the expected PPI interfaces between the sequential enzymes, putative PPI interfaces were found between enzymes that were distal along the DNPB pathway (Table 1). Indeed, a dense network of interactions was found, with probable PPI interfaces (at least partially), indicative of a direct physical interaction, connecting every enzyme pair. Surprisingly, some enzymes appeared to use the same surface regions to participate in multiple PPIs. For example, high-confidence putative PPI interfaces was Attorney Docket No.14017-0126WO1 / 2024-141 observed between ADSL and all other purinosome enzymes. The same C-terminal region of ADSL (AAs 451-484) was used to interact with distal enzymes PPAT and PFAS, as well as enzymes that catalyzed immediate upstream (PAICS) and downstream (ATIC) reactions; however, the putative GART-binding site on ADSL (AAs 391-402) was distinct. Validating PPI interfaces via in vivo photocrosslinking While many of the putative purinosome PPI interfaces (e.g., between sequential enzymes) identified by SpARC-map were consistent with structural requirements for substrate channeling, others (e.g. the multispecific C-terminus region of ADSL) were quite unexpected. Therefore, it was sought to further validate SpARC-map identifications using site-specific photocrosslinking in live human cells. Using genetically encoded unnatural amino acid incorporation, p-azido-L-phenylalanine (azF), a phenylalanine derivative, can be inserted into the amino acidsequence of a protein (Chin et al., Jour of the Amer Chem Soc 124:9026-9027(2002);Chatterjee et al., Proc Natl Acad Sci USA, 110:11803-11808(2013)). Upon UV-activation, the aryl azide group on azF can attack nearby double bonds, primary amines, C-H and N-H groups, etc., to form a covalent bond (26). Thus azF, when inserted into a PPI interface of a bait protein, can photocrosslink prey proteins that interact with the bait specifically via that interface (Figure 5A); its small reactive radius (<0.3 nm) and short lifetime (~1 ns post-photoactivation) minimizes nonspecific crosslinking due to random molecular collisions. To validate whether ADSL451-484is a multispecific PPI interface between ADSL and other purinosome enzymes, a mutant ADSL was constructed containing the substitution Y466azF, as well as a 2×StepTag epitope tag fused to its N-terminus, and was expressed in HEK293T cells cultured in purine-depleted media. Following brief UV irradiation (for photocrosslinking samples), the cells were lysed and anti-StrepTag affinity purification was used to pull-down ADSLY466azF(plus its crosslinked adducts and co-precipitants) for analysis using protein tandem mass spectrometry (Figure 5B). Pulled- down proteins should include those that interact specifically with ADSL but not via the C-terminal region, and those that interact with ADSL specifically via ADSL451-484. Pull- down of the latter, but not the former, should increase following UV-irradiation. Attorney Docket No.14017-0126WO1 / 2024-141 Of the four DNPB enzymes (PPAT, PFAS, PAICS, ATIC) found by SpARC-map to interact with ADSL via ADSL451-484, three (PFAS, PAICS, and ATICS) were low / undetectable in non-irradiated controls, but were abundant in UV-irradiated samples. They were among the top 20 most significantly UV-enriched proteins (out of 1717 proteins detected). PPAT was only detected upon UV-irradiation, but at low levels that did not result in statistically significant differences between UV-irradiated versus non- irradiated samples. GART, which SpARC-map indicated interacts with ADSL but via a distinct region, was detected at high levels in both UV-irradiated and non-irradiated samples, with no statistically significant difference between the two (Figures 5C-D). Bait map. High-confidence interfaces (p<0.003) are highlighted in bold. Attorney Docket No.14017-0126WO1 / 2024-141 Example 2: Supplemental materials for “Multiplex mapping of protein-protein interaction interfaces” METHODS AND MATERIALS Structure and design of SpARC vectors The SpARC vectors pSEP4N and pSEP4C (Figure 6) harbor low-copy pBR322 origin of replication and ampicillin-resistance gene. Transcriptional suppressors lacI and tetR(separated by an internal ribosome binding site) were expressed under the control ofthe lacIq promoter. Expression of neoRN-bait protein fusion (in pSEP4N) or bait protein-neoRN fusion (in pSEP4C) was controlled by IPTG-inducible trc promoter / lac operator.Expression of the prey peptide-neoRC fusion was controlled by the aTc-inducible PletO-1promoter. Generation of prey peptide fragments Sequence-independent enzymatic digestion (DNA Fragmentase, NEB M0348)was used to fragment dsDNA containing prey CDS (without stop codon). For multiplexedexperiments, an equimolar mix of dsDNA containing different prey was fragmented. For each experiment, a time-series was conducted to ascertain the time required such that the bulk of fragments was between 100-300 bp in length. Digested dsDNA fragments were purified (QIAquick PCR purification kit, Qiagen 28104) and blunted (Quick Blunting Kit, NEB E1201) for subsequent use in ligation reactions. Library construction First, the bait protein CDS was cloned into appropriate SpARC vector using the multiple cloning sites (MCS). p21, PCNA, GART, PFAS, PAICS, ADSL, ATIC were cloned as bait into pSEP4N. PPAT was cloned as bait into pSEP4C. Then, the bait- containing SpARC vector was cut at NotI and SbfI restriction sites flanking the library insertion site, dephosphorylated (Shrimp Alkaline Rhosphatase, NEB M0371), blunted (DNA Polymerase I, Large Fragment, NEB M0210), and purified. The digested / blunted vector backbone was then mixed with prey CDS fragments at a molar ratio ~1:7, and ligated (T4 DNA Ligase, NEB M0202). Finally, the ligation mix was transformed intohome-made chemically competent E. coli (strain NEBStable) (Yang et al., MicrobioSpectr 10:e02497-02422(2022)). Attorney Docket No.14017-0126WO1 / 2024-141 Library pre-selection To eliminate vectors with inserts that cannot result in a peptide fused in-framewith the neoRC fragment (and therefore cannot reconstitute kanamycin resistance), the ligated SpARC library was pre-selected by outgrowing transformed bacteria in the presence of low concentrations of kanamycin. For p21 (bait)-PCNA (CDS fragments peptides) and PCNA (bait)-p21 (CDS fragments) libraries, transformed bacteria were pre- selected on LB agar with 100 ng / mL aTc, 100 µg / mL ampicillin, and 4 µg / mL kanamycin; no IPTG was added, relying on leaky expression from the IPTG-inducible promoter for bait expression. For DNPB enzyme (bait)-DNPB enzymes (CDS fragment) libraries, transformed bacteria were pre-selected on LB agar with 1 mM IPTG, 100 ng / mL aTc, 100 µg / mL ampicillin, and 5 µg / mL (GART, PAICS, ADSL, and ATIC), 3 µg / mL, (PPAT), and 2 µg / mL (PFAS) kanamycin. In all cases, ~1% of transformants plated onto LB agar containing only 100 µg / mL ampicillin for counting, and as a sample of the unselected input library which was then sequence for quality control. In all cases, pre-selection eliminated between 85-95% of the SpARC library as ligated. Following outgrowth, colonies were collected, pooled, and the hosted SpARC vectors were extracted (QIAprep Spin Miniprep Kit, QIAGEN 27106). Library selection (final) Prior to final selection, the kill curve of the (pre-selected) SpARC library being screened was measured. The pre-selected SpARC library was transformed into bacteria and plated onto a series of LB agar plates containing 100 ng / mL aTc, 100 µg / mL ampicillin, IPTG (1 mM for DNPB libraries, no ITPG for p21 and PCNA libraries), with varying concentrations of kanamycin (5-640 µg / mL). In parallel, the parental SpARC vector containing only the bait protein, but no peptide inserts was subjected to identical selection conditions. As final selection conditions, 2-3 kanamycin concentrations were chosen under which the bait-only parental SpARC vector cannot survive, but the SpARC library still yielded surviving colonies. A large-scale transformation of the (pre-selected) SpARC library was conducted with 5-50 fold redundancy (2.5-5 million transformants), and the transformants were plated onto the LB agar plates under these selection conditions. Attorney Docket No.14017-0126WO1 / 2024-141 Library (bait- Size (as ligated) Size (pre- Selection Condition (final) prey) selected) Sequencing Post final selection, surviving SpARC library colonies were pooled, collected, and the host SpARC vectors were extracted. The prey peptide coding sequence was PCR amplified using barcoded NGS sequencing primers that annealed on invariant sequences that flanked the prey peptide insertion site. PCR amplification was for 12-cycles, and the purified PCR product was sequenced using the AmpliconEZ service from Azenta Life Sciences. This service was based on Illumina’s MiSeq2 platform and yielded >50k (guaranteed minimum, >200k typical) 250 bp paired-end reads per sample. Prey peptide identification and mapping Paired-end NGS reads were trimmed, quality-filtered, and merged; and gapped sequences were discarded. For each merged NGS read, the longest open reading frame Attorney Docket No.14017-0126WO1 / 2024-141 that was fused in-frame with neoRC was identified, and denoted it as the peptide coding sequence (pCDS). ATG, TTG, and GTG were accepted as start codons (Hecht et al., Nucleic Acids Res, 45:3615-3626(2017)). The pCDS was then mapped against prey CDS(s); pCDSs that were less than 15 bps in length, or that were concatemers (either of CDS fragments from different preys, or of noncontiguous fragments of the same prey) were discarded. All NGS processing and mapping were done using a custom Julia workflow built using BioSequences.jl and FASTX.jl packages. Estimating the random peptide background pCDSs that were inverted and / or frameshifted with respect to the prey CDS to which they were mapped to, were designated as encoding random peptides. All suchrandom species i were identified and their abundances {ni } were quantified. Theresulting abundances {ni} were fitted to a generalized Pareto distribution (GPD) using maximum likelihood estimation as implemented in the Julia package Extremes.jl (Jalbertet al., Journ of Stat Soft, 109(2024)). The result is the background abundance distributiondue to random peptide-bait binding, PNS(ni) = GPD(ni; , , ). Then, given an NGSspecies, with abundance n, that contained a mappable pCDS encoding an in-frame preypeptide, its p-value was computed as Alarge p-value indicates the abundance of said NGS species is similar to whatone expects from a random peptide, and the encoded prey peptide is unlikely to be a specific binder of the bait. In vivo protein photocrosslinking HEK293T cells were seeded in 15-cm dishes and cultured in purine-depleted media (DMEM, 10% v / v dialyzed FBS with 10 kDa molecular weight cutoff) until 40- 60% confluency. Cells were then transfected (Xfect DNA Transfection Reagent, Takara 631318) with p2×StrepTag-ADSLY466TAG(driving the expression of 2×StrepTag - ADSLY466TAGfusion under the CMV promoter, 35 µg) and pMAH-POLY (expressing theunnatural amino acid incorporation system (Chen et al., J Am Chem Soc, 135:14940- Attorney Docket No.14017-0126WO1 / 2024-141 14943(2013); 30 µg), and maintained in purine-depleted media supplemented with 1.8 mM 4-azido-L-phenylalanine (Chem-Impex 06162) for 48 hours before collection. For UV-irradiated samples, cell culture media was removed, and cells were washed once with PBS buffer. Cells were then immersed under 7 mL of PBS and irradiated (lid off) for 3 minutes using the germicidal UV lamp (260 nm peak emission) inside a biological safety cabinet (distance between sample and light source ~ 30 cm). For non-irradiated samples, cells were washed and collected. Cells were harvested and lysed in lysis buffer (50 mM Tris pH 7.8, 135 mM KCl, 15 mM NaCl, 5 mM MgCl2, 5% v / v glycerol, 1% v / v Triton X-100, 1 mM DTT, 2 mM ATP, 1 mM l-glutamine, supplemented with protease and phosphatase inhibitors (PierceA32955, 78428). Cell lysate was clarified by centrifugation (10 minutes × 10,000 g, 4°C), and the supernatant was collected and incubated with anti-StrepTag magnetic beads (IBA Lifesciences 2-5090-010). The beads were then sequentially washed 1× with lysis buffer, 1× with buffer B (lysis buffer without triton X-100 and protease inhibitor), 2×with buffer C (buffer B with 500 mM NaCl), and 2× with buffer B. Washed beads were transferred to a new tube and resuspended in buffer B.1 / 8 of the beads were boiled to obtain the protein sample for western blot analysis; the remainder was treated with trypsin (Pierce Trypsin Protease, MS Grade, 90057) for on-bead protein digestion followed by iodoacetamide treatment; digested peptides were purified using C18 spin column (Pierce 89870). Experiment was repeated to obtain two biological replicates per treatment condition (UV-irradiated vs non-irradiated). Mass spectrometry and data analysis Purified peptides were analyzed using liquid chromatography electrospray ionization tandem mass spectrometry on a Thermo Orbitrap Eclipse instrument. Peptide species were identified and quantified (label-free quantification) using Proteome Discoverer 1.4. To identify differentially enriched protein species, we start from statistical null model that differences in protein enrichment between irradiated and non- irradiated samples are simply the inherent variation expected for biological replicates, i.e., UV-irradiation has no effect. The followings were assumed: firstly, the conditional inter-replicate distribution for the abundance of protein species i, Prep(ni|ni, avg), given its inter-replicate mean ni, avg, depends only on the inter-replicate mean; and, secondly, this Attorney Docket No.14017-0126WO1 / 2024-141distribution is the same for proteins i, j, if their abundances are the same, i.e., if ni, avg = nj,avg = navg, then Prep(ni|ni, avg) = Prep(nj|nj, avg) = Prep(n|navg). Prep(n|navg) was estimated bybinning together proteins with inter-replicate means, and use Gaussian kernel density estimation, as implemented by the KernelDensity.jl package of the JuliaStats distribution, to approximate the resulting abundance distribution across replicates.Finally, to evaluate whether a protein species i exhibits UV-dependentenrichment / depletion, given its abundance in non-irradiated samples is ni, avg, UV-, theestimated Prep(n|navg = ni, avg, UV-) was used to compute its p-value. A small p-values indicates that the abundance of the protein in UV-irradiated samples much higher than expected from inter-replicate variation, given its abundance in non-irradiated samples. Modeling kanamycin resistance in SpARC-map Reconstituting kanamycin resistance. Consider a bacterial host in an environment with kanamycin concentration [kan]out. The kanamycin concentration [kan]ininside the host cell depends on two processes: the transport of kanamycin from the external environment into the cell, and the degradation of kanamycin by reconstituted neoR inside the cell. The evolution of [kan]in was modeled as (2.1) The first term describes the influx of kanamycin from the external environmentinto the cell, characterized by an influx rate µ. The second term describes the degradationof kanamycin by (reconstituted) neoR catalytic rate kcat. The inactivation of kanamycin byneoRwas assumed to follow Michaelis-Menten kinetics, such that (2.2) and that the reconstitution of intact neoRis entirely due to the 1:1 binding (withdissociation constant KD) of the bait protein and prey peptide to which the neoRN andneoRC fragments were fused, Attorney Docket No.14017-0126WO1 / 2024-141 (2.3) In a steady state, the kanamycin concentration inside the cell is given by (2.4)where it was assumed that [kan]in < KM is low so the first order kinetics apply. The hostcell can survive only if [kan]inis below some critical value, which was assumed as equalto the minimum inhibitory concentration of kanamycin for E. coli, [kan]MIC = 3.4 µM. Interms of the binding affinity KD between the bait protein and the prey peptide, thecondition for host survival is given by (2.5) The selection threshold KD* depends on several kinetic parameters: for the split neoRsystem, kcat / KM 0.1 µM-1 s-1(Paschon et la., J Mol Biol, 353:26-37(2005)); for into E coli, µ 0.002 s-1(Nakae et al., Antimicro agents and Chemo,22:554-559(1982)). Taking [kan]ext / [kan]MIC= 5 – 100, and assuming [bait] ~ [prey] = 0.1 – 10 µM (equivalent to ~ 0.01 – 1 mg of a 30 kDa recombinant protein expressed per literof culture), yields a tunable range for KD* between 5 nM and 1 mM.In reality, even for cells that host the identical bait / prey pair, cell-to-cell variation in bait and prey expression (and therefore neoRreconstitution), kanamycin uptake, andkanamycin sensitivity means KD* is not a single hard threshold, but will vary from cell tocell. The probability of survival given psurv(KD|KD, avg*), where KD, avg* is the averageselection threshold across all cells, will be a sigmoidal function with limits psurv 1when KD KD, avg*, psurv 0 when KD KD, avg* (Figures 7). For cells transformed witha SpARC library with many bait-binding prey peptides (most of which are randombinders) with varying KD, characterized by some distribution Plib(KD), the NGS speciesabundance extracted from the post-selection cell population will be directly proportionalto psurv(KD|KD, avg*). However, if selection is too weak (large KD, avg*), it may be difficult Attorney Docket No.14017-0126WO1 / 2024-141to distinguish strong versus moderate binding peptides because their psurv will be nearlyequivalent. Conversely, very strong selection (small KD, avg*) can drive the library toextinction. Optimal selection is when the decline of psurv(KD|KD, avg*) coincides with theleft wing of Plib(KD). Here, it is noted that the fraction of transformed cells that survives agiven selection condition (characterized by KD, avg*) is simply(2.6)fsurv(KD, avg*) 0(no selection), fsurv 0 when KD, avg* (maximum selection); where fsurv(KD, avg*)declines most rapidly demarcates the range of optimal selection conditions. C-terminal tagging of ADSL compromised its ability to rescue ADSL-deficient crADSLcells. ADSL deficient HeLa (crADSL) (Baresova et al., Mol Genet Metab, 119:270-277(2016)) requires adenine supplementation in culture media for survival and proliferation. The ability of CMV-driven expression of exogenous ADSL in crADSL to rescue the deficient phenotype was assessed by transfecting crADSL with plasmids expressing 2×Strep-ADSL or ADSL-2×Strep fusion. Transfected crADSL was transferred to purine-depleted media and monitored for survival and proliferation. After two weeks, crADSL expressing exogenous 2×StrepTag-ADSL had approximately twice as many surviving cells as crADSL expressing exogenous ADSL-2×StrepTag (Figure 8A), suggesting that the 2×StrepTag fusion on the ADSL C-terminus negatively impacted its function. After extended culture, both rescue experiments yielded colonies of cells (with many few colonies in the ADSL-2×StrepTag) that re-integrated the exogenous ADSL into their genome; these can survive and proliferate indefinitely in purine-depleted media. This result indicates both 2×StrepTag-ADSL and ADSL-2×StrepTag retained ADSL’s essential enzymatic activity. However, when examined for protein expression, the crADSL::ADSL-2×StrepTag stable rescue expressed (re-integrated) ADSL at ~35× that of wildtype HeLa, while ADSL expression in crADSL:: 2×StrepTag-ADSL was comparable to that of wildtype HeLa (Figure 8B). Since the C-terminus of ADSL does not participate in known active or ligand binding sites, the overexpression of ADSL in crADSL::ADSL-2×StrepTag was interpreted as a compensation for disturbed PPIs in Attorney Docket No.14017-0126WO1 / 2024-141 ADSL-2×StrepTag, and mirrored a very similar effect that was previously observed forPAICS (He et al., J Biol Chem, 298:101853(2022)).AlphaFold modeling of PPAT-GART complex Using the AlphaFold Server (golgi.sandbox.google.com), 305 randomly seeded AlphaFold3 (AF3) models of the PPAT-GART complex consisting of 4 copies of PPAT and 2 copies of GART were generated. After discarding models containing clashes (16 / 305), for each remaining model, all its atomic contacts (defined as atomic separation < 0.4 nm) between PPAT and GART monomers were identified. As shown in Figure 9A, there was no clear consensus among these models as to the structure of the PPAT-GART complex. On PPAT, AAs 470-500, which overlapped the SpARC-map identified GART- interacting region (AAs 468-504), was found in ~10% of models as the part of PPAT's interface with GART. On GART, ~2% of the models showed AAs 14-55 and 159-232, regions identified by SpARC-map as PPAT-binding, as parts of GART's interface with PPAT. Figure 9B shows one example of an AF3 model that is compatible with SpARC- map results. The majority of AF3 models identified GART's interface with PPAT on GART's AIR synthase (AAs 434-809) and / or GAR transformylase (AAs 809-1010) domains, in configurations inconsistent with SpARC-map results, and incompatible with substrate channeling between PPAT and GART. Example 3: Affinity screen of synthetic peptide libraries with SpARC (SpARC screen) In SpARC-map, the bait-binding surface were mapped on the prey protein by screening a library of prey peptide fragments, and identified those that bind the bait. This Example describes using the same SpARC modality to screen libraries of synthetic peptides, and to discover novel peptides that can bind to the bait protein with high affinity. Such affinity peptides can be used as affinity reagents for detecting and / or purifying the bait protein from biological sources, and potentially as drugs if on binding the bait, they disrupt the bait's native PPIs and associated biological activities. Affinity optimization through templated mutagenesis Prey peptides on the bait-prey PPI interface are potential competitive disruptors of the bait-prey interaction because they bind the same region on the bait protein as the Attorney Docket No.14017-0126WO1 / 2024-141 intact prey protein. However, this is unlikely to be practical without modifications to achieve higher bait-binding affinity. A two-stage strategy was used to enhance the bait-binding affinity of the wildtype prey peptide. First, starting with the wildtype prey peptide as a template, a SpARC library containing all possible single-residue mutations was generated. The mutant peptide library was selected to identify those that retain their affinity for the bait. Mutants that do not retain affinity (i.e., are depleted from the unselected SpARC library) are likely to have mutations on key residues that anchor the bait-prey interaction on their specific PPI interface (Figure 10A). Such residues were identified as the scaffold that should be retained. Next, a second SpARC library of mutant peptides was constructed, all of which retains the same scaffold, but now includes all possible single-residue mutation elsewhere (Figure 10B). This second library was selected to identify mutant peptides that bind the bait more tightly than the original prey peptide. The process can be iterated multiple times to yield tighter binding mutant peptides. Since these mutant peptides bind on the same region on the bait protein as the prey protein, they are potentially highly specific competitive inhibitors of the bait-prey interaction. For weakly bound complexes such as the purinosome, a modest enhancement in binding affinity can be sufficient for the optimized peptides to become effective competitive PPI antagonists. Synthetic evolution of tight binding peptides As a parallel strategy, it was sought to identify and evolve wholly synthetic peptides that can bind the bait protein, or a specific region thereof, without reference to naturally occurring templates. A SpARC library was constructed that expresses both the bait (or bait fragment) and random peptides (10-20 AAs) encoded by degenerate DNA sequences. This library was subjected to selection, and identified peptide hits that showed affinity for the bait (Figure 11). It is probable that screening from a random library may only yield peptides with a modest affinity for the bait. Therefore, it was sought to enhance the affinity of hit peptides through synthetic evolution. The post-selection SpARC library was enriched in peptides with increased bait affinity. Random mutations were introduced using error- prone PCR to amplify the peptide-encoding regions of the post-selection SpARC library, Attorney Docket No.14017-0126WO1 / 2024-141 and used mutated peptide sequences to create the next-generation SpARC library. The sequence of selection and mutation can be iterated for multiple rounds of evolution. With each round of evolution, the selection stringency can be gradually increased, and push the SpARC library towards ever higher bait-binding affinities. This process, also known as “simulated annealing,” mirrors that of antibody affinity maturation in immune response.
[0002] Attorney Docket No. 14017-0126W01 / 2024-141
[0003] OTHER EMBODIMENTS
[0004] It is to be understood that while the invention has been described in conjunction with the detailed description thereof, the foregoing description is intended to illustrate and not limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.
Claims
Attorney Docket No.14017-0126WO1 / 2024-141 OTHEREMBODIMENTSIt is to be understood that while the invention has been described in conjunction with the detailed description thereof, the foregoing description is intended to illustrate and not limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.Attorney Docket No.14017-0126WO1 / 2024-141 WHATISCLAIMEDIS:
1. A nucleic acid construct comprising a promotor sequence operably linked to anucleic acid comprising (a) a nucleotide sequence that encodes a N-terminal fragment of an aminoglycoside phosphotransferase (neoR) polypeptide (a neoRN polypeptide), and (b) a nucleotide sequence that encodes a bait polypeptide.
2. The nucleic acid construct of claim 1, wherein said promoter is an induciblepromoter.
3. The nucleic acid construct of claim 1, wherein said promoter is an isopropyl- -D-thiogalactopyranoside (IPTG) promoter.
4. The nucleic acid construct of any one of claims 1-3, wherein said nucleic acidconstruct is an expression vector.
5. The nucleic acid construct of claim 4, wherein said expression vector is a bacterialexpression vector.
6. The nucleic acid construct of any one of claims 1-5, wherein said neoRNpolypeptide comprises, consists essentially of, or consist of an amino acid sequence set forth in SEQ ID NO:1.
7. A nucleic acid construct comprising a promotor sequence operably linked to anucleic acid comprising (a) a nucleotide sequence that encodes a C-terminal fragment of an aminoglycoside phosphotransferase (neoR) polypeptide (a neoRC polypeptide), and (b) a nucleotide sequence that encodes a prey polypeptide.
8. The nucleic acid construct of claim 7, wherein said promoter is an induciblepromoter.
9. The nucleic acid construct of claim 7, wherein said promoter is ananhydrotetracycline (aTc) promoter.Attorney Docket No.14017-0126WO1 / 2024-141 10. The nucleic acid construct of any one of claims 7-9, wherein said nucleic acid construct is an expression vector.
11. The nucleic acid construct of claim 10, wherein said expression vector is a bacterial expression vector.
12. The nucleic acid construct of any one of claims 7-11, wherein said neoRC polypeptide comprises, consists essentially of, or consist of an amino acid sequence set forth in SEQ ID NO:
2.
13. The nucleic acid construct of any one of claims 7-12, wherein said prey polypeptide is synthetic polypeptide.
14. The nucleic acid construct of any one of claims 7-12, wherein said prey polypeptide is a recombinant polypeptide.
15. A method for detecting protein-protein interactions, said method comprising: (1) providing a host cell with: (a) a nucleic acid construct of any one of claims 1-6 such that said host cell expresses a first fusion polypeptide comprising said neoRN polypeptide and said bait polypeptide, (b) a nucleic acid construct of any one of claims 7-14 such that said host cells expresses a second fusion polypeptide comprising said neoRC polypeptide and said prey polypeptide, and wherein an interaction between said bait polypeptide and said prey polypeptide reconstitutes a neoRpolypeptide; and (2) providing said host cell with an aminoglycoside antibiotic; and (3) detecting aminoglycoside antibiotic resistance, wherein resistance to said aminoglycoside antibiotic indicates a protein-protein interaction between said bait polypeptide and said prey polypeptide, and wherein a lack of resistance to said