Methods and compositions for target complex detection and quantification

The system addresses the limitations of current protein interaction mapping by using barcode combinations to determine complex protein interactions, enabling accurate abundance measurements and interaction maps for disease diagnosis.

WO2025262551A1PCT designated stage Publication Date: 2025-12-26NATIONAL UNIVERSITY OF SINGAPORE
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/056100
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-18
Filing Date
2025-06-13
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Current technologies for mapping complex protein interactions are limited by their binary nature, requiring genetic manipulation and extensive processing, leading to incomplete and indirect profiling, and often miss weak interactions.

Method used

A system and method utilizing barcode combinations to determine the distribution and abundance of target complexes, including a processor and memory device for executing operations to analyze barcode contributions and probabilities, and a composition of complexes with binding domains, flexible linkers, and encoding polynucleotides to identify and quantify protein interactions.

Benefits of technology

Enables comprehensive and direct profiling of complex protein interactions, providing accurate abundance measurements and interaction maps, and facilitating disease diagnosis and prognosis by quantifying target complexes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025056100_26122025_PF_FP_ABST
    Figure IB2025056100_26122025_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein is TETRIS (tandem elongation of templated DNA repeats for analysis of interacting proteins), a DNA barcoding platform enabling comprehensive, in situ mapping of protein interactions of different orders within cells. TETRIS uses bivalent molecular nanostructures, each recognizing a target protein and carrying a DNA barcode that fuses with neighboring units in an interaction-driven, target-agnostic manner. This approach linearizes molecular information within individual protein complexes and enables multiparametric decoding of protein constituents and spatial interactions. TETRIS was employed to map and quantify complex protein interactions and to monitor higher-order protein dynamics during cellular processes. In clinical studies, TETRIS accurately diagnosed cancer subtypes and distinguished disease aggressiveness by identifying higher-order protein interactions. Additionally, a combinatorial barcoding method is provided, wherein a system determines barcode distributions in a sample, attributes contributions of target complexes, and quantifies complex abundance via conversion based on these contributions.
Need to check novelty before this filing date? Find Prior Art

Description

PATENT Attorney Docket No.097520-1511678 Client Ref. No.2023-371-03 Methods and Compositions for Target Complex Detection And Quantification CROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority and the benefit of U.S. Provisional Application No. 63 / 661,140, filed on June 18, 2024, and U.S. Provisional Application No.63 / 661,252, filed June 18, 2024. The aforementioned provisional applications are herein incorporated by reference in their entirety for all purposes. BACKGROUND

[0001] Protein–protein interactions play a central role in mediating biological processes and their dysregulation is widely implicated in diseases1–3. Multiple proteins interact, often forming different- ordered macromolecular assemblies, to steer reactions, organize processes, and orchestrate myriad cellular activities. The study of ensemble protein interactions – the protein interactome – can provide critical insights on functional mechanisms and motivate new drug developments4,5.

[0002] Despite the importance, comprehensive mapping of complex, different-ordered protein interactions remains challenging due to limitations of current technologies6. Conventional approaches measure binary protein interactions and require genetic manipulation and / or extensive processing, thereby causing incomplete and indirect profiling. For example, resonance-energy transfer assays employ the proximity binding of proteins to provide experimental evidence7,8, but they examine only binary protein interactions, in a pairwise approach. To improve the throughput, several scalable techniques have been developed. These include yeast two-hybrid and other protein-complementation assays9,10, which determine interactions through reconstitution of reporter proteins. Such methods remain binary, are genetic and can yield considerable false results due to protein tagging and engineered environment. Alternatively, mass spectrometry-based proteomics allows high-throughput identification of interaction partners in native cell lysates. New enrichment and / or labeling processes (e.g., antibody-based affinity purification11,12, size-exclusion chromatography13or proximity labeling by engineered enzymes14) can be applied before mass spectrometry to identify various target protein interactions. Nevertheless, as the approach requires extensive sample processing, it faces several limitations. For example, affinity-based mass spectrometry tends to miss weak protein interactions which are often lost during cell lysis and remains binary in its evaluation15. SUMMARY

[0003] In one aspect, this disclosure provides a system comprising: a processor; and a memory device including instructions executable by the processor for causing the processor to perform 1 KILPATRICK TOWNSEND 797562351operations comprising: determining a distribution of barcode combinations in a sample, wherein each barcode combination of the barcode combinations includes at least one barcode of a corresponding target; determining a contribution of a target complex to the distribution of the barcode combinations; and determining, based on a conversion that uses the contribution, an abundance of the target complex in the sample.

[0004] In some embodiments, the memory device further includes instructions executable by the processor for causing the processor to determine the abundance of the target complex by performing operations comprising: determining a longest combination barcode of the barcode combinations, wherein the longest combination barcode corresponds to a largest target complex; determining, based on an amount of the longest combination barcode in the distribution and a probability of barcode formation, a first amount of each of multiple barcode sub-combinations in the longest combination barcode contributed by the largest target complex; and determining an abundance of the largest target complex in the sample based on the first amount of the largest combination barcode and the probability of barcode formation. In some embodiments, the memory device further includes instructions executable by the processor for causing the processor to determine the abundance of the target complex by perform operations comprising: removing the first amount of each of the multiple barcode sub-combinations in the longest combination barcode contributed by the largest target complex from the distribution to generate an updated distribution; determining a next longest combination barcode in the updated distribution, wherein the next longest combination barcode corresponds to a next largest target complex; determining, based on an amount of the next longest combination barcode in the updated distribution and the probability of barcode formation, a second amount of each of the multiple barcode sub-combinations in the next longest combination barcode contributed by the next largest target complex; and determining an abundance of the next largest target complex in the sample based on the second amount of the next largest combination barcode and the probability of barcode formation. In some embodiments, the operations further comprise iteratively determining abundances of target complexes until a remaining longest combination barcode includes only a single barcode. In some embodiments, the probability of barcode formation corresponds to a likelihood of a single linkage forming in a link barcode molecule. In some embodiments, the link barcode molecule incorporates a unique molecular identifier (UMI). In some embodiments, the operations further comprise determining the abundance of the largest target complex using a combinatorics equation. In some embodiments, the operations further comprise determining an interaction map of constituents in target complexes based on the abundance of the target complex. In some embodiments, a number of constituents in the target complex is in a range of 1 to 100. In some embodiments, an amount of a link barcode molecule that is generated is a function 2 KILPATRICK TOWNSEND 797562351of a barcode combination length, an abundance and a length of an originating target complex, and a probability of barcode formation. In some embodiments, one or more of targets are proteins, nucleic acids, lipids, small molecule compounds, and derivatives thereof. In some embodiments, the at least one barcode is a nucleic acid.

[0005] In another aspect, this disclosure provides a method of disease diagnosis or prognosis, wherein the method comprises: determining an abundance of a target complex in a sample with the system previously described; determining that the abundance of the target complex is different than a control greater than a threshold difference; and diagnosing a patient as having a disease or having an increased risk of developing the disease based on the abundance of the target complex being different than the control greater than the threshold difference.

[0006] In some embodiments, the control is an amount of the target complex in a normal sample.

[0007] In another aspect, this disclosure provides one or more non-transitory computer-readable storage media storing instructions that, upon execution executable by one or more processors of a system, cause the system to perform operations comprising: determining a distribution of barcode combinations in a sample, wherein each barcode combination of the barcode combinations includes at least one barcode of a corresponding target; determining a contribution of a target complex to the distribution of the barcode combinations; and determining, based on a conversion that uses the contribution, an abundance of the target complex in the sample.

[0008] In some embodiments, the operations further comprise determining the abundance of the target complex by: determining a longest combination barcode of the barcode combinations, wherein the longest combination barcode corresponds to a largest target complex; determining, based on an amount of the longest combination barcode in the distribution and a probability of barcode formation, a first amount of each of multiple barcode sub-combinations in the longest combination barcode contributed by the largest target complex; and determining an abundance of the largest target complex in the sample based on the first amount of the largest combination barcode and the probability of barcode formation. In some embodiments, the operations further comprise determining the abundance of the target complex by: removing the first amount of each of the multiple barcode sub-combinations in the longest combination barcode contributed by the largest target complex from the distribution to generate an updated distribution; determining a next longest combination barcode in the updated distribution, wherein the next longest combination barcode corresponds to a next largest target complex; determining, based on an amount of the next longest combination barcode in the updated distribution and the probability of barcode formation, a second amount of each of the multiple barcode sub-combinations in the next longest combination barcode contributed by the next largest target 3 KILPATRICK TOWNSEND 797562351complex; and determining an abundance of the next largest target complex in the sample based on the second amount of the next largest combination barcode and the probability of barcode formation. In some embodiments, the operations further comprise iteratively determining abundances of target complexes until a remaining longest combination barcode includes only a single barcode. In some embodiments, the probability of barcode formation corresponds to a likelihood of a single linkage forming in a link barcode molecule. In some embodiments, the link barcode molecule incorporates a unique molecular identifier (UMI). In some embodiments, the operations further comprise determining the abundance of the largest target complex using a combinatorics equation. In some embodiments, the operations further comprise determining an interaction map of constituents in target complexes based on the abundance of the target complex. In some embodiments, a number of constituents in the target complex is in a range of 1 to 100. In some embodiments, an amount of a link barcode molecule that is generated is a function of a barcode combination length, an abundance and a length of an originating target complex, and a probability of barcode formation. In some embodiments, one or more of targets are proteins, nucleic acids, lipids, small molecule compounds, and derivatives thereof. In some embodiments, the at least one barcode is a nucleic acid.

[0009] In another aspect, this disclosure provides a composition of N complexes identifying N targets, each of the N complexes identifying a different target, wherein each complex comprises: 1) a binding domain, wherein the binding domain identifies a target, 2) a flexible linker, and 3) an encoding polynucleotide, wherein the encoding polynucleotide comprises a modified unit modification at an internal position, wherein the encoding polynucleotide is conjugated to the flexible linker via the modified unit modification,wherein the encoding polynucleotide comprises a unique identifier that identifies the binding domain and wherein the unique identifier identifies the target, wherein the binding domains of different complexes identify different targets, wherein the unique identifiers of different complexes identify different targets, thereby the unique identifiers of the N complexes identify the N targets.

[0010] In some embodiments, N is in a range of 1 to 100. In some embodiments, the modified unit is a phosphoramidite with an amine group. In some embodiments, the flexible linker is a polyethylene glycol (PEG). In some embodiments, each of the one or more or all of the N targets are proteins, lipids, nucleic acids, small molecule compounds, and derivatives thereof. In some embodiments, the encoding polynucleotide comprises an overlap oligonucleotide. In some embodiments, the overlap oligonucleotides in all the complexes have the same nucleic acid sequence. In some embodiments, the binding domain is an antibody, a peptide, a nucleic acid, a small molecule, or an aptamer. In some embodiments, the encoding polynucleotide is unligatable. In some embodiments, the encoding polynucleotide is ligatable. In some embodiments, the overlap oligonucleotide has a length of 0 to 4 KILPATRICK TOWNSEND 797562351100 nucleotides. In some embodiments, unique identifier has a length of 5 to 100 nucleotides. In some embodiments, where internal nucleotide of each encoding polynucleotide is located at least 5 nucleotides away from the closer of the two ends of the encoding polynucleotide. In some embodiments, the linker has a length in a range from 2 nm to 100 nm, for example, in a range from 9.6 nm to 66.3 nm. In some embodiments, the PEG linker has 2-200 units.

[0011] In another aspect, this disclosure provides a method of measuring one or more of N targets in a sample, wherein N is an integer, wherein the method comprises: contacting the sample with the N complexes disclosed herein; contacting the N complexes with N complementary polynucleotides, each of the N complexes identifying a different target, wherein each complementary polynucleotide comprises a barcode being complementary to the unique identifier of one of the N complexes, and wherein the N complexes comprise N barcodes, and wherein optionally, each complementary polynucleotide comprises a UMI, and quantifying the one or more of the N barcodes, thereby measuring the abundance of the one or more of the N targets.

[0012] In some embodiments, the encoding polynucleotide in each complex comprises an overlap oligonucleotide, and each complementary polynucleotide has a universal complementary overlap, wherein the universal complementary overlap is complementary to the overlap oligonucleotide in each complex. In some embodiments, measuring the amount of the one or more of the N barcodes is by one or more of sequencing and quantitative PCR.

[0013] In another aspect, this disclosure provides a method of quantifying interactions and / or abundance of N targets in a sample, wherein N is an integer, wherein the method comprises: contacting the sample with the N complexes disclosed herein, contacting the N complexes with N complementary polynucleotides, wherein each complementary polynucleotide comprises a barcode being complementary to the unique identifier in one of the N complexes, wherein the barcode identifies the target the one of the unique identifier identifies, wherein each complementary polynucleotide has ligatable ends, wherein optionally, each complementary polynucleotide comprises a UMI. The method further comprises generating link nucleic acid molecules in a reaction mixture, wherein at least one of the link nucleic acid molecules comprising the sequences of M of N complementary polynucleotides and thus comprising M of the N barcodes, optionally separating the link nucleic acid molecules from the complexes. In some embodiments, the method further comprises quantifying the amount of the link nucleic acid molecules comprising the M barcodes, wherein M is an integer no greater than N.

[0014] In some embodiments, each complementary polynucleotide has a universal complementary overlap wherein the universal complementary overlap is complementary to the overlap 5 KILPATRICK TOWNSEND 797562351oligonucleotide in each complex. In some embodiments, the link nucleic acid molecules comprise one or more copies of at least one of the M barcodes. In some embodiments, generating link nucleic acid molecules in situ comprises ligating the complementary polynucleotides. In some embodiments, generating link nucleic acid molecules in situ comprises extension using a DNA polymerase. In some embodiments, the method further comprises quantifying one or more of the N targets, wherein for each target, quantifying the target by summing the barcode identifying the target in individual form and combination form. In some embodiments, the sample is a cell, a cell lysate, or a purified / enriched biological sample.

[0015] In another aspect, this disclosure provides a method of disease diagnosis or prognosis, the method comprising: quantifying interaction and / or abundance of N targets within a cell from a patient sample with the method described above, diagnosing the patient as having a disease or have increased risk of developing the disease if the amount of the link nucleic acid molecules comprising the M barcodes is greater than a control, wherein the targets promote the disease.

[0016] In another aspect, the disclosure provides a method of disease diagnosis or prognosis, the method comprising: quantifying interaction and / or abundance of N targets within a cell from a patient sample with the method described above, diagnosing the patient as having a disease or have increased risk of developing the disease if the amount of the link nucleic acid molecules comprising the M barcodes is lower than a control, wherein the targets are suppressors of the disease.

[0017] In some embodiments, the control is the amount of link nucleic acid molecules comprising the M barcodes in a normal sample.

[0018] In another aspect, this disclosure provides a method of monitoring disease progression in a patient, the method comprising: quantifying interaction and / or abundance of N targets within a cell from a patient sample with the method disclosed herein at a first time point and a second time point during a time period, wherein the second time point is later than the first time point, determining that the disease has progressed if the amount of the link nucleic acid molecules comprising the M barcodes at the second time point is different from the amount of the link nucleic acid molecules at the first time point.

[0019] In another aspect, this disclosure provides a method of monitoring a patient’s response to a therapy, the method comprising: quantifying interaction and / or abundance of N targets within a cell from a patient sample with the method disclosed herein at a first time point and a second time point during a time period where the patient received a therapy, wherein the second time point is later than the first time point, and determining that the patient is responsive to the therapy if the amount of the 6 KILPATRICK TOWNSEND 797562351link nucleic acid molecules at the second time point is different from the amount of the link nucleic acid molecules at the first time point.

[0020] In some embodiments, the method detects the interaction of two or more targets. In some embodiments, the N targets comprise one or more targets in FIG.4.

[0021] In another aspect, this disclosure provides a kit for quantifying interactions and / or abundance of two or more of N targets, wherein the kit comprises N complexes disclosed herein and N complementary polynucleotides, wherein N is an integer, wherein each complementary polynucleotide comprises a complementary identifier being complementary to the unique identifier of one complex, and wherein each complementary polynucleotide has ligatable ends.

[0022] In some embodiments, the polynucleotide of each complex comprises an overlap oligonucleotide, wherein each complementary polynucleotide has a universal complementary overlap that is complementary to the overlap oligonucleotide.

[0023] In another aspect, the disclosure provides a method of producing a complex of the composition disclosed above comprises: providing the binding domain and the encoding polynucleotide,optionally in a molar ratio in a range from 0.5 to 6, modifying the internal nucleotide of the encoding polynucleotide with an amino functional group, treating the encoding polynucleotide with modified internal nucleotide with an SPDP–PEG–NHS ester to produce a reactive encoding polynucleotide, treating the binding domain with sulfo-SMCC to produce a reactive binding domain, and contacting the reactive encoding polynucleotide with the reactive encoding binding domain, thereby conjugating the binding domain to the encoding polynucleotide and forming the complex. BRIEF DESCRIPTION OF THE DRAWINGS FIG.1A-1E. TETRIS for direct profiling of complex protein interactions.

[0024] FIG. 1A. (a) TETRIS assay schematics. The technology is designed to measure complex protein interactions in cells and comprises three steps. First, TETRIS units – antibodies conjugated with DNA strands through a PEG linker – are used to label proteins in cells. Each DNA is conjugated at an internal position and bears a barcode that encodes target protein identity and a universal overlap that facilitates interaction with neighboring TETRIS units. Second, DNA barcodes are formed in situ through bidirectional growth. DNA complements with 5’ phosphorylation are added. Each complement bears a unique complementary identifier, to record protein identity, and a repetitive overhang, to hybridize with neighboring TETRIS units. This addition thus enables bidirectional chaining of adjacent TETRIS units, in an interaction-driven and target-agnostic manner. Upon DNA ligation of the complements, TETRIS barcodes that explicitly capture complex protein interactions (i.e., molecular constituents and spatial interactions) are covalently formed. Finally, barcodes are 7 KILPATRICK TOWNSEND 797562351liberated from protein-bound TETRIS units and their full distribution analyzed to decode different- ordered protein interactions. FIG. 1B. (b) TETRIS analysis workflow to decode complex protein interactions. As identical protein complexes can generate a mixture of fully- and partially-formed barcodes, and the same barcodes are produced from different-ordered protein interactions (i.e., heterogeneous and degenerate barcoding), we perform large-scale combinatorics transformation of the generated barcode distribution to decode complex protein interactions. TETRIS barcodes are sequenced, clustered and transformed to quantify the originating protein complexes and map different-ordered protein interactions (see Figures 8 and 20 for details). Protein interactions are presented as network graphs to illustrate the abundance and composition of different protein complexes. Protein markers are defined as peripheral nodes, protein complexes as central nodes, and each confirmed TETRIS interaction in a protein complex as a line. FIG. 1C. (c) Schematics of different-ordered protein interactions in cancer and control cells. TETRIS enables direct and comprehensive mapping of different-ordered protein interactions in cells. While both cancer and control specimens show similar amounts of lower-order protein complexes (e.g., 2-mer protein complexes), cancer cells show more diverse and more abundant higher-order protein complexes (e.g., 4-mer protein complexes) than control cells. FIG.1D. (d) Clinical applications. We applied TETRIS to evaluate the abundance of individual proteins and complex protein interactions in biopsy samples of breast cancer patients. Protein interactions were profiled in situ, across different subcellular compartments in whole cells. As compared to individual proteins and lower-order protein interactions, the TETRIS-identified higher-order protein complexes can more effectively distinguish patient disease aggressiveness. FIG.1E. (e) Terminology used in describing the TETRIS technology. FIG.2A-2F. Assay design and evaluation.

[0025] FIG. 2A. (a) Bidirectional extendibility of barcode units. Sticky- and blunt-end DNA barcode units were ligated and analyzed through gel electrophoresis. While the blunt-end units generated only short dimer products, the optimized universal overlaps by the sticky-end units formed much longer barcode products. FIG.2B. (b) Spatial activity profiles of TETRIS units. TETRIS units prepared with short or long PEG linkers were used to measure interacting proteins (humanized anti- EGFR antibody (anti-EGFR) and EGFR with C-terminal DDK tag (EGFR-DDK)) and non- interacting proteins (humanized anti-HER2 antibody (anti-HER2) and EGFR-DDK). Protein separation was determined from published values67. To vary the distance between the non-interacting proteins, we prepared mixtures containing different amounts of BSA spacer molecules before protein immobilization onto high-binding plates and used the protein hydrodynamic sizes to estimate average separation between the non-interacting protein pair. The immobilized proteins were treated with TETRIS units (anti-human IgG1 Fc, and anti-DDK) and the generated barcodes were analyzed by 8 KILPATRICK TOWNSEND 797562351TaqMan assay. The short-PEG TETRIS units formed barcodes only when the target proteins were interacting, and showed a sharp signal decrease with increasing protein separation. FIG. 2C. (c) Detection sensitivity. TETRIS measurements were performed by titrating a known quantity of model protein complexes (humanized anti-EGFR–EGFR-DDK) assembled on polystyrene beads and measuring the resultant barcodes. For comparison, proximity extension assay (PEA) and sandwich enzyme-linked immunosorbent assay (ELISA) were performed with a pair of antibodies targeting distinct protein components in the model protein complex (anti-human IgG1 Fc and anti-DDK). TETRIS achieved ~150-fold and ~1000-fold improvement in its limit of detection (LOD) as compared to PEA and ELISA, respectively. The LOD is defined as the intrapolated concentration at 3× s.d. of a no-target control (indicated by the dotted line for TETRIS). FIG. 2D. (d) Correlation between TETRIS and ELISA. TETRIS correlated well with gold-standard ELISA (R2= 0.9875). FIG. 2E. (e) Experimental validation of barcode distribution. TETRIS barcodes were generated from pure protein complexes assembled on polystyrene beads (two-protein complex: anti-EGFR–EGFR-DDK; and three-protein complex: anti-EGFR–EGFR-DDK–anti-DDK). The experimental barcode distribution matched that by numerical prediction derived of known quantities of protein complexes (see Figure 20 for derivation details). FIG. 2F (f) Measurement of protein complex mixtures. Mixtures comprising known ratios of protein complexes (two-protein complex: anti-EGFR–EGFR- DDK; three-protein complex: anti-EGFR–EGFR-DDK–anti-DDK) were subjected to the TETRIS assay and the generated DNA barcodes were analyzed by TaqMan assay (top). Using combinatorics transformation that accounts for different barcode combinations and their barcode formation probabilities, the measured DNA barcode signals (top) were converted to protein complex signals (bottom). Note that incomplete DNA barcodes were detected due to ligation inefficiency. All measurements were performed in biological triplicate and the data are presented as mean ± s.d. Measurements were normalized against that of IgG isotype control antibodies in b–f. The data were normalized to the mean of the maximum signal in a–d, and to the mean of the highest-ordered barcode signal in e. PEA, proximity extension assay. ELISA, enzyme-linked immunosorbent assayAll measurements were performed in biological triplicate and the data are presented as mean ± s.d. Measurements were normalized against that of IgG isotype control antibodies in b–f. The data were normalized to the mean of the maximum signal in a–d, and to the mean of the highest-ordered barcode signal in e. PEA, proximity extension assay. ELISA, enzyme-linked immunosorbent assay. FIG.3A-3E. TETRIS measurement of cellular protein complexes.

[0026] FIG.3A. (a) Structure of EGFR–GRB2–SOS1, a natural three-protein complex known to have unique interactions that can be activated upon EGFR activation. FIG.3B. (b) Measurement of cellular protein interactions and individual protein abundances. Human cancer cells (MDA-MB-231 9 KILPATRICK TOWNSEND 797562351and H1975) were treated with EGF (50 ng / ml, 2 min) to induce protein interaction changes. The protein complexes in treated and control cells were measured directly in fixed cells by TETRIS (top) or in cell lysates by conventional cross-link co-immunoprecipitation (co-IP) (bottom). FIG.3C. (c) Reliable TETRIS detection. TETRIS correlated well with gold-standard cross-link co-IP for the measurements of individual proteins and protein complexes (R2 = 0.9851). FIG.3D. (d) Treatment- induced composition changes in protein interactions. Using an adapted TETRIS assay that enables simultaneous analysis of target proteins and protein interactions, we measured changes in the total abundance of EGFR (top) and p53 protein (bottom) and their respective protein interaction composition. “Others” refer to the amount of EGFR (top) or p53 (bottom) that are present as monomeric proteins or complexed with other proteins not under investigation here. FIG. 3E. (e) Responsiveness of higher-order protein interactions. In both the EGFR (top) and p53 (bottom) systems, across all cell lines tested, the three-protein complexes showed the most significant fold increases upon cellular stimulation. All measurements were performed in triplicate and normalized against that of IgG isotype control antibodies. The data are presented as mean ± s.d. in b–d, and as row-normalized mean in e. Co-IP, co-immunoprecipitation. FIG.4A-4D. Complex protein interactions of different orders.

[0027] FIG. 4A. (a) Diversity of complex protein interactions of different orders. We employed the massively-parallel TETRIS approach to evaluate interactions among 34 proteins in cancer (BT474) and control (MCF10A) cells. Detailed processing is presented in Figure 9. For each cell line, four interaction maps were constructed based on different-ordered protein interactions (2-mer, 3-mer, 4-mer, and 5-mer protein complexes, respectively). Notably, the cancer cells demonstrate more diverse higher-order protein complexes (3-mer and above). All protein interactions are presented as network graphs, by defining protein markers as peripheral nodes, protein complexes as central nodes, and each confirmed TETRIS interaction in a protein complex as a line. Each interaction map illustrates the top 10% most abundant protein complexes identified. FIG. 4B. (b) Abundance of individual proteins and different-ordered protein complexes. We performed pairwise comparison of different proteins and protein complexes to evaluate their relative amounts in cancer (BT474) vs. control (MCF10A) cells. While the cancer cells showed similar amounts of individual proteins or low-order protein complexes (2-mer protein complexes) with respect to the control cells, they consistently showed more abundant higher-order protein complexes (3-mer and ≥4-mer protein complexes) across different complexes identified, which were formed by two or more of the targets selected from the group consisting of GRB2, HER2, SHC, ER, EGFR, FOXA1, LMO4, GATA3, PR, BRCA1, TIF2, VAV2, SCRIB, NOS1AP, PTPN1, SRC, VANGL1, WASF3, NCKAP1, CYFIP1, AP2M1, ERRFI1, UBS3B, P85B, c-Met, HER3, SPRY2, CBL, PLCG1, STAT3, RAF1, PTPN11, 10 KILPATRICK TOWNSEND 797562351CtIP, and SOS1. Proteins and protein complexes are sorted as a waterfall plot based on their measured abundance in the cancer cells. FIG. 4C. (c) Comparison of complex abundance. As compared to lower-order protein complexes, higher-order protein complexes demonstrated a significantly higher fold difference (p < 1×10–10) in the cancer cells relative to the control cells. P values were determined by two-tailed unpaired Student’s t-test. FIG.4D. (d) Breast cancer subtyping. By profiling ER, PR and HER2-based protein interactions, TETRIS accurately identified the molecular subtypes of different breast cancer cell lines. BT474 (HER2+, ER+, PR+), SKBR3 (HER2+, ER–, PR–), MDA- MB-231 (HER2–, ER–, PR–) and MCF7 (HER2–, ER+, PR+). All measurements were performed in triplicate and normalized against that of IgG isotype control antibodies. The data are presented as mean in a and c, as mean ± s.d. in b, and as column-normalized mean in d. N.D., not detected. FIG.5A-5E. Clinical profiling of multiprotein complexes.

[0028] (a) TETRIS analysis of multiprotein complexes in clinical samples. Patient-matched cancer, border and control fine needle aspiration biopsies (n = 48) were obtained from breast cancer patients. TETRIS was applied to profile protein markers and their different-ordered protein complexes. (b) Receiver operating characteristic (ROC) curves. TETRIS-derived logistic regression models demonstrated accurate cancer detection. (c) Breast cancer subtyping. Based on their clinical pathology classification, cancer samples were identified as luminal (ER+or PR+), non-luminal HER2- positive (ER–, PR–, HER2+) and triple-negative (ER–, PR–, HER2–). TETRIS profiling of ER, PR and HER2 complexes accurately classified the patient cancer samples into distinct molecular subtypes. (d) TETRIS differentiation of disease aggressiveness. We measured the abundance of individual proteins and protein interactions in cancer aspiration biopsies (n = 24) and computed distinct composite scores through cross-trained regression models for individual proteins, 2-mer and 3-mer protein complexes, respectively. Disease aggressiveness was determined through clinical evaluation of tissue morphology (clinical grade 3, more aggressive; clinical grade 1 and 2, less aggressive). (e) Statistical analyses of the composite scores. In each pair of the bar graphs for which the composite scores are compared, the bar on the left represents a more aggressive cancer, and the one on the right represents a less aggressive cancer. The composite score derived from 3-mer protein complexes showed the best ability to distinguish disease aggressiveness, while that derived from individual proteins could not. The boxes indicate the 25th–75thpercentiles and the whiskers denote the maximum and minimum. P values were determined by two-tailed unpaired Mann–Whitney U test. All measurements were performed in triplicate and normalized against that of IgG isotype control antibodies. AUC, area under the curve. 11 KILPATRICK TOWNSEND 797562351FIG.6A-6E. The TETRIS unit.

[0029] (a) Schematic representation of the TETRIS unit preparation. Single-stranded DNA modified with amino functional groups at an internal position are reacted with SPDP–PEG–NHS ester and the products are activated by reducing with tris(2-carboxyethyl)phosphine (TCEP). Antibodies are activated with sulfosuccinimidyl 4-(N-maleimidomethyl)cyclohexane-1-carboxylate (sulfo- SMCC). The activated antibodies and PEG-modified DNA are mixed to form the hybrid TETRIS units. (b) Complements for bidirectional barcode growth. To generate TETRIS barcodes, single- stranded complementary DNA (complements) are added to bind and chain adjacent TETRIS units. Each complement contains a unique complementary identifier, which hybridizes with protein- encoding sequences of TETRIS units, and a repetitive complementary overhang, which binds to the universal overlap of adjacent TETRIS units. (c) Enzymatic and chemical ligation. For enzymatic ligation, complements are linked by T4 DNA ligase through phosphodiester linkage. For chemical ligation, complements are linked by click chemistry through triazole linkage. TETRIS units were mixed with appropriate DNA complements to achieve sticky-end and blunt-end ligations. Ligations were performed enzymatically or chemically. Arrows indicate bidirectional barcode growth. (d) Interaction-driven barcode growth. TETRIS units (anti-human IgG1 Fc, anti-DDK) were closely- spaced through model interacting proteins (humanized anti-EGFR antibody and EGFR with C- terminal DDK tag; anti-EGFR–EGFR-DDK), and distantly-spaced through non-interacting proteins (humanized anti-HER2 antibody and EGFR with C-terminal DDK tag) and BSA spacer molecules. Closely-spaced TETRIS units formed barcodes while distantly-spaced TETRIS units formed negligible barcodes. Barcode signals were measured by TaqMan assays. P values were determined by two-tailed unpaired Student’s t-test. (e) Barcode recovery optimization. Water (H2O), 100 mM hydrochloric acid (HCl) or 100 mM sodium hydroxide (NaOH) was added to liberate barcodes from protein-bound TETRIS units. Sodium hydroxide demonstrated the best recovery and was applied for subsequent experiments. P values were determined by one-way ANOVA. All measurements were performed in biological triplicate and normalized to the mean of the maximum signal. The data are presented as mean ± s.d. FIG.7. TETRIS assay workflow.

[0030] Model protein complexes (immobilized on a solid support) or biological protein complexes (in fixed and permeabilized cells) are incubated with TETRIS units (antibody–PEG–ssDNA, 3 μg / ml) for 2 h at room temperature or 1 h at 4 °C, respectively. Unbound TETRIS units are then removed, before incubating the protein complexes with DNA complements (ssDNA complements, 500 nM) for 1 h at room temperature. After washing to remove unbound DNA complements, T4 DNA ligase (20,000 U / ml) is introduced to the sample and incubated for 10 min at room temperature to enable 12 KILPATRICK TOWNSEND 797562351complement ligation into DNA barcodes. Following enzyme inactivation (65 °C for 10 min), sodium hydroxide (NaOH, 100 mM) is added to liberate DNA barcodes from the bound TETRIS units. The DNA barcode solution is then collected, purified through a desalting column to remove NaOH, and counted through qPCR or NGS. The measured DNA barcode distribution is finally analyzed through combinatorics data transformation to decode the composition and abundance of the originating protein complexes. FIG.8A-8B. Illustration of TETRIS combinatorics transformation.

[0031] (a) Barcode generation from a known amount of protein complexes. For a sample mixture containing protein A (or Target 1 or T1), complexes AB (or T1T2) and ABC (or T1T2T3) at 1:1:1 ratio, the overall barcode distribution can be derived by summing the contributions by different protein complexes. Note that for any specific protein complex, the amount of its fully-formed TETRIS barcodes reflects the abundance of the originating protein complex. (b) Decoding of protein complexes from an ensemble measured barcode distribution. The measured barcode distribution is apportioned, according to the projected contributions by different originating protein complexes, as determined by the amount of fully-formed barcodes in an iterative approach. The apportioned barcode signals are then used to determine the abundance and composition of protein complexes in the heterogeneous mixture. FIG.9. High-throughput multiplexed TETRIS workflow.

[0032] The integrated workflow consists of four key steps. In the first step, TETRIS barcodes are generated from complements bearing unique molecular identifiers (UMIs). In the second step, the UMI-incorporated barcodes are subjected to magnetic bead-based purification (see Figure 29 for details) to enrich for long barcodes. Following sequencing library preparation, the barcodes are characterized by gel electrophoresis (see Fig. 30 for details). Size control oligonucleotides are included throughout the selection process (see Fig. 31 for details) to account for enrichment efficiency. Following next-generation sequencing, amplified barcode reads are deduplicated according to their UMIs. In the third step, the measured barcode distribution is transformed to decode the composition and abundance of originating protein complexes. In the final step, the protein interaction data are used to construct multiprotein interaction maps and visualized as network graphs. The peripheral circles indicate the identity and amount of interacting proteins. The central nodes indicate the amount of protein complexes, and their branching lines denote confirmed TETRIS interactions by the constituent proteins. 13 KILPATRICK TOWNSEND 797562351FIG.10A-10C. Different-ordered protein complexes in cancer cells.

[0033] (a) Diversity of protein interactions in cells. The massively-parallel TETRIS approach was applied to evaluate complex interactions among 34 proteins in breast cancer cell lines (SKBR3, MDA-MB-231 and MCF7). For each cell line, four interaction maps were constructed based on the 2-mer, 3-mer, 4-mer and 5-mer protein complexes, respectively. Each map illustrates the top 10% most abundant protein complexes identified. (b) Abundance of individual proteins and different- ordered protein complexes in cells. We performed pairwise comparison of different proteins and protein complexes to evaluate their relative amounts in cancer vs. control cells. While the cancer cells showed similar amounts of individual proteins or low-order protein complexes (2-mer protein complexes) with respect to the control cells, they showed more abundant higher-order protein complexes (3-mer and ≥4-mer protein complexes) across different complexes identified, which were formed by two or more of the targets selected from the group consisting of GRB2, HER2, SHC, ER, EGFR, FOXA1, LMO4, GATA3, PR, BRCA1, TIF2, VAV2, SCRIB, NOS1AP, PTPN1, SRC, VANGL1, WASF3, NCKAP1, CYFIP1, AP2M1, ERRFI1, UBS3B, P85B, c-Met, HER3, SPRY2, CBL, PLCG1, STAT3, RAF1, PTPN11, CtIP, and SOS1. Proteins and protein complexes are sorted as a waterfall plot based on their measured abundances in the cancer cells. (c) Examples of protein complexes with decreased abundance in cancer cells compared to control cells. Some protein complexes (e.g., PR–SOS1, LMO4–CYFIP1, TIF2–UBS3B) demonstrated significantly lower abundance in cancer cells than in control cells. PR–SOS1 represents a complex comprising PR and SOS1. LMO4–CYFIP1 represents a complex comprising LMO4 and CYFIP1. TIF2–UBS3B represents a complex comprising TIF2 and UBS3B. In each pair of the bar graphs for which the signals are compared, the one on the left is from a cancer sample, while the one on the right is from a control sample. Values were determined by two-tailed unpaired Student’s t-test. All measurements were performed in technical triplicate and normalized against that of IgG isotype control antibodies. The data are presented as mean in a and as mean ± s.d. in b and c. FIG.11. TETRIS sensitivity and specificity characteristics.

[0034] Comparison of TETRIS performance with published binary protein interaction assays73,76-78. PRS, positive reference set. NRS, negative reference set. FIG.12. Clinical analysis of individual proteins and different-ordered protein complexes in cancer and matched control patient samples.

[0035] Abundance comparisons of individual proteins, 2-mer and 3-mer protein complexes in cancer vs. matched control patient samples. Even for individual protein markers (e.g., WASF3, NCKAP1 and CYFIP1) that showed a similar abundance in cancer and control samples, protein 14 KILPATRICK TOWNSEND 797562351complexes comprising the same proteins (e.g., WASF3–NCKAP1, WASF3–CYFIP1, WASF3– NCKAP1–CYFIP1) could distinguish the samples. Signal levels were normalized with respect to the mean marker abundance by the matched control samples. P values were determined by two-tailed unpaired Mann–Whitney U test. All measurements were performed in biological triplicate and normalized against that of IgG isotype control antibodies. FIG.13. Clinical analysis of individual proteins and different-ordered protein complexes in cancer patient samples of different disease aggressiveness.

[0036] Abundance comparisons of individual proteins, 2-mer and 3-mer protein complexes in more aggressive (grade 3) vs. less aggressive (grade 1–2) cancer samples. While individual protein abundances universally failed to differentiate disease aggressiveness, protein interactions comprising the same proteins (2-mer and 3-mer protein complexes) demonstrated better abilities to classify the clinical groups. Signal levels were normalized with respect to the mean marker abundance by the less aggressive cancer patient samples. P values were determined by two-tailed unpaired Mann–Whitney U test. All measurements were performed in biological triplicate and normalized against that of IgG isotype control antibodies. FIG.14A-14B. Universal overlap optimization.

[0037] (a) Overlap optimization. Duplex double-stranded DNA (dsDNA) units with universal overlaps of varying lengths (0–10 nucleotides) were enzymatically ligated at different DNA concentrations (high concentration: 100 nM, low concentration: 5 nM) to evaluate barcode formation at different inter-unit spacing. Ligation yields (left) were obtained from denaturing polyacrylamide gel electrophoresis (right). Note that only the top strand is ligatable (with 5’ phosphate) while the bottom strand is non-ligatable (without 5’ phosphate). Zone I of the gel image denotes the region containing the reactants and Zone II denotes the region containing the ligation products. Different- zoned intensities were used to calculate the ligation reaction yield as below.

[0038] An overlap length of 6 nucleotides was chosen for subsequent experiments as it generated high barcode signal with closely-spaced units (high concentration) and low background ligation with far-spaced units (low concentration). (b) Extendibility of sticky- and blunt-end DNA units. Sticky- and blunt-end DNA units were ligated and examined through denaturing gel electrophoresis. While the blunt-end duplex units generated only dimer barcodes, the sticky-end units formed much longer 15 KILPATRICK TOWNSEND 797562351barcodes. Measurements were performed in biological triplicate and the data are presented as mean ± s.d. FIG.15A-15C. Density tuning of TETRIS units.

[0039] Density of DNA strands on TETRIS units. TETRIS units were prepared by reacting antibodies with varying molar excess of PEG-modified DNA strands. The DNA strands were modified with 6-FAM to enable DNA density measurements. (b) Real-time binding kinetics of unmodified antibody and TETRIS units with varying density of DNA strands. (c) KD values of unmodified antibody and TETRIS units with varying density of DNA strands. KDvalues were obtained from the real-time binding kinetics. All measurements were performed in biological triplicate and the data are presented as mean ± s.d. in a and c. The data were normalized to the mean of the maximum signal in b. FIG.16A-16B. Barcode generation by intra- and inter-antibody ligation.

[0040] TETRIS units bearing varying density of DNA strands were differentially spaced to evaluate (a) intra-antibody ligation and (b) inter-antibody ligation. Intra-antibody ligation refers to the formation of barcodes from the same TETRIS unit, while inter-antibody ligation refers to barcode formation from distinct neighboring units. Exclusive intra-antibody ligation was obtained by spatially separating the TETRIS units with large amounts of BSA spacers during plate adsorption. Inter- antibody ligation signals were obtained by subtracting the normalized total ligation signal of closely- spaced TETRIS units with the respective normalized intra-antibody ligation signal. A conjugation density of 0.97 (~1) DNA strand per antibody was chosen for subsequent experiments as it generated low intra-antibody ligation and high inter-antibody ligation signals. All measurements were performed in biological triplicate and normalized against the amount of antibodies. The data were normalized to the mean of the maximum signal and are presented as mean ± s.d. FIG.17A-17B. Size tuning of TETRIS units.

[0041] a. Hydrodynamic diameter of unmodified antibody and TETRIS units with varying length of PEG linkers. Measurements were performed by dynamic light scattering analysis. P values were determined by two-tailed unpaired Student’s t-test. b. Optimization of DNA length. TETRIS units conjugated with different-sized DNA strands were used to measure model protein complexes (humanized anti-EGFR–EGFR-DDK). The TETRIS unit bearing 46-nucleotide DNA oligonucleotide demonstrated the best signal-to-noise ratio and was used for subsequent experiments. All measurements were performed in biological triplicate and the data are presented as mean ± s.d. Measurements were normalized against that of IgG isotype control antibodies, and then to the mean of the maximum signal in b. 16 KILPATRICK TOWNSEND 797562351FIG.18. TETRIS DNA hybridization conditions.

[0042] We evaluated various hybridization conditions for TETRIS DNA complements. A low amount of model interacting proteins (anti-EGFR–EGFR-DDK, 73 amol) were first labeled with specific antibody–DNA conjugates (TETRIS units), before being incubated with different DNA tags (on-target DNA complements as well as off-target DNA complements). DNA hybridization was performed in different buffers and at different annealing temperatures. After ligation of the TETRIS DNA barcodes, we measured the resultant barcode formation to evaluate performance by the various hybridization conditions. 1× annealing buffer at 20 °C (room temperature) demonstrated the best signal-to-noise ratio and was applied for subsequent experiments. P values were determined by two- tailed unpaired Student’s t-test. All measurements were performed in biological triplicate and normalized against that of IgG isotype control antibodies. In each pair of the bar graphs for which signals are compared, the bar on the left represents on-target DNA complement, and the one on the right represents off-target DNA complement. The data were normalized to the mean of the maximum signal and are presented as mean ± s.d. FIG.19. Low TETRIS signals in various negative controls.

[0043] Using the model interacting proteins (humanized anti-EGFR antibody and EGFR with C- terminal DDK tag; anti-EGFR–EGFR-DDK), we performed the following positive and negative control experiments. In the positive set, respective TETRIS units (anti-human IgG1 Fc and anti-DDK) were used. Various negative controls were performed accordingly: 1) IgG isotype control: TETRIS units prepared from IgG isotype control antibodies were applied for the measurement.2) Free DNA: the TETRIS units were replaced with unconjugated antibodies and unlinked oligonucleotides.3) No- target control: the target proteins were absent in the sample. 4) Non-interacting proteins: non- interacting proteins (humanized anti-HER2 antibody and EGFR with C-terminal DDK tag) were plated and measured with their respective TETRIS units. 5) Epitope-blocked control: despite target protein interaction (anti-EGFR–EGFR-DDK), the binding site of one of the TETRIS units used (anti- EGFR, clone 528) was blocked by this protein interaction. All measurements were performed in biological triplicate and normalized against that of IgG isotype control. The data are presented as mean ± s.d. FIG.20. Mathematical model for TETRIS barcode generation.

[0044] To build a mathematical model of TETRIS DNA barcode generation from different-ordered protein complexes, we consider the barcode formation probability (p), which accounts for TETRIS unit binding, complement hybridization and ligation of closely-spaced TETRIS units, and is determined experimentally. A given amount of identical protein complexes can thus generate a 17 KILPATRICK TOWNSEND 797562351mixture of fully-and partially-formed barcodes; the amount of respective barcodes generated is a function of the barcode length, the originating protein complex abundance and p. We develop equations to quantify the distribution of DNA barcode generation, from single proteins, two-protein and three-protein complexes, respectively, and further generalize the equation for different-ordered protein complexes (any given length of protein complexes). These equations serve as the basis for the development of the combinatorics transformation that converts an ensemble measured TETRIS barcode distribution into different-ordered protein interactions of origin, even in heterogeneous mixtures of protein complexes. FIG.21A-21B. Model protein complexes.

[0045] (a) Experimental characterization of model protein complexes. The model three-protein complex (anti-EGFR–EGFR-DDK–anti-DDK) was formed on a plate surface using EGFR–DDK, anti-DDK antibody (anti-DDK) and anti-EGFR antibody (anti-EGFR) that were fluorescently-labeled with Alexa Fluor 405, 488 and 647, respectively. Fluorescence signals were measured at different time points after complex formation and respective protein amounts were determined by intrapolating standard concentration curves. P values were determined by one-way ANOVA. (b) Barcode formation efficiency. Model two-protein complexes (anti-EGFR–EGFR-DDK) were assembled on a plate and quantified spectroscopically. The prepared complexes were subjected to the TETRIS assay and the generated barcodes were quantified by qPCR to determine the barcode formation efficiency. As a negative control, a comparable mixture of non-interacting proteins, namely humanized anti- HER2 antibody (anti-HER2) and EGFR-DDK, were used and subjected to the TETRIS assay (TETRIS units: anti-human IgG1 Fc, anti-DDK). The measured barcode formation efficiency was applied to the subsequent experiments. All measurements were performed in biological triplicate. Measurements were normalized against that of IgG isotype control antibodies b. The data were normalized to the mean of the maximum signal and are presented as mean ± s.d. FIG.22. Amplification efficiencies of TETRIS barcodes.

[0046] TaqMan qPCR calibration curves of TETRIS barcodes with their respective probe and primer sets. The slope, PCR efficiency calculated from the slope and goodness of fit (R2) are presented for each barcode. All sets show >90% efficiency. Sequences of the barcodes and their respective probe and primer sets are presented in Table 2. All measurements were performed in biological triplicate and the data are presented as mean ± s.d. FIG.23A-23B. Experimental validation on natural protein complexes.

[0047] Human cancer cells (MDA-MB-231 and H1975) were treated with EGF (50 ng / ml, 2 min) to induce protein interaction changes. (a) Measurement of cellular individual protein abundance. 18 KILPATRICK TOWNSEND 797562351Individual protein abundance of the treated and control cells was measured directly in fixed cells by TETRIS (top) or in cell lysates by conventional cross-link co-immunoprecipitation (co-IP) (bottom). Both the TETRIS and co-IP measurements showed that the EGF stimulation did not alter the abundance of individual proteins. P values were determined by two-tailed unpaired Student’s t-test. (b) TETRIS DNA barcode signals before data transformation. The measured DNA barcode distribution was analyzed through combinatorics data transformation to decode the composition and abundance of protein complexes and individual proteins. All measurements were performed in biological triplicate and normalized against that of IgG isotype control antibodies. In each pair of the bar graphs for which the signals are compared, the bar on the left represents control sample, and the one on the right represents EGF-treated sample. The data were normalized to the mean of the maximum signal in b. The data are presented as mean ± s.d. FIG.24A-24B. Immunofluorescence images of target protein markers in cytoplasm and nucleus.

[0048] (a) Single-channel and merged fluorescence images of EGF-treated A431 cells. The cells were stained with antibodies against EGFR and its complex partners, GRB2 and SHC, indicating the protein markers’ cytoplasmic localization. (b) Single-channel and merged fluorescence images of oxaliplatin-treated HCT116 cells. The cells were stained with antibodies against p53 and its complex partners, MDM2 and MDMX, showing the protein markers’ nuclear localization. All cells were counterstained with nuclear dye Hoechst 33342. FIG.25A-25E. Cellular response to targeted treatment.

[0049] (a) Western blotting analysis of EGF treatment. EGF-treated and control A431 cells were lysed and immunoblotted for EGFR signaling pathway markers (p-EGFR, p-SHC, p-Gab1, p-PLCγ1, p-p44 / 42 MAPK, p-Akt) and GAPDH as a loading control. (b) Flow cytometry analysis of EGF treatment. A431 cells treated with 50 ng / ml of EGF for 2 min at 37 °C and control A431 cells were lysed and immunocaptured using polystyrene beads conjugated with anti-EGFR antibodies. The beads were then stained with antibodies against EGFR and phosphotyrosine. In each pair, the bar on the left represents a control sample, and the one on the right represents an EGF-treated sample. (c) Flow cytometry analysis of oxaliplatin treatment. HCT116 cells treated with 20 μM oxaliplatin for different durations at 37 °C and control HCT116 cells were stained with antibodies against p53, MDM2 and MDMX. (d) Gating strategy for flow cytometry analyses. (e) Full-length unprocessed blots for a. All measurements were performed in biological triplicate and normalized against that of IgG isotype control antibodies. The data are presented as mean ± s.d. in b and c. 19 KILPATRICK TOWNSEND 797562351FIG.26A-26B. Time-dependent protein interaction changes.

[0050] (a) In cells undergoing targeted treatments, TETRIS was applied at different time points to measure interaction changes of EGFR protein (in A431 cells treated with EGF; top) and p53 protein (in HCT116 cells treated with oxaliplatin, bottom). (b) In both systems, the higher-order three-protein complexes showed the earliest and most significant fold increases compared with untreated control samples. P values were determined by two-tailed unpaired Student’s t-test. All measurements were performed in triplicate and normalized against that of IgG isotype control antibodies. The data are presented as mean ± s.d. FIG.27A-27B. Protein complex signals at different absolute protein levels.

[0051] (a) Measurements of total EGFR and EGFR–SHC complex by TETRIS and conventional assays. In four cell lines known to have different EGFR expression levels (H1975, MDA-MB-231, HCC827 and A431), we applied TETRIS to measure their total EGFR abundance and EGFR–SHC complex signals. The TETRIS measurements correlated well with conventional co- immunoprecipitation measurements. The levels of protein complexes (EGFR–SHC) measured by TETRIS do not follow the trend of the absolute protein levels (total EGFR) and corresponded well with gold-standard co-immunoprecipitation. (b) TETRIS measurements of total EGFR and EGFR– SHC–GRB2 complex during cellular stimulation. We applied TETRIS to measure total EGFR expression and EGFR–SHC–GRB2 signals in untreated cells (0 min) and at different time points after 50 ng / ml EGF treatment (2 min and 10 min). While there were no apparent changes in total EGFR expression for each cell line, TETRIS captured dynamic increase and decrease in EGFR–SHC–GRB2 complex levels. All measurements were performed in biological triplicate and normalized against that of IgG isotype control antibodies. The data were normalized to the mean of the maximum signal and are presented as mean ± s.d. FIG.28A-28C. TETRIS units for next-generation sequencing.

[0052] (a) Modified TETRIS unit design for next-generation sequencing. As compared with the original design, unique molecular identifiers (UMIs) were built into the complements for accurate bioinformatic identification of PCR duplicates. The identifier was thus shortened to 20 nucleotides, to adapt the barcode length for next-generation sequencing. The modified design was employed to measure plated (b) model two-protein and (c) model three-protein complexes and demonstrated comparable performance with the original design (without UMI). P values were determined by two- tailed unpaired Student’s t-test. All measurements were performed in biological triplicate and normalized against that of IgG isotype control antibodies. The data were normalized to the mean of the maximum signal and are presented as mean ± s.d. 20 KILPATRICK TOWNSEND 797562351FIG.29. Barcode size enrichment.

[0053] Equimolar amounts of all size controls (2-mer, 3-mer, 4-mer and 5-mer size controls) were mixed and subjected to size selection with SPRIselect beads. The beads were applied at different ratios to evaluate their size-enrichment performance. After elution, the recovery of each size control was determined by qPCR. A bead-to-sample ratio of 5.5 was chosen for size selection to enrich for long barcodes. All measurements were performed in biological triplicate, and the data are presented as mean ± s.d. FIG.30. Barcode size characterization.

[0054] To validate the library preparation, the barcode samples were analyzed through Bioanalyzer. Samples were size-selected by polyacrylamide gel electrophoresis before running on the Bioanalyzer. A standard sample containing all size controls was run concurrently as reference. The data were normalized to the maximum signal. FIG.31A-31C. Size controls.

[0055] a. Amplification efficiencies of size controls. SYBR qPCR calibration curves of size controls with their respective primer sets. The slope, PCR efficiency calculated from the slope and goodness of fit (R2) are presented for each size control. All sets show >88% efficiency. b. Specificity of primer sets for size controls. Each size control was subjected to SYBR qPCR analysis with all four primer sets to examine the signal specificity. Negligible crosstalk was observed. Sequences of the size controls and their respective primer sets are presented in Table 6. c. Recovery of size controls. Size controls were enriched and quantified by qPCR. All measurements were performed in biological triplicates, and the data are presented as mean ± s.d in a and c and as row-normalized mean in b. FIG.32A-32B. Barcode distribution combinatorics.

[0056] TETRIS DNA barcodes generated from a. human breast cancer cell line BT474, and b. control breast cell line MCF10A were analyzed, sorted according to their length (e.g., 5-mer, 4-mer and 3-mer DNA barcodes) and sequence identity. These barcodes were then used to map to different- ordered protein complexes of origin. The mapping of DNA barcodes to protein complexes reflects heterogeneous and degenerate barcoding, that identical protein complexes can generate a mixture of fully-formed and partially-formed TETRIS barcodes and identical barcodes can be produced from different protein complexes. As compared to the control cell line, the cancer cell line measured more abundant and more diverse higher-order protein complexes (e.g., 4-mer and 5-mer complexes). 21 KILPATRICK TOWNSEND 797562351FIG.33. Comparisons with public databases.

[0057] For the top 10% most abundant two-protein complexes identified by TETRIS, 38–44% of the complexes are represented in public databases (PINA68, STRING69, BioGRID70, IntAct71and CancerNET72). The TETRIS performance is comparable with other large-scale protein interaction studies, such as the development of new affinity-based mass spectrometry technologies, where ~22% of their high-confidence protein complexes identified were represented in public databases12. PPI, protein–protein interaction. FIG.34. Positive and negative reference sets.

[0058] The high-confidence (HC) and general (G) positive reference set (PRS) and negative reference set (NRS) were constructed against known interactions available in five public databases (PINA68, STRING69, BioGRID70, IntAct71and CancerNET72) and a recent protein interaction database developed specifically for breast cancer by Kim et al.12FIG.35. Overlaps and differences of the detected protein–protein interactions in BT474 and MCF10A cells.

[0059] The TETRIS-identified protein–protein interactions (PPIs) were compared between the two cell lines. More overlaps than differences were observed in the identified PPIs. This is consistent with known cell line features that both cell lines are derived from breast tissues but have some degree of differences in protein expression profiles. FIG.36. Computing environment for combinatorial transformation of barcoding for target complex detection and quantification. FIG.37. Example flow of a process for combinatorial transformation of barcoding for target complex detection and quantification. FIG.38. Computing components for combinatorial transformation of barcoding for target complex detection and quantification. DETAILED DESCRIPTION INTRODUCTION

[0060] Disclosed herein is a DNA barcoding approach that leverages a hybrid molecular nanostructure as a bivalent encoder and the use thereof to comprehensively analyze hierarchical interactions among targets, e.g., proteins, which are difficult to capture using current measurement technologies. Each unit recognizes a target (e.g., a protein) and carries a templated DNA barcode that bilaterally fuses with neighboring units, enabling explicit and comprehensive mapping of different- 22 KILPATRICK TOWNSEND 797562351ordered protein interactions within cells. This interaction-driven and target-agnostic approach linearizes molecular information within individual protein complexes, facilitating multiparametric decoding of diverse protein interactions. By mapping different-ordered target complexes directly in whole cells and measuring the responsive dynamics of higher-order target complexes during cellular activities, we find in some cases, the abundance of higher-order target complexes are significantly different in cancer cells as compared to in control cells. For example, cancers cells may exhibit increased diversity and abundance of certain higher-order target complexes (for example, Figure 10B) while exhibit decreased abundance of certain other higher-order target complexes (for example, Figure 10C). In breast cancer biopsy from patients, this approach accurately diagnosed cancer subtypes and identified higher-order protein interactions associated with disease aggressiveness.

[0061] The technology, termed tandem elongation of templated DNA repeats for analysis of interacting Targets, e.g., proteins (TETRIS), enables in situ encoding and multiparametric decoding – molecular constituents and spatial interactions – of different-ordered target (e.g., protein) complexes and explicit and comprehensive mapping of different-ordered target interactions directly in cells, e.g., without cell lysis or extraction of targets from the cells. To encode, TETRIS leverages hybrid molecular nanostructures as bivalent building blocks. Each unit recognizes a target (e.g., a protein) and carries a templated DNA barcode that fuses bilaterally with neighboring units, in an interaction- driven and target-agnostic manner. Only in interacting targets (proteins), closely-spaced TETRIS units connect in situ to enable bidirectional elongation of DNA barcodes (a schematic representation is shown in Figure 2b); the formed barcodes thus explicitly linearize comprehensive molecular information (i.e., components and interactions) within individual 3D target complexes. Further, combinatorics transformation can be performed on the entire distribution of formed DNA barcodes to accurately measure complex protein interactions of different orders (i.e., composition and abundance).

[0062] This technology is a significant advancement in the field of disease diagnosis and patient care. Specifically, as further disclosed below, while diseased sample (for example, cancer cells) and control samples (for example normal cells) show similar relative amounts of individual targets and lower-order target complexes (e.g., 2-mer target complexes), but may show very different amounts of higher-order target complexes (e.g., 4-mer target complexes). In some cases, diseased samples show more diverse and more abundant higher-order target complexes than the control. In some cases, diseased samples show less diverse and less abundant higher-order target complexes (for example, tumor suppressor complexes) than the control. In one particular study, when applied to assess fine needle aspiration biopsies (n = 72) from breast cancer patients, TETRIS not only accurately diagnosed cancer molecular subtypes, but also revealed higher-order target interactions to effectively distinguish 23 KILPATRICK TOWNSEND 797562351patient disease aggressiveness. As such, this technology empowers biomedical researchers and clinical professionals to assess different-ordered target complexes across various subcellular locations in scarce clinical cell samples, allows more accurate diagnosis and monitoring drug responsiveness, and improves patient care. TERMINOLOGY

[0063] All technical and scientific terms used herein, unless otherwise defined below, are intended to have the same meaning as commonly understood by one of ordinary skill in the art. References to techniques employed herein are intended to refer to the techniques as commonly understood in the art, including variations on those techniques and / or substitutions of equivalent techniques that would be apparent to one of skill in the art. While the following terms are believed to be well understood by one of ordinary skill in the art, the following definitions are set forth to facilitate explanation of the presently disclosed subject.

[0064] Recitation of ranges of values herein are merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, unless otherwise indicated herein, and each separate value is incorporated into the specification as if it were individually recited herein. For example, if a range is stated as 1 to 50, it is intended that values such as 2 to 40, 10 to 30, or 1 to 3%, etc., are expressly enumerated in this specification. These are only examples of what is specifically intended, and all possible combinations of numerical values between and including the lowest value and the highest value enumerated and all individual values in the range are to be considered to be expressly stated in this disclosure.

[0065] As used in herein, the singular forms “a”, “an” and “the” include plural referents unless the content clearly dictates otherwise. Thus, for example, reference to “an antibody” optionally includes a combination of two or more such molecules, and the like.

[0066] The term “about” as used herein refers to the usual error range for the respective value readily known to the skilled person in this technical field, for example ± 20%, ± 10%, or ± 5%, are within the intended meaning of the recited value.

[0067] The term “patient” used interchangeably with “subject,” means any animal, including any vertebrate or mammal, and, in particular, a human, and can also be referred to, e.g., as an individual or patient.

[0068] The term “antibody” includes, but is not limited to, synthetic antibodies, monoclonal antibodies, recombinantly produced antibodies, multispecific antibodies (including bi-specific antibodies), human antibodies, humanized antibodies, chimeric antibodies, single-chain Fvs (scFv), 24 KILPATRICK TOWNSEND 797562351Fab fragments, F(ab′) fragments, disulfide-linked Fvs (sdFv) (including bi-specific sdFvs), and anti- idiotypic (anti-Id) antibodies, and epitope-binding fragments of any of the above. The antibodies provided herein may be monospecific, bispecific, trispecific or of greater multi-specificity. Multispecific antibodies may be specific for different epitopes of a polypeptide or may be specific for both a polypeptide as well as for a heterologous epitope, such as a heterologous polypeptide or solid support material.

[0069] The terms “protein” and “polypeptide” are used interchangeably and refer to any polymer of amino acids (dipeptide or greater) linked through peptide bonds or modified peptide bonds. Polypeptides of less than about 10-20 amino acid residues are commonly referred to as “peptides.” The polypeptides of the invention may comprise non-peptidic components, such as carbohydrate groups. Carbohydrates and other non-peptidic substituents may be added to a polypeptide by the cell in which the polypeptide is produced and will vary with the type of cell. Polypeptides are defined herein, in terms of their amino acid backbone structures; substituents such as carbohydrate groups are generally not specified, but may be present, nonetheless.

[0070] The term “sample” refers to any sample comprising or being tested for the presence of a target of a drug of interest. Such a sample includes samples derived from or containing cells, organisms (bacteria, viruses), lysed cells or organisms, cellular extracts, nuclear extracts, components of cells or organisms, extracellular fluid, media in which cells or organisms are cultured in vitro, blood, plasma, serum, gastrointestinal secretions, urine, ascites, homogenates of tissues or tumors, synovial fluid, feces, saliva, sputum, cyst fluid, amniotic fluid, cerebrospinal fluid, peritoneal fluid, lung lavage fluid, semen, lymphatic fluid, tears, pleural fluid, nipple aspirates, breast milk, external sections of the skin, respiratory, intestinal, and genitourinary tracts, and prostatic fluid. A sample can be a viral or bacterial sample, a sample obtained from an environmental source, such as a body of polluted water, an air sample, or a soil sample, as well as a food industry sample. A sample can be a biological sample which refers to the fact that it is derived or obtained from a living organism. The organism can be in vivo (e.g. a whole organism) or can be in vitro (e.g., cells or organs grown in culture). A “biological sample” also refers to a cell or population of cells or a quantity of tissue or fluid from a subject. Most often, a sample has been removed from a subject, but the term “biological sample” can also refer to cells or tissue analyzed in vivo, i.e., without removal from the subject. Often, a “biological sample” will contain cells from a subject, but the term can also refer to non-cellular biological material, such as non-cellular fractions of blood, saliva, or urine. The biological sample may be from a resection, bronchoscopic biopsy, or core needle biopsy of a primary, secondary or metastatic tumor, or a cellblock from pleural fluid. In addition, fine needle aspirate biological samples are also useful. In one embodiment, a biological sample is primary ascites cells. Biological samples 25 KILPATRICK TOWNSEND 797562351also include explants and primary and / or transformed cell cultures derived from patient tissues. A biological sample can be provided by removing a sample of cells from subject but can also be accomplished by using previously isolated cells or cellular extracts (e.g. isolated by another person, at another time, and / or for another purpose). Archival tissues, such as those having treatment or outcome history may also be used. Biological samples include, but are not limited to, tissue biopsies, scrapes (e.g. buccal scrapes), whole blood, plasma, serum, urine, saliva, cell culture, or cerebrospinal fluid. The samples analyzed by the compositions and methods described herein may have been processed for purification or enrichment of extracellular vesicles contained therein. In one embodiment, the sample is blood.

[0071] The term “drug,” refers to any molecule that can be used in a patient for diagnostics or therapeutic purposes. In some embodiments the drug is a non-protein chemical compound. In some embodiments the drug is a protein, for example, a polypeptide, an antibody or a functional fragment thereof.

[0072] As used herein, and unless otherwise indicated, the term “complementary,” when used to describe a first nucleotide sequence in relation to a second nucleotide sequence, refers to the ability of a polynucleotide comprising the first nucleotide sequence to hybridize and form a duplex structure under certain conditions with a polynucleotide comprising the second nucleotide sequence, as will be understood by the skilled person. Such conditions can, for example, be stringent conditions, where stringent conditions may include: 400 mM NaCl, 40 mM PIPES pH 6.4, 1 mM EDTA, 50°C or 70°C for 12-16 hours followed by washing. Other conditions, such as physiologically relevant conditions as may be encountered inside an organism, can apply. The skilled person will be able to determine the set of conditions most appropriate for a test of complementarity of two sequences in accordance with the ultimate application of the hybridized nucleotides.

[0073] Complementary sequences refer to regions of polynucleotides in which a first nucleotide sequence is capable of base-pairing with a second nucleotide sequence, either along the entire length or over a portion of the length of one or both sequences.. Such sequences can be referred to as “complementary” with respect to each other herein. However, where a first sequence is referred to as “substantially complementary” with respect to a second sequence herein, the two sequences can be complementary, or they may include one or more, but generally not more than about 5, 4, 3, or 2 mismatched base pairs within regions that are base-paired. For two sequences with mismatched base pairs, the sequences will be considered “substantially complementary” as long as the two nucleotide sequences bind to each other via base-pairing. 26 KILPATRICK TOWNSEND 797562351

[0074] “Complementary” sequences, as used herein, may also include, or be formed entirely from, non-Watson-Crick base pairs and / or base pairs formed from non-natural and modified nucleotides, in as far as the above embodiments with respect to their ability to hybridize are fulfilled. Such non- Watson-Crick base pairs includes, but are not limited to, G:U Wobble or Hoogstein base pairing.

[0075] The term “substantially the same” when referring to the two nucleic acid molecules, it means that that two nucleic acid molecules share at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity over the entire length of the longer nucleic acid molecule of the two.

[0076] The term “substantially the same” when referring to the length of two nucleic acid molecules, it means that the length of the shorter nucleic acid molecule is at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the length of the longer nucleic acid molecule of the two.

[0077] As used herein, “barcode” or “barcode sequence” refers to any unique sequence label that can be coupled to at least one nucleotide sequence for, e.g., later identification of the at least one nucleotide sequence.

[0078] The term “in combinatorial form,” or “in combination form,” refers to a barcode is present with one or more different barcodes in a single nucleic acid molecule, i.e., a link nucleic acid molecule that is formed by joining two or more complementary polynucleotides. The two or more barcodes in combinatorial form are collectively referred to as a combination barcode. In this disclosure, a combination barcode is represented by barcodes grouped with an underline, for example, B1B2B3.

[0079] The term “in individual form,” refers to that a barcode is not associated with any other barcodes in a single nucleic acid molecule. For example, B1, B2, and B3 represent three barcodes in individual form. DETAILED DESCRIPTION

[0080] Provided herein are complexes (aka. TETRIS units) that can be used to label targets in a biological sample such that the targets can be detected using methods disclosed herein. 1. Target

[0081] A target described herein refers to any molecule of interest, for example, a nucleic acid, a lipid, a small molecule compound, or a protein. In terms of location the target can be within a cell or outside a cell. In some embodiments the target is a DNA molecule. In some embodiments the target 27 KILPATRICK TOWNSEND 797562351is an RNA molecule, for example, an mRNA molecule. The target can also be a derivative of any of the molecules described above. Throughout the disclosure, where a method or composition involving a protein target is described, it is meant that the method or composition is also applicable to other types of targets, e.g., nucleic acids, lipids, etc.

[0082] In some embodiments, the target is a protein. In some embodiments the target is a modified protein, such as a glycoprotein, lipoprotein, as for the proteins methylated proteins. In some embodiments, the target is selected from the group consisting of proteins listed in Figure 4: GRB2, HER2, SHC, ER, EGFR, FOXA1, LMO4, GATA3, PR, BRCA1, TIF2, VAV2, SCRIB, NOS1AP, PTPN1, SRC, VANGL1, WASF3, NCKAP1, CYFIP1, AP2M1, ERRFI1, UBS3B, P85B, c-Met, HER3, SPRY2, CBL, PLCG1, STAT3, RAF1, PTPN11, CtIP, and SOS1. In some embodiments, the target is selected from the group consisting of PR, LMO4, CYFIP1, TIF2, and UBS3B. 2. Sample

[0083] The method the composition described herein is used to quantify targets in a sample. Such samples include but are not limited to biosamples, such as tissues, isolated cells or cell cultures, bodily fluids (including, but not limited to, blood, urine, serum, lymph, saliva, anal and vaginal secretions, perspiration, and semen); and environmental samples, such as air, agricultural, water and soil samples, etc. in some embodiments, the sample can be a cell lysate, the purified enriched biological sample. The sample may be obtained from one individual or from multiple individuals (i.e., a population). A sample may contain a mixture of cells or even organisms, such as, a human saliva sample that includes human cells and bacterial cells; a mouse xenograft that includes mouse cells and cells from a transplanted human tumor; etc. In some cases, the sample is a cancer biopsy. 3. Complexes (aka. TETRIS units or TETRIS complexes)

[0084] The TETRIS technology disclosed herein can be used to measure complex target interactions directly in whole cells. Each of the TETRIS unit comprises a binding domain, flexible linker, and a polynucleotide. The binding domain specifically binds a target, therefore linking the rest of complex, e.g, the polynucleotide, to the target. Illustrative configurations of the complex are shown in Figure 1A and Figure 6B. a) Binding domain

[0085] In some cases, the binding domain comprises an aptamer, and the binding domain has a length of 20 to 200 nucleotides, 20 to 150 nucleotides, e.g., 20 to 100 nucleotides, 20 to 80 nucleotides, 20 to 60 nucleotides, or 25 to 55 nucleotides. In some embodiments, the binding region comprises an antibody, an ScFv, a Fab, or the like. Antibodies that can specifically bind to various 28 KILPATRICK TOWNSEND 797562351targets (for example, the targets in Figure 4) are well known, and many of which are commercially available. In some embodiments the binding domain is a peptide, for example, a receptor that can bind to the target. In some embodiments, the binding domain is a small molecule that can bind the target, for example, afatinib, erlotinib, which binds to EGFR. b) Flexible linker

[0086] The flexible linker used in the complex typically has a length in the range from 2 nm to 100 nm, for example, in a range from 9.6 nm to 66.3 nm, from 10 nm to 50 nm, from 20 nm to 40 nm. In some embodiments the flexible linker is polyethylene glycol (PEG). In some embodiments, the flexible linker comprises multiple units of a polyethylene glycol, i.e., (PEG)n, wherein the n is an integer within the range of 2-200, for example, 3-150, 4-100, 10-80, 10-50, etc. In some embodiments, the flexible linker is Gly-Gly-Gly-Ser-Gly-Gly-Gly. Other flexible linkers include linear or branched polymers, cleavable linkers, polynucleotides, polypeptides, polysaccharides. Additional flexible linkers that can also be used for the method and compositions disclosed herein are described in, e.g., US11981708 B2, Tables 2 and 3. The entire disclosure of said patent is incorporated herein by reference. c) Polynucleotides comprising unique identifiers

[0087] A complex disclosed herein comprises a polynucleotide comprising a unique identifier is conjugated to the binding domain via to the flexible linker described above. Such a polynucleotide is also referred to as a unique identifier-containing polynucleotide or an encoding polynucleotide in this disclosure. The unique identifier is a unique sequence identifier and encodes the identity of the target by virtue of linking to the binding domain that recognizes the target. In some embodiments, the unique identifier has a length of 5-100 nucleotides, for example 10-70 nucleotides, 15-6018-50 nucleotides, 20-48 nucleotides, 30-40 nucleotides, for example, 18, 26, 40, 48, and 60 nucleotides.

[0088] The encoding polynucleotide is conjugated to the flexible linker at a modified unit at an internal position. In some embodiments, internal position is located at least 5 nucleotides, at least 6 nucleotides at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides away from the closer of the two ends of the polynucleotide to facilitate hybridization on both ends of the encoding polynucleotide (Fig.17A). In some embodiments, the distance between the 5’ end and the internal position is substantially the same as the distance between the 3’ end and the internal position.

[0089] In some embodiments, the encoding polynucleotide contains a modified unit at an internal position through which it can form a covalent bond with a flexible linker. In some embodiments, the 29 KILPATRICK TOWNSEND 797562351modified unit is a nucleotide that has been modified by adding a functional group to facilitate the covalent bond formation. In some embodiments, the modified unit is a nucleoside or nucleotide derivative (e.g., a phosphoramidite) that has been modified by adding a functional group, wherein the nucleoside or nucleotide derivative is a unit integral to the encoding polynucleotide. In some embodiments, the functional group is an amino functional group (-NH2). In some embodiments, a nucleotide at the internal position (“internal nucleotide”) is modified by adding an amino functional group (-NH2). In one illustrative embodiment, the internal nucleotide is modified by adding an amino functional group, such that this modified nucleotide can react with an SPDP–PEG–NHS ester, thus linking the encoding polynucleotide to the PEG linker. In some embodiments, the modified unit is a phosphoramidite with an amino group (-NH2). Exemplary encoding polynucleotides (“identifier A,” “identifier B,” and “identifier C,”) and their modifications are shown in Table 2. In Table 2, the term “iUniAmM” refers to an insertion of a phosphoramidite with an amino functional group (-NH2) between the two flanking nucleotides. For example, ACTCGTGCCACCCGTGCTTCCAG / iUniAmM / ATTTACCTCAGGAGAGTACCACA, shows that a phosphoramidite with an amino group (-NH2) is inserted between G and A.

[0090] Optionally, the encoding polynucleotides each comprises an oligonucleotide, where is referred to as overlap oligonucleotide in disclosure. An overlap oligonucleotide promotes the ligation of the complementary polynucleotides, which comprising sequences that are complementary to the overlap oligonucleotides, when they are hybridized to the adjacent encoding polynucleotides . In some embodiments, the overlap oligonucleotides in all complexes share substantially the same nucleic acid sequence; such overlap oligonucleotides are referred to as universal overlap or universal overlap oligonucleotide in this disclosure. One illustrative example of an overlap oligonucleotide is shown in Figure 1A, left panel. The lengths of the overlap oligonucleotides may vary. In some embodiments, an overlap oligonucleotide has a length of 1 to 100 nucleotides, e.g., 1-50 nucleotides, 1-20 nucleotides, or 3 to 15 nucleotides, 2-10 nucleotides, or about 6 nucleotides. See Figure 14A. The overlap oligonucleotide can be either at the 3’ end or the 5’ end of the encoding polynucleotide. In one embodiment, the overlap oligonucleotide is at the 5’ end of the encoding polynucleotide. Table 2 shows exemplary encoding polynucleotides with an overlap oligonucleotide sequence “ACTCGT” at the 5’ end.

[0091] In some embodiments the encoding polynucleotides are ligatable. In some embodiments the encoding polynucleotides are unligatable. One exemplary approach of producing an encoding polynucleotide with unligatable ends is to remove the phosphate group from the nucleotide at the 5’ end of the encoding polynucleotide, using e.g., a phosphatase, so that it cannot form any phosphodiester bond with another nucleotide. Since the goal of the TERTIS technology is to analyze 30 KILPATRICK TOWNSEND 797562351combination barcode in the link nucleic acid molecules, which are formed by joining complementary polynucleotides (e.g., through ligation), it is advantageous to minimize the formation of other kind of ligation products, such as those formed the encoding polynucleotides themselves, to simplify the combination barcode analysis. 4. Preparation of the complexes (aka. TETRIS units or TETRIS complexes)

[0092] A complex disclosed herein can be prepared by any method that can introduce covalent bonds between two or more molecules, for example, between (1) the binding domain and the flexible linker and (2) between the flexible linker and the encoding polynucleotide. One exemplary method is described in example 1, entitled “General Methods.” For example, a binding domain, such as an antibody, can be activated by reacting with a chemical such as, sulfo-SMCC. An encoding nucleotide comprising an internal nucleotide modified with the amino functional group can be activated by reacting with SPDP–PEG–NHS ester that can react with sulfo-SMCC. SPDP–PEG–NHS ester on the internal nucleotide of the encoding polynucleotide art reacts with sulfo-SMCC in the target binding domain to form a chemical bond, therefore conjugating the target binding domain with the encoding polynucleotide. Other types of linkers can also be used in place of PEG in the methods above. In some embodiments, the method of preparing the complexes using the encoding polynucleotide to the binding domain (e.g., an antibody) in a molar ratio. In some embodiments, the ratio is in the range from 0.5 to 6, for example, from 0.6 to 5, from 0.7 to 4, from 0.8 to 3, from 0.9 to 2, or about 1. See FIG.16B. 5. Complementary polynucleotides and combination barcodes

[0093] Also provided in this disclosure are complementary polynucleotides that can hybridize to the encoding polynucleotides described above. Complementary polynucleotides are also referred to as complements in aspect of the disclosure, for example, in Figure 1A. These complementary polynucleotides comprise nucleic acid sequences that are complementary to the unique identifiers of the encoding polynucleotides of the complexes (i.e., TERTIS units or TETRIS complexes) and these sequences are referred to as barcodes. As such, just like the unique identifiers in the complexes, the barcodes can also be used to identify the targets. Optionally, the complementary polynucleotide comprises a unique molecule index (UMI), which uniquely tag each molecule during amplification. It can also be used to correct errors and increase accuracy when analyzing barcodes. Figure 9 shows schematic illustrations of complementary polynucleotides with UMIs.

[0094] In some embodiments, the complementary polynucleotides comprise sequences complementary to the overlap oligonucleotide in the complexes described above. In some 31 KILPATRICK TOWNSEND 797562351embodiments, the sequences complementary to the overlap oligonucleotides, referred to herein as complementary overlaps or complementary overhangs, have the same or substantially the same length as the respective overlap oligonucleotides. Exemplary complementary polynucleotides are shown in Table 1, for example Complement A, Complement B, and also Table 2, complement A, Complement B, and complement C.

[0095] In some embodiments, the complementary polynucleotides have ligatable ends so that they can be joined to form a link nucleic acid molecule under conditions permissible for the ligation or extension. Joining of the complementary polynucleotides can be performed through various methods, such as, chemical reactions or enzymatic reactions. Enzymatic ligations typically are performed in a reaction mixture comprising a ligation buffer and a ligase and the complementary polynucleotides. The complementary polynucleotides comprise a 5’ phosphate group and 3’ hydroxyl group so that they can be joined with one another under suitable conditions for ligation or extension. Chemical ligations are typically performed in a reaction mixture comprising complementary polynucleotides modified with reactive groups at 3’ or 5’ ends, such that they can be joined together under conditions permissible for the chemical ligation of the reactive groups to occur. Nonlimiting examples of reactive groups are: azide and alkyne groups; aryl azides and benzophenone derivatives, which can be ligated by light illumination (typically UV light for 10-30 mins); pH-sensitive protecting groups at the ends, such as hydrazones and aconityl derivatives, which can be ligated by lowering the pH to <6.5. In one illustrative example, complementary polynucleotides modified with azide at the 3’ end and complementary polynucleotides modified with alkyne groups at the 5’ end, and the complementary polynucleotides modified with azide at the 3’ end can be ligated by mixing them with copper(ii) sulfate (e.g., 150 μM; Sigma-Aldrich), tris(3-hydroxypropyltriazolylmethyl)amine (e.g., 1 mM; Lumiprobe) and sodium ascorbate (e.g., 1.5 mM; Sigma-Aldrich) and incubated for 2 h at room temperature. The resulting link nucleic acid molecule contains all the barcodes (e.g., B1, B2, and B3, etc.) from the complementary polynucleotides that form the link nucleic acid molecule. These barcodes together form a combination barcode, for example, B1B2B3. In some embodiments, a combination barcode comprises sequences of the complementary overlap (“O”) interspersed between the individual barcodes, for example, OB1OB2OB3. For brevity, a combination barcode represented in the form of B1B2B3includes both the form without the overlap sequence and the form with the overlap sequence (e.g., OB1OB2OB3). Examples of combination barcodes are shown in Table 2, in which a combination barcode representing protein complex ACB is shown as follows: ACGAGTGAAACTCACACTGGCGAGATGCAGCAGTCTACTTCCTCCAACGAGTGGTCCATTGACGCTA CATGGATAGCTGACGAAGAGTCTGTACGAGTTGTGGTACTCTCCTGAGGTAAATCTGGAAGCACGGG TGGC, 32 KILPATRICK TOWNSEND 797562351in which the double underlined sequence, ACGAGT, is the complementary overlap sequence, and the bold letters represent the barcodes identifying protein A, protein C, and protein B, respectively, in that order. Together, the sequence represents a combination barcode ACB. Detecting this combination barcode in a reaction in which the sample is treated with the TETRIS technology indicates that proteins A, C, and B interact with one another in the sample, and that protein C is more closely associated with protein A than protein B is. 6. Joining the complementary polynucleotide in situ

[0096] The complementary polynucleotides bound to the complexes are placed under conditions to allow them to form a link nucleic acid molecule. In some embodiments the complementary polynucleotides bound to the complexes are ligated to form a link nucleic acid molecule in the presence of reagents necessary for the ligation, for example a ligase and suitable buffers. If the encoding polynucleotides in the complexes comprise the overlap oligonucleotide, the ligation of the encoding polynucleotides link the sticky ends of the two adjacent encoding polynucleotides, i.e., a sticky end ligation. If the encoding polynucleotides in the complexes do not comprise the overlap oligonucleotide, they can be ligated under conditions suitable for blunt end ligation. Both type of ligations can be used in the methods disclosed herein although, sticky end ligations typically can occur in much higher efficiency than blunt end ligations.

[0097] Besides ligation through enzymatic reactions, complementary polynucleotides can also be joined to chemical reaction, for example, by click chemistry through triazole linkage. As an illustrative example, complementary polynucleotides can be modified with azide and alkyne modifications at the 3’- and 5’-ends, respectively (Table 1), followed by incubating with copper(ii) sulfate, tris(3-hydroxypropyltriazolylmethyl)amine, and sodium ascorbate for a sufficient time to allow formation of the triazole linkage through which the complementary polynucleotides are joined to form combination barcodes. The incubation time can be 0.5h to 8h, for example, 1-4 hours, or 1-2 hours.

[0098] In some embodiments, the complementary polynucleotides are joined by extension in the presence of a DNA polymerase with exonuclease activity. In some embodiments, the DNA polymerase may possess 5’-3’ exonuclease activity, 3’-5’ exonuclease activity, or both. DNA polymerases with such properties include strand-displacement polymerases and nick-translating polymerases with exonuclease activities. Nonlimiting examples of DNA polymerases that can be used herein include DNA Polymerase I, phi29 DNA Polymerase, T7 DNA Polymerase, Bst DNA Polymerase, Taq DNA Polymerase, as well as their fragments or derivatives or engineered versions, provided that these fragments or derivatives or engineered versions retain polymerase activity and 33 KILPATRICK TOWNSEND 797562351exonuclease activity. Under this approach, one of the complementary polynucleotides is extended by the DNA polymerase using the polynucleotide (in a complex or TETRIS unit) it is hybridized to as a template, and simultaneously with the extension, the DNA polymerase removes nucleotides from the 5’ of the adjacent downstream complementary polynucleotide that is hybridized to the polynucleotide in the adjacent complex, effectively displacing the downstream complementary polynucleotide. The strand-displacement extension continues to displace all complementary polynucleotides hybridized to polynucleotides that identify targets that interact and form a multi-target complex in the sample to form a continuous DNA strand, which is a link nucleic acid molecule comprising a combination barcode with all barcodes in these complementary polynucleotides. In some embodiments, a link nucleic acid molecule disclosed above further comprises one or more UMIs.

[0099] Optionally, the link nucleic acid molecules from above comprising the combination barcodes can be released by denaturing the reaction mixture, which separates the link nucleic acid molecules from the polynucleotides in the TETRIS units they are hybridized to. In some embodiments, denaturing is performed by adding a chemical for example, sodium hydroxide, to the reaction mixture. In some embodiments, denaturing is performed by heating the reaction mixture. The link nucleic acid molecules are then collected and / or purified before subjecting to the combination barcode analysis as further described below. Optionally, the link nucleic acid molecules can be amplified before performing the combination barcode analysis described below. 7. Characterizing and quantifying combination barcodes

[0100] Various methods can be used to quantify and / or characterize the combination barcodes. Exemplary methods include sequencing and quantitative PCR. a) Sequencing

[0101] In some embodiments, the link nucleic acid molecules or amplified link nucleic acid molecules comprising the combination barcodes are sequenced using methods known in the art, including for example without limitation, polymerase-based sequencing-by-synthesis (e.g., HiSeq 2500 system, Illumina, San Diego, CA), ligation-based sequencing (e.g., SOLiD 5500, Life Technologies Corporation, Carlsbad, CA), ion semiconductor sequencing (e.g., Ion PGM or Ion Proton sequencers, Life Technologies Corporation, Carlsbad, CA), zero-mode waveguides (e.g., PacBio RS sequencer, Pacific Biosciences, Menlo Park, CA), nanopore sequencing (e.g., Oxford Nanopore Technologies Ltd., Oxford, United Kingdom), pyrosequencing (e.g., 454 Life Sciences, Branford, CT), or other sequencing technologies. Some of these sequencing technologies are short- read technologies, but others produce longer reads, (e.g., the GS FLX+ (454 Life Sciences; up to 1000 bp), PacBio RS (Pacific Biosciences; approximately 1000 bp) and nanopore sequencing (Oxford 34 KILPATRICK TOWNSEND 797562351Nanopore Technologies Ltd.; 100 kb)). For haplotype phasing, longer reads are advantageous and require much less computation, although they tend to have a higher error rate and errors in such long reads may need to be identified and corrected according to methods set forth herein before haplotype phasing.

[0102] In some approaches, sequencing is performed using combinatorial probe-anchor ligation (cPAL) as described in, for example, US 20140051588, U.S. 20130124100, both of which are incorporated herein by reference in their entirety for all purposes.

[0103] In some approaches, sequencing is performed using DNBseq sequencers. The link nucleic acid molecules are denatured to produce single-stranded molecules. These circles are then used to make DNA nanoballs (DNBs) for DNBseq sequencers. In some approaches, the link nucleic acid molecules are sequenced on Illumina or other systems that do not require circularization. a) Quantitative PCR

[0104] In some approaches, quantitative PCR (qPCR) can be performed to detect the presence of and / or determine the amount of the barcodes in individual form or combination form using primers with sequences that are substantially the same as the unique identifiers such that the primers can hybridize to the barcodes. Exemplary primers for performing qPCR in determining the presence of barcodes for characterizing the multi-target complex is shown in Table 2. 8. Quantifying abundance of individual targets using TETRIS units (complexes)

[0105] Embodiments of this disclosure provide methods of measuring one or more targets in a sample in one single reaction through quantifying the combination barcodes as described above. In some embodiments, the TETRIS technology can be performed in situ, i.e., without the cell lysis or extraction of targets. In some embodiments, the method can simultaneously analyze N targets, where N can be any integer. In some embodiments, N is any integer in the range of 1 to 100, for example, 1 to 50, e.g., 2-40, or 3-34. The method comprises contacting a sample comprising N complexes, each complex having the features as described above and each complex identifying a different target. The method further comprises contacting the N complexes with N complementary polynucleotides as described in the above (see the section entitled “Complementary polynucleotides and combination barcodes”), each complementary polynucleotide comprising a barcode that is complementary to the unique identifier of a corresponding complex. Thus, the N complexes comprise N different barcodes. Optionally, each complementary polynucleotide comprises a UMI. The method further comprises 35 KILPATRICK TOWNSEND 797562351quantifying the one or more of the N barcodes, thereby measuring the abundance of the one or more of the N targets. In some embodiments, the N complexes and the N complementary polynucleotides are incubated with the sample simultaneously. In some other embodiments, the N complexes are incubated with the sample in a reaction mixture before the N complementary polynucleotides are added to the reaction mixture. The method comprises further quantifying the one or more of the N barcodes, thereby measuring the abundance of the one or more of the N targets in the sample. 9. Quantifying interaction and abundance of targets using TETRIS units (complexes)

[0106] Embodiments of the application also provide methods and compositions to detect and quantify the interactions between the two or more targets. In some embodiments, the method can detect and quantify the interactions among N targets, where N is an integer range of 1 to 100, for example, 1 to 50, e.g., 2-40, and 3-34. The method comprises contacting the sample with N complexes, where each complex uniquely identifies one of the N different targets, and contacting the N complexes with N complementary polynucleotides, thus forming a reaction mixture. The complementary polynucleotide comprises ligatable ends and can be joined under suitable conditions when they brought into in proximity by virtue of binding to targets that interact with one another. The reaction mixture comprising these complementary polynucleotides, which remain hybridized to the encoding polynucleotides in the complexes, are then placed in conditions that allow the complementary polynucleotides to be joined in situ, thereby producing a link nucleic acid molecule comprising a combination barcode formed by the individual barcodes present in the complementary polynucleotides.

[0107] Targets that interact with one another forming a multi-target complex bring the TETRIS complexes that bind to the targets in proximity and in turn bring the complementary polynucleotides in proximity. Under the conditions as disclosed above, complementary polynucleotides each comprising a barcode, are joined to result in formation of various forms of combination barcodes, e.g., 2-mer, 3-mer, 4-mer, ... or M-mer, and so on. M is an integer. A 2-mer combination barcode comprises two (2) barcodes, identifying two (2) different targets that form a multi-target complex, for example, B1B2,and B2B3. If not all N targets in the sample to be characterized interact one another, then the M is a number smaller than N. If all targets in the sample interact with one another, the M is the same as N, i.e. all N barcodes are included in the link nucleic acid molecule.

[0108] Analyzing the composition of the combination barcode in the link nucleic acid molecules can thus reveal the nature of interaction between the targets. As one illustrative example, to detect a sample comprising N targets: T1, T2, T3, ... Tm, T(m+1), ..., and Tn, the method labels the targets with 36 KILPATRICK TOWNSEND 797562351N complexes comprise unique identifiers U1, U2, U3, . . . Um, U(m+1), . . . Un, and complementary polynucleotides comprising corresponding barcodes B1, B2, B3, . . . Bm, B(m+1), . . . Bn, i.e., B1is complementary to U1, . . . , Bm is complementary to Um, . . ., Bn is complementary to Un etc. After joining the complementary polynucleotides to form combination barcodes, detection of the presence of a combination barcode comprising two barcodes indicate the presence of a 2-mer complex in the sample, and detection of the presence of a combination barcode comprising three barcodes indicate the presence of a 3-mer complex in the sample, and so on. Detection of a combination barcode B3B4B5. . . Bm(where M is less than or equal to N) indicates that targets T3, T4, T5, . . . , and Tminteract with one another and that they form a M-mer complex in the sample. If B1and B2are only detected in individual form, i.e., in separate molecules, not present together in a single link nucleic acid molecule (i.e., they are not present in combination form), then these two targets likely do not interact with any of the other targets being analyzed. By quantifying the number of barcodes in each of the combination barcodes in the link nucleic acid molecules, one can determine the identity of the targets that form complexes in the sample.

[0109] Not only the TETRIS technology can be used to detect the multi-target complex formation in the sample, it can also quantify the number of complexes in the sample. This is because the amount of link nucleic acid molecules comprising a particular combination barcode positively correlates with the number of complexes formed by the targets that are identifiable by the individual barcodes constitute the combination barcode. For example, quantifying the number of link nucleic acid molecules comprising a combination barcode B1B2B3 provides the amount of multi-target complex T1T2T3 in the sample. When all conditions are equal, a sample with a higher number of link nucleic acid molecules comprising a combination barcode B1B2B3has a higher number of multi-target complex T1T2T3.

[0110] In some embodiments, the method can also be used to quantify individual targets T1, T2, T3, ... Tm, T(m+1), ..., and Tnin the sample. In practice, joining the complementary polynucleotide may produce a distribution of combination barcodes and the individual barcodes, i.e., 1-mer barcode, 2- mer barcodes, 3-mer barcodes, and so on. See Figure 2A, 2E, and 2F. The abundance of target T1 can be calculated by summing the amount of the barcode present in individual form and the amount of the barcode present as part of a combination barcode. As an illustrative example, the amount of target T1 can be represented by the summing the amount of B1, the amount of B1B2, the amount of 2-mer combination barcodes B1B2, 3-mer combination barcode B1B2B3, and 4-mer combination barcode B1B2B3B4, (or B1B1B2B3), and M-mer combination barcode B1B2B3... Bm, and so on.

[0111] In some embodiments, each barcode constituting the combination barcode is distinct from any other barcode constituting the combination barcode, for example, B1B2B3B4. In some 37 KILPATRICK TOWNSEND 797562351embodiments, a combination barcode comprises more than one copy of an individual barcode, for example, B1B1B2B3, which indicates that T1forms homodimers in the sample. When calculating the T1 amount, the amount of B1B1B2B3 will be accounted for twice as it contains a homodimer of B1. 10. Using TETRIS for disease diagnosis

[0112] As demonstrated in this disclosure, e.g., Figure 1C-1D, pathological samples, for example, tumor tissues, generally exhibit increased diversity and abundance of higher-order multi-target complexes compared to normal tissues. Consequently, the quantitative difference between a pathological sample and a normal sample is more pronounced for the higher order protein complex (for example, a 3-mer complex) than for the lower order protein complex (for example, a 2-mer complex) as illustrated in Figure 4B. Thus, the methods disclosed herein can be used to detect and quantify the abundance of higher-order multi-target complexes for disease diagnosis and / or prognosis with high accuracy.

[0113] In some embodiments, the method comprises quantifying interaction and / or abundance of N targets from a patient sample with the method disclosed above and diagnosing the patient as having a disease or have an increased risk of developing the disease if the amount of the link nucleic acid molecules, each comprising the M barcodes (where M is less than or equal to N), is significantly different a control sample, for example, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 100%, or at least 200% different than the control sample. For purpose of this disclosure, a control sample can be a sample from a healthy individual. A control sample used herein is typically of the same tissue origin as the sample being tested. In some embodiments, the method further comprises collecting the sample from a patient.

[0114] In some embodiments, the N targets are one or more of those molecules that promote or contribute to the development of the disease, i.e., pathogenic molecules, for example carcinogenic proteins, and the method comprises quantifying interaction and / or abundance of N targets from a patient sample with the method disclosed above and diagnosing the patient as having a disease or have an increased risk of developing the disease if the amount of the link nucleic acid molecules, each comprising the M barcodes (where M is less than or equal to N), is greater than a control sample, for example, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 100%, or at least 200% greater than the control sample. In some embodiments, the method further comprises collecting the sample from a patient. In some embodiments, the N targets are one or more of those targets listed in Figure 4. 38 KILPATRICK TOWNSEND 797562351

[0115] In some embodiments, the N targets are one or more of those molecules that suppress development of the disease, for example, tumor suppressor proteins, and the method comprises quantifying interaction and / or abundance of N targets from a patient sample with the method disclosed above and diagnosing the patient as having a disease or have an increased risk of developing the disease if the amount of the link nucleic acid molecules, each comprising the M barcodes (where M is less than or equal to N), is less than a control sample, for example, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 100%, or at least 200% lower than the control sample. In some embodiments, the method further comprises collecting the sample from a patient. In some embodiments, the N targets are one or more of those molecules listed in Figure 10C.

[0116] In some embodiments, the disease is a cancer. The term “cancer” as used herein, include all types of cancer, neoplasm, or malignant tumors found in mammals, including leukemia, carcinomas, and sarcomas. Exemplary cancers include cancer of the brain, breast, cervix, colon, head & neck, liver, kidney, lung, non-small cell lung, melanoma, mesothelioma, ovary, sarcoma, stomach, uterus, and medulloblastoma. Additional examples include Hodgkin’s Disease, Non-Hodgkin’s Lymphoma, multiple myeloma, neuroblastoma, ovarian cancer, rhabdomyosarcoma, primary thrombocytosis, primary macroglobulinemia, primary brain tumors, cancer, malignant pancreatic insulinoma, malignant carcinoid, urinary bladder cancer, premalignant skin lesions, testicular cancer, lymphomas, thyroid cancer, neuroblastoma, esophageal cancer, genitourinary tract cancer, malignant hypercalcemia, endometrial cancer, adrenal cortical cancer, neoplasms of the endocrine and exocrine pancreas, and prostate cancer.

[0117] In some embodiments, the disease is a neurodegenerative disease, a cardiovascular disease, a metabolic disease, an infectious disease. 11. Using TETRIS to monitor a disease progression or evaluate drug efficacy or drug response

[0118] It is noted that typically as a disease as described above progresses, the amount of higher order target complex formation also increases. As such, the TETRIS method can be used to compare disease severity between two different samples. In scenarios where the different samples are collected from two different patients, the method can be used to determine which patient has a more severe form of disease and thus more likely need immediate medical intervention. In scenarios where the samples are collected from the same patient at a different time points of the illness, the method can be used to determine whether the patient disease has worsened or improved. Thus, embodiments of the disclosure provide methods of monitoring disease progression in a patient, the method comprising quantifying interaction and / or abundance of M targets within a cell from a patient sample with the 39 KILPATRICK TOWNSEND 797562351method at a first time point and a second time point during a time period, wherein the second time point is later than the first time point. In cases where the targets are pathogenic, the method further comprises determining that the disease has progressed if the amount of the link nucleic acid molecules, each comprising the M barcodes, at the second time point is greater than the amount of the link nucleic acid molecules, each comprising the M barcodes, at the first time point. In cases where the targets are suppressors of the disease, the method further comprises determining that the disease has progressed if the amount of the link nucleic acid molecules, each comprising the M barcodes, at the second time point is less than the amount of the link nucleic acid molecules, each comprising the M barcodes, at the first time point.

[0119] The TETRIS technology can also be used to assess a patient’s response to a treatment or to evaluate the efficacy of a drug in treating the disease. In some embodiments, the method comprises contacting the sample collected from a patient at the first time point of the treatment with the N complexes. The amount of the link nucleic acid molecules each comprising the M barcodes the first time point is compared with the amount of the link nucleic acid molecules each comprising the M barcodes at the second time point of the treatment period, where M is equal to or less than N. In cases where the targets are pathogenic, the patient is determined as responsive to the drug (i.e., the drug is effective in treating the patient) if the amount at the second time point is less than the amount at the first time point, for example, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 100%, or at least 200% less. In cases where the targets are suppressors of the disease, the patient is determined as responsive to the drug (i.e., the drug is effective in treating the patient) if the amount at the second time point is greater than the amount at the first time point, for example, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 100%, or at least 200% greater. 12. Kits

[0120] Embodiments of the disclosure also provide kits for quantifying interactions and / or abundance of N targets, where N is an integer in the range of 1 to 100, for example, 1 to 50, e.g., 2- 40, or 3-34. The kit comprises N complexes, each complex comprising: 1) a binding domain, wherein the binding domain identifies a target, 2) a flexible linker, and 3) a polynucleotide, where the polynucleotide comprises an internal nucleotide conjugated to the flexible linker, where the polynucleotide comprises a unique identifier that uniquely identifies the binding domain and thus also uniquely identifies the target that the binding domain specifically binds. The unique identifiers of different complexes identify different targets, and the unique identifiers of the N complexes identify N different targets. 40 KILPATRICK TOWNSEND 797562351

[0121] In some embodiments, a kit disclosed above further comprises N complementary polynucleotides, where each complementary polynucleotide comprises a barcode being complementary to the unique identifier of one complex, and each complementary polynucleotide has ligatable ends. In some embodiments, the polynucleotide of each complex comprises an overlap oligonucleotide, and each complementary polynucleotide has a complementary overlap that is complementary to the overlap oligonucleotide. In some embodiments, the kit further comprises buffer and enzymes that are suitable for joining one or more complementary polynucleotides. Nonlimiting examples of enzymes include ligases, DNA polymerases, etc. 13. Combinatorial transformation

[0122] FIG.36 illustrates an example of a computing environment for combinatorial transformation of barcoding for target complex detection and quantification. The computing environment includes a sequencing system 3610 in communication with a computer system 3620. In some instances, the functionalities of the computer system 3620 can be implemented as program code that the sequencing system 3610 executes (e.g., the computer system 3620 can be implemented as a component of the sequencing system 3610). The sequencing system 3610 can receive a sample 3605 of link nucleic acid molecules or amplified link nucleic acid molecules comprising combination barcodes. The sequencing system 3610 sequences the sample 3605 using methods known in the art, and as described herein above, for example in FIG.9, to generate barcode sequences.

[0123] In some embodiments, the computer system 3620 receives the barcode sequences. From the barcode sequences, the computer system 3620 can determine a distribution 3622 of barcode combinations in the sample 3605. The distribution 3622 indicates an amount of each barcode combination in the sample 3605, where each barcode combination includes at least one barcode of a target. For example, if a target is a target complex including proteins A (T1), B (T2), and C (T3), the barcode combinations include each combination of the barcodes that correspond to the proteins A, B, and C. For instance, barcode B1can correspond to T1, barcode B2can correspond to T2, and barcode B3 can correspond to T3, so the barcode combinations can include B1, B2, B3, B1B2, B1B3, B2B3, and B1B2B3. So, the distribution 3622 can includes measured amounts of each of the barcode combinations of B1, B2, B3,B1B2, B1B3,B2B3, and B1B2B3in the sample 3605. A number of constituents in a target complex can be in a range of 1 to 100, for example, 1–15, 16–40, 41–70, 71– 90, or 91–100.

[0124] Upon determining the distribution 3622, a combinatorics module 3624 can determine a contribution of a target complex to the distribution 3622 of the barcode combinations. The contribution corresponds to an amount of a barcode combination contributed by the target complex 41 KILPATRICK TOWNSEND 797562351to the distribution 3622. From the contribution, the computer system 3620 can ultimately determine an abundance 3626 of the target complex in the sample 3605. To determine the contribution and the abundance 3626, and as illustrated in FIG.8, the combinatorics module 3624 can perform an iterative process of detecting a longest combination barcode of the barcode combinations in the distribution 3622, determining an abundance of a target complex that corresponds to the longest combination barcode, and updating the distribution 3622 until a remaining longest combination barcode includes only a single barcode.

[0125] Referring to the clustering illustrated in FIG.8b, the computer system 3620 determines the longest barcode combination in the distribution 3622 to be ABC (e.g., B1B2B3), which corresponds to an originating complex of ABC and also a largest target complex ABC (e.g., T1T2T3). Based on an amount of the longest combination barcode in the distribution 3622 and a probability of barcode formation (p), the combinatorics module 3624 determines an amount of each barcode sub- combinations in the longest combination barcode contributed by the largest target complex, which corresponds to the contribution of the largest target complex to the distribution 3622. The probability of barcode formation corresponds to a likelihood of a single linkage forming in a link barcode molecule. An amount of a link barcode molecule comprising a combination barcode that is generated is a function of a barcode combination length, an abundance and a length of an originating target complex, and the probability of barcode formation. In some instances, the link barcode molecule comprises a UMI. The combinatorics module 3624 then performs a conversion using a combinatorics equation to determine the abundance 3626 of the largest target complex based on the amount of the barcode sub-combinations in the longest combination barcode and the probability of barcode formation. Combinatorics equations for different sizes of target complexes are shown in FIG.8 and FIG.20, where T corresponds to the amount of a barcode sub-combination and M corresponds to the abundance 3626 of the target complex.

[0126] Once the abundance 3626 of the largest target complex is determined, the combinatorics module 3624 can generate an updated distribution by removing the amount of the barcode sub- combinations in the longest combination barcode contributed by the largest target complex from the distribution 3622. The computer system 3620 determines a next longest combination barcode in the updated distribution. The next longest combination barcode can correspond to a longest combination barcode that is present in the updated distribution. For instance, the computer system 3620 can determine that the next longest barcode combination in the updated distribution to be AB (e.g., B1B2), which corresponds to an originating complex of AB and also a next largest target complex AB (e.g., T1T2). Based on an amount of the next longest combination barcode in the updated distribution and the probability of barcode formation, the combinatorics module 3624 determines an amount of each 42 KILPATRICK TOWNSEND 797562351barcode sub-combinations in the next longest combination barcode contributed by the next largest target complex, which corresponds to the contribution of the next largest target complex to the distribution 3622. The combinatorics module 3624 then performs a conversion using a combinatorics equation to determine the abundance 3626 of the next largest target complex based on the amount of the barcode sub-combinations in the next longest combination barcode and the probability of barcode formation. This process continues until a remaining longest combination barcode is a single barcode.

[0127] In an example, the computer system 3620 determines an interaction map of constituents in the target complexes based on the abundance of the target complex. Constituents that are closer together in the interaction map may be more likely to bind. Similarly, permutations of barcode combinations may be analyzed to determine which constituents tend to be next to each other in a link barcode molecule.

[0128] As previously described, pathological samples, for example, tumor tissues, generally exhibit increased diversity and abundance of higher-order multi-target complexes relative to normal tissues. As a result, the amount difference between a pathological sample and a normal sample is greater with the higher order protein complex (for example, a 3-mer complex) than with the lower order protein complex (for example, a 2-mer complex). Thus, quantifying the abundance of target complexes can be used for disease diagnosis and / or prognosis.

[0129] In some embodiments, the method involves determining an abundance of a target complex in a sample as disclosed herein and determining that the abundance of the target complex is different than a control sample by greater than a threshold difference. The threshold difference may be, for example, 10% different, 20% different, 30% different, 40% different, or 50% different than the control sample. Based on the abundance of the target complex being greater than a threshold difference different from the control, a patient can be diagnosed as having a disease (e.g., cancer, a neurodegenerative disease, a cardiovascular disease, a metabolic disease, an infectious disease, etc.) or as having an increased risk of developing the disease. For purpose of this disclosure, a control sample can be a sample from a healthy individual. A control sample used herein is typically of the same tissue origin as the sample being tested. In some embodiments, the method further involves collecting the sample from a patient. In some embodiments, the targets are one or more of those targets listed in FIG.4.

[0130] FIG. 37 illustrates an example flow of a process for combinatorial transformation of barcoding for target complex detection and quantification. A computer system (e.g., computer system 3620 in FIG.36) is described as performing the operations of the example flow. Some or all of the instructions for performing the operations can be implemented as hardware circuitry and / or stored as 43 KILPATRICK TOWNSEND 797562351computer-readable instructions on a non-transitory computer-readable medium of the computer system. As implemented, the instructions represent modules that include circuitry or code executable by processor(s) of the computer system. The use of such instructions configures the computer system to perform the specific operations described herein. Each circuitry or code in combination with the relevant processor(s) represent a means for performing a respective operation(s). While the operations are illustrated in a particular order, it should be understood that no particular order is necessary and that one or more operations may be omitted, skipped, performed in parallel, and / or reordered.

[0131] In an example, the flow includes operation 3702, where the computer system determines a distribution of barcode combinations in a sample. The distribution indicates an amount of each of multiple barcode combinations that are present in the sample. The distribution is determined based on a sequencing of the sample.

[0132] In an example, the flow includes operation 3704, where the computer system determines a contribution of a target complex to the distribution. If the target complex is a largest target complex, then the computer system determines the contribution of the target complex based on an amount of the combination barcode that corresponds to the target complex and a probability of barcode formation. The contribution is an amount of each of multiple barcode sub-combinations in the combination barcode contributed by the target complex. If the combination barcode is B1B2B3, then the barcode sub-combinations can include B1, B2, B3,B1B2, B1B3,B2B3, and B1B2B3.The distribution can then be updated by removing the amount of each of the multiple barcode sub-combinations from the distribution so that the contributions of other target complexes can be determined.

[0133] In an example, the flow includes operation 3706, where the computer system determines an abundance of the target complex in the sample. The computer system determines the abundance based on a conversion that uses the contribution of the target complex. The conversion involves using a combinatorics equation that includes the probability of barcode formation to convert from the amount of barcode sub-combinations to the abundance of the target complex.

[0134] FIG. 38 illustrates components of a computer system 3800 according to various embodiments. The computer system 3800 may be an example of the computer system 3620 and / or the sequencing system 3610 in FIG. 36. The computer system 3800 can include any suitable combination of components to process barcode sequences to determine abundances of target complexes in a sample. For example, in the illustrated embodiment, the computer system 3800 includes one or more processors 18, read only memory (ROM) 20, random access memory (RAM) 22, a wireless receiver 24, one or more input devices 26, and a communication bus 28, which provides a communication interconnection path for the components of the computer system 3800. The ROM 44 KILPATRICK TOWNSEND 79756235120 can store basic operating system instructions for an operating system of the controller. The RAM 22 can store barcode sequence information received from a sequencing system and program instructions to process the barcode sequence information to determine the target complexes and the abundance of barcode sequences in a sample.

[0135] Other variations are within the spirit of the present disclosure. Thus, while the disclosed techniques are susceptible to various modifications and alternative constructions, certain illustrated embodiments thereof are shown in the drawings and have been described above in detail. It should be understood, however, that there is no intention to limit the invention to the specific form or forms disclosed, but on the contrary, the intention is to cover all modifications, alternative constructions, and equivalents falling within the spirit and scope of the invention, as defined in the appended claims.

[0136] The use of the terms “a” and “an” and “the” and similar referents in the context of describing the disclosed embodiments (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The terms “comprising,” “having,” “including,” and “containing” are to be construed as open-ended terms (i.e., meaning “including, but not limited to,”) unless otherwise noted. The term “connected” is to be construed as partly or wholly contained within, attached to, or joined together, even if there is something intervening. Recitation of ranges of values herein are merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, unless otherwise indicated herein, and each separate value is incorporated into the specification as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”) provided herein, is intended merely to better illuminate embodiments of the invention and does not pose a limitation on the scope of the invention unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the invention.

[0137] Preferred embodiments of this disclosure are described herein, including the best mode known to the inventors for carrying out the invention. Variations of those preferred embodiments may become apparent to those of ordinary skill in the art upon reading the foregoing description. The inventors expect skilled artisans to employ such variations as appropriate, and the inventors intend for the invention to be practiced otherwise than as specifically described herein. Accordingly, this invention includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Moreover, any combination of the above-described 45 KILPATRICK TOWNSEND 797562351elements in all possible variations thereof is encompassed by the invention unless otherwise indicated herein or otherwise clearly contradicted by context.

[0138] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein. EXEMPLARY EMBODIMENTS Embodiment 1. A composition of N complexes identifying N targets, each of the N complexes identifying a different target, wherein each complex comprises: 1) a binding domain, wherein the binding domain identifies a target, 2) a flexible linker, and 3) an encoding polynucleotide, wherein the encoding polynucleotide comprises a modified unit at an internal position, wherein the encoding polynucleotide is conjugated to the flexible linker via the modified unit, wherein the encoding polynucleotide comprises a unique identifier that identifies the binding domain and wherein the unique identifier identifies the target, wherein the binding domains of different complexes identify different targets, wherein the unique identifiers of different complexes identify different targets, thereby the unique identifiers of the N complexes identify the N targets. Embodiment 2. The composition of embodiment 1, wherein N is in a range of 1 to 100. Embodiment 3. The composition of embodiment 1, wherein modified unit is a phosphoramidite with an amine group. Embodiment 4. The composition of embodiment 1, wherein the flexible linker is a polyethylene glycol (PEG). Embodiment 5. The composition of embodiment 1, wherein the N targets are proteins, lipids, nucleic acids, small molecule compounds, and derivatives thereof. Embodiment 6. The composition of embodiment 1, wherein the encoding polynucleotide comprises an overlap oligonucleotide. Embodiment 7. The composition of embodiment 6, wherein the overlap oligonucleotides in all the complexes have the same nucleic acid sequence. 46 KILPATRICK TOWNSEND 797562351Embodiment 8. The composition of embodiment 1, wherein the binding domain is an antibody, a peptide, a nucleic acid, a small molecule, or an aptamer. Embodiment 9. The composition of embodiment 1, wherein the encoding polynucleotide is unligatable. Embodiment 10. The composition of embodiment 1, wherein the encoding polynucleotide is ligatable. Embodiment 11. The composition of embodiment 1, wherein overlap oligonucleotide has a length of 0 to 100 nucleotides. Embodiment 12. The composition of embodiment 1, where unique identifier has a length of 5 to 100 nucleotides. Embodiment 13. The composition of embodiment 1, where internal nucleotide of each encoding polynucleotide is located at least 5 nucleotides away from the closer of the two ends of the encoding polynucleotide. Embodiment 14. The composition of embodiment 1, where the linker has a length in a range from 2 nm to 100 nm, for example, in a range from 9.6 nm to 66.3 nm. Embodiment 15. The composition of embodiment 4, where the PEG linker has 2-200 units. Embodiment 16. A method of measuring one or more of N targets in a sample, wherein N is an integer, wherein the method comprises: 1) contacting the sample with the N complexes of embodiment 1, 2) contacting the N complexes with N complementary polynucleotides, each of the N complexes identifying a different target, wherein each complementary polynucleotide comprises a barcode being complementary to the unique identifier of one of the N complexes, and wherein the N complexes comprise N barcodes, wherein optionally, each complementary polynucleotide comprises a UMI, and quantifying the one or more of the N barcodes, thereby measuring the abundance of the one or more of the N targets. Embodiment 17. The method of measuring one or more of N targets in a sample of embodiment 16, wherein the encoding polynucleotide in each complex comprises an overlap oligonucleotide, and wherein each complementary polynucleotide has a universal complementary overlap, wherein the universal complementary overlap is complementary to the overlap oligonucleotide in each complex. 47 KILPATRICK TOWNSEND 797562351Embodiment 18. The method of measuring one or more of N targets in a sample of embodiment 16, wherein measuring the amount of the one or more of the N barcodes is by one or more of sequencing and quantitative PCR. Embodiment 19. A method of quantifying interactions and / or abundance of N targets in a sample, wherein N is an integer, wherein the method comprises: contacting the sample with the N complexes of embodiment 1, contacting the N complexes with N complementary polynucleotides, wherein each complementary polynucleotide comprises a barcode being complementary to the unique identifier in one of the N complexes, wherein the barcode identifies the target the one of the unique identifier identifies, wherein each complementary polynucleotide has ligatable ends, wherein optionally, each complementary polynucleotide comprises a UMI, generating link nucleic acid molecules in a reaction mixture, wherein at least one of the link nucleic acid molecules comprising the sequences of M of N complementary polynucleotides and thus comprising M of the N barcodes, optionally separating the link nucleic acid molecules from the complexes, quantifying the amount of the link nucleic acid molecules comprising the M barcodes, wherein M is an integer no greater than N. Embodiment 20. The method of embodiment 19, wherein each complementary polynucleotide has a universal complementary overlap wherein the universal complementary overlap is complementary to the overlap oligonucleotide in each complex. Embodiment 21. The method of embodiment 19, wherein the link nucleic acid molecules comprises one or more copies of at least one of the M barcodes. Embodiment 22. The method of embodiment 19, wherein generating link nucleic acid molecules in situ comprises ligating the complementary polynucleotides. Embodiment 23. The method of embodiment 19, wherein generating link nucleic acid molecules in situ comprises extension using a DNA polymerase. Embodiment 24. The method of embodiment 19, wherein the method further comprises quantifying one or more of the N targets, wherein for each target, quantifying the target by summing the barcode identifying the target in individual form and combination form. 48 KILPATRICK TOWNSEND 797562351Embodiment 25. The method of embodiment 19, wherein the sample is a cell, a cell lysate, or a purified / enriched biological sample. Embodiment 26. A method of disease diagnosis or prognosis, the method comprising: quantifying interaction and / or abundance of N targets within a cell from a patient sample with the method of embodiment 19, diagnosing the patient as having a disease or have increased risk of developing the disease if the amount of the link nucleic acid molecules comprising the M barcodes is greater than a control, wherein the targets promote the disease. Embodiment 27. A method of disease diagnosis or prognosis, the method comprising: quantifying interaction and / or abundance of N targets within a cell from a patient sample with the method of embodiment 19, diagnosing the patient as having a disease or have increased risk of developing the disease if the amount of the link nucleic acid molecules comprising the M barcodes is lower than a control, wherein the targets are suppressors of the disease. Embodiment 28. The method of embodiment 26, wherein the control is the amount of link nucleic acid molecules comprising the M barcodes in a normal sample. Embodiment 29. A method of monitoring disease progression in a patient, the method comprising: quantifying interaction and / or abundance of N targets within a cell from a patient sample with the method of embodiment 19 at a first time point and a second time point during a time period, wherein the second time point is later than the first time point, determining that the disease has progressed if the amount of the link nucleic acid molecules comprising the M barcodes at the second time point is different from the amount of the link nucleic acid molecules at the first time point. Embodiment 30. A method of monitoring a patient’s response to a therapy, the method comprising: quantifying interaction and / or abundance of N targets within a cell from a patient sample with the method of embodiment 19 at a first time point and a second time point during a time period where the patient received a therapy, wherein the second time point is later than the first time point, determining that the patient is responsive to the therapy if the amount of the link nucleic acid molecules at the second time point is different from the amount of the link nucleic acid molecules at the first time point. 49 KILPATRICK TOWNSEND 797562351Embodiment 31. The method of detecting a tumor of embodiment 26, wherein the method detects the interaction of two or more targets. Embodiment 32. The method of detecting a tumor of embodiment 26, wherein the N targets comprise one or more targets in FIG.4. Embodiment 33. A kit for quantifying interactions and / or abundance of two or more of N targets, wherein the kit comprises N complexes of embodiment 1 and N complementary polynucleotides, wherein N is an integer, wherein each complementary polynucleotide comprises a complementary identifier being complementary to the unique identifier of one complex, wherein each complementary polynucleotide has ligatable ends. Embodiment 34. The kit of embodiment 33, wherein the polynucleotide of each complex comprises an overlap oligonucleotide, wherein each complementary polynucleotide has a universal complementary overlap that is complementary to the overlap oligonucleotide. Embodiment 35. A method of producing a complex of the composition of embodiment 1 by conjugating the binding domain to the encoding polynucleotide through a flexible linker, where the method comprises: providing the binding domain and the encoding polynucleotide in a molar ratio in a range from 0.5 to 6, modifying the internal nucleotide of the encoding polynucleotide with an amino functional group, treating the encoding polynucleotide with modified internal nucleotide with an SPDP– PEG–NHS ester to produce a reactive encoding polynucleotide, treating the binding domain with sulfo-SMCC to produce a reactive binding domain, and contacting the reactive encoding polynucleotide with the reactive encoding binding domain, thereby forming the complex. Embodiment 36. A system comprising: a processor; and a memory device including instructions executable by the processor for causing the processor to perform operations comprising: 50 KILPATRICK TOWNSEND 797562351determining a distribution of barcode combinations in a sample, wherein each barcode combination of the barcode combinations includes at least one barcode of a corresponding target; determining a contribution of a target complex to the distribution of the barcode combinations; and determining, based on a conversion that uses the contribution, an abundance of the target complex in the sample. Embodiment 37. The system of embodiment 36, wherein the memory device further includes instructions executable by the processor for causing the processor to determine the abundance of the target complex by performing operations comprising: determining a longest combination barcode of the barcode combinations, wherein the longest combination barcode corresponds to a largest target complex; determining, based on an amount of the longest combination barcode in the distribution and a probability of barcode formation, a first amount of each of multiple barcode sub-combinations in the longest combination barcode contributed by the largest target complex; and determining an abundance of the largest target complex in the sample based on the first amount of the largest combination barcode and the probability of barcode formation. Embodiment 38. The system of embodiment 37, wherein the memory device further includes instructions executable by the processor for causing the processor to determine the abundance of the target complex by performing operations comprising: removing the first amount of each of the multiple barcode sub-combinations in the longest combination barcode contributed by the largest target complex from the distribution to generate an updated distribution; determining a next longest combination barcode in the updated distribution, wherein the next longest combination barcode corresponds to a next largest target complex; determining, based on an amount of the next longest combination barcode in the updated distribution and the probability of barcode formation, a second amount of each of the multiple barcode sub-combinations in the next longest combination barcode contributed by the next largest target complex; and determining an abundance of the next largest target complex in the sample based on the second amount of the next largest combination barcode and the probability of barcode formation. Embodiment 39. The system of embodiment 38, wherein the operations further comprise: 51 KILPATRICK TOWNSEND 797562351iteratively determining abundances of target complexes until a remaining longest combination barcode includes only a single barcode. Embodiment 40. The system of embodiment 37, wherein the probability of barcode formation corresponds to a likelihood of a single linkage forming in a link barcode molecule. Embodiment 41. The system of embodiment 40, wherein the link barcode molecule incorporates a unique molecular identifier (UMI). Embodiment 42. The system of embodiment 37, wherein the operations further comprise: determining the abundance of the largest target complex using a combinatorics equation. Embodiment 43. The system of embodiment 36, wherein the operations further comprise: determining an interaction map of constituents in target complexes based on the abundance of the target complex. Embodiment 44. The system of embodiment 36, wherein a number of constituents in the target complex is in a range of 1 to 100. Embodiment 45. The system of embodiment 36, wherein an amount of a link barcode molecule that is generated is a function of a barcode combination length, an abundance and a length of an originating target complex, and a probability of barcode formation. Embodiment 46. The system of embodiment 36, wherein one or more of targets are proteins, nucleic acids, lipids, small molecule compounds, and derivatives thereof. Embodiment 47. The system of embodiment 36, wherein the at least one barcode is a nucleic acid. Embodiment 48. A method of disease diagnosis or prognosis, the method comprising: determining an abundance of a target complex in a sample with the system of embodiment 36; determining that the abundance of the target complex is different than a control greater than a threshold difference; and diagnosing a patient as having a disease or having an increased risk of developing the disease based on the abundance of the target complex being different than the control greater than the threshold difference. 52 KILPATRICK TOWNSEND 797562351Embodiment 49. The method of embodiment 48, wherein the control is an amount of the target complex in a normal sample. Embodiment 50. One or more non-transitory computer-readable storage media storing instructions that, upon execution executable by one or more processors of a system, cause the system to perform operations comprising: determining a distribution of barcode combinations in a sample, wherein each barcode combination of the barcode combinations includes at least one barcode of a corresponding target; determining a contribution of a target complex to the distribution of the barcode combinations; and determining, based on a conversion that uses the contribution, an abundance of the target complex in the sample. Embodiment 51. The one or more non-transitory computer-readable storage media of embodiment 50, wherein the operations further comprise determining the abundance of the target complex by: determining a longest combination barcode of the barcode combinations, wherein the longest combination barcode corresponds to a largest target complex; determining, based on an amount of the longest combination barcode in the distribution and a probability of barcode formation, a first amount of each of multiple barcode sub-combinations in the longest combination barcode contributed by the largest target complex; and determining an abundance of the largest target complex in the sample based on the first amount of the largest combination barcode and the probability of barcode formation. Embodiment 52. The one or more non-transitory computer-readable storage media of embodiment 51, wherein the operations further comprise determining the abundance of the target complex by: removing the first amount of each of the multiple barcode sub-combinations in the longest combination barcode contributed by the largest target complex from the distribution to generate an updated distribution; determining a next longest combination barcode in the updated distribution, wherein the next longest combination barcode corresponds to a next largest target complex; determining, based on an amount of the next longest combination barcode in the updated distribution and the probability of barcode formation, a second amount of each of the multiple barcode sub-combinations in the next longest combination barcode contributed by the next largest target complex; and 53 KILPATRICK TOWNSEND 797562351determining an abundance of the next largest target complex in the sample based on the second amount of the next largest combination barcode and the probability of barcode formation. Embodiment 53. The one or more non-transitory computer-readable storage media of embodiment 52, wherein the operations further comprise: iteratively determining abundances of target complexes until a remaining longest combination barcode includes only a single barcode. Embodiment 54. The one or more non-transitory computer-readable storage media of embodiment 51, wherein the probability of barcode formation corresponds to a likelihood of a single linkage forming in a link barcode molecule. Embodiment 55. The one or more non-transitory computer-readable storage media of embodiment 54, wherein the link barcode molecule incorporates a unique molecular identifier (UMI). Embodiment 56. The one or more non-transitory computer-readable storage media of embodiment 51, wherein the operations further comprise: determining the abundance of the largest target complex using a combinatorics equation. Embodiment 57. The one or more non-transitory computer-readable storage media of embodiment 50, wherein the operations further comprise: determining an interaction map of constituents in target complexes based on the abundance of the target complex. Embodiment 58. The one or more non-transitory computer-readable storage media of embodiment 50, wherein a number of constituents in the target complex is in a range of 1 to 100. Embodiment 59. The one or more non-transitory computer-readable storage media of embodiment 50, wherein an amount of a link barcode molecule that is generated is a function of a barcode combination length, an abundance and a length of an originating target complex, and a probability of barcode formation. Embodiment 60. The one or more non-transitory computer-readable storage media of embodiment 50, wherein one or more of targets are proteins, nucleic acids, lipids, small molecule compounds, and derivatives thereof. 54 KILPATRICK TOWNSEND 797562351Embodiment 61. The one or more non-transitory computer-readable storage media of embodiment 50, wherein the at least one barcode is a nucleic acid. EXAMPLES 1. General methods

[0139] Cell culture. All human cell lines were obtained from American Type Culture Collection. A431, MDA-MB-231, HCT116, SKBR3, BT474 and MCF7 were grown in Dulbecco’s modified essential medium (HyClone) supplemented with 10% fetal bovine serum (FBS, Gibco) and 1% penicillin–streptomycin (Gibco). HCC827, H1975 and A2780 were grown in the Roswell Park Memorial Institute 1640 medium (HyClone) supplemented with 10% FBS and 1% penicillin– streptomycin. MCF10A was grown in mammary epithelial cell growth medium (Lonza). All cell lines were tested and free of mycoplasma contamination (MycoAlert Mycoplasma Detection Kit, Lonza, LT07-418).

[0140] Preparation of TETRIS units. Antibodies used in this study (Table 4) were selected based on the following criteria: 1) high-affinity monoclonal antibodies; 2) have well-characterized binding epitope; 3) are well-validated by the manufacturer (with documentations available on the manufacturer’s website) and / or published studies for relevant applications (e.g. immunofluorescence and flow cytometry). To prepare TETRIS units, antibodies were activated by reacting with sulfosuccinimidyl 4-(N-maleimidomethyl)cyclohexane-1-carboxylate (sulfo-SMCC, Pierce) at 50- fold molar excess in phosphate-buffered saline (PBS, Gibco) for 2 h at room temperature. Excess sulfo-SMCC was then removed through Zeba desalting columns (Pierce). DNA strands modified with internal amino groups (Integrated DNA Technologies; Table 2) were reacted with 50-fold molar excess SPDP–PEG–NHS ester (Quanta BioDesign, Nanocs). The DNA product was purified through desalting and activated by tris(2-carboxyethyl)phosphine reducing gel (Pierce) for 1 h at room temperature. After reduction, the activated product was concentrated by Amicon Ultra centrifugal filters (Millipore). The concentrations of the activated antibodies and PEG-modified DNA were measured by absorbance spectroscopy (NanoDrop One, Thermo Fisher). The activated antibodies and excess PEG-modified DNA were then mixed and incubated overnight at 4 °C to form the TETRIS units, before purification through Amicon Ultra centrifugal filters. The concentrations of the TETRIS units were measured by bicinchoninic acid assay (BCA assay, Pierce). To determine the density of DNA strands on antibodies, TETRIS units were prepared using DNA strands modified with both amino group and 6-FAM (Integrated DNA Technologies; Table 1). Fluorescence measurements were performed on the resultant products using a Spark 10M microplate reader (SparkControl, version 2.1) (Tecan). The hydrodynamic diameter of antibodies and TETRIS units were measured by Zetasizer 55 KILPATRICK TOWNSEND 797562351Nano ZS instrument (Zetasizer Nano Software, version 3.30) (Malvern). Their binding kinetics were measured by bio-layer interferometry (Pall Fortebio), through real-time assessment of their binding to immobilized protein antigens for 300 s (Octet Data Acquisition Software, version 9.0; Octet Data Analysis Software, version 9.0).

[0141] Enzymatic ligation and chemical ligation. For enzymatic ligation, TETRIS units and complementary DNA strands (complements) with phosphorylated 5’-ends (Integrated DNA Technologies; Table 1) were incubated with T4 DNA ligase (20,000 U / ml, New England Biolabs / NEB) for 10 min at room temperature. The ligase was inactivated by heating at 65 °C for 10 min. For chemical ligation, TETRIS units and DNA complements with azide and alkyne modifications at the 3’- and 5’-ends, respectively (Integrated DNA Technologies; Table 1), were mixed with copper(ii) sulfate (150 μM, Sigma-Aldrich), tris(3-hydroxypropyltriazolylmethyl)amine (1 mM, Lumiprobe) and sodium ascorbate (1.5 mM, Sigma-Aldrich) and incubated for 2 h at room temperature. After washing, the generated barcodes were liberated by adding sodium hydroxide solution (100 mM, Sigma-Aldrich), purified with desalting columns and analyzed by TaqMan assay.

[0142] Barcode evaluation via TaqMan assay. For TaqMan quantitative polymerase chain reaction (qPCR) analysis of the liberated TETRIS barcodes, 2 μl of sample mixture from the TETRIS assay was mixed with 300 nM of primers and TaqMan probe (Integrated DNA Technologies; Table 2) in TaqMan Fast Advanced Master Mix (Applied Biosystems) to a final volume of 10 μl. TaqMan qPCR analysis was performed on a QuantStudio 5 real-time PCR system (QuantStudio Design and Analysis Software, version 1.5) (Applied Biosystems) with the following cycling protocol: 1 cycle of 50 °C for 2 min, 1 cycle of 95 °C for 2 min, 40 cycles of 95 °C for 1 s and 60 °C for 20 s. Signals (cycle threshold, Ct) were normalized against that of IgG isotype control TETRIS units. The amplification efficiency of all primer and probe sets were validated by a serial dilution of their respective DNA templates.

[0143] Overlap optimization. To optimize the length of the repeating overlap, we mixed equimolar amounts of identifiers with DNA complements to form constructs bearing either blunt ends or sticky ends with overlaps of different lengths, from 2 to 10 nucleotides (Integrated DNA Technologies; Table 1). These constructs were enzymatically ligated (T4 DNA ligase, 20,000 U / ml) at different DNA concentrations (high concentration: 100 nM, low concentration: 5 nM) to evaluate barcode formation at different inter-unit spacing. Following heat inactivation, 10% denaturing polyacrylamide gel electrophoresis (PAGE) was conducted to determine the yield of barcode generation. The gel was stained with SYBR Gold and imaged using a ChemiDoc Touch imaging system (Image Lab Acquisition Software, version 2.3.0.07; Image Lab Analysis Software, version 6.0.1) (Bio-Rad). 56 KILPATRICK TOWNSEND 797562351

[0144] Model protein assembly. To prepare model protein complexes, bait proteins were first adsorbed onto solid supports (high-binding plates or polystyrene beads), before being incubated with excess prey proteins to ensure complete complex formation. For example, for the formation of two- protein complex (anti-EGFR–EGFR-DDK), bait protein EGFR with C-terminal DDK tag (EGFR- DDK, 3 μg / ml, OriGene) was mixed with spacer molecules (bovine serum albumin, BSA, Sigma) at 1,000-fold molar excess to allow good inter-complex distancing and adsorbed onto either 96-well plate (MaxiSorp, Thermo Scientific) or polystyrene beads (1 μm, Spherotech) overnight at 4 °C. The immobilization surfaces were then blocked with 2% w / v BSA in PBS for 2 h at room temperature. Prey protein humanized anti-EGFR antibody (Cetuximab) was added in excess to the adsorbed bait protein and incubated for 2 h at room temperature to ensure complete formation of the two-protein complexes. The immobilization supports were washed with 0.05% v / v Tween-20 (Sigma) in PBS (PBST) to remove excess unbound proteins. To prepare the model three-protein complex (anti- EGFR–EGFR-DDK–anti-DDK), EGFR-DDK was adsorbed and blocked as described above. Excess humanized anti-EGFR antibody (Cetuximab) and anti-DDK tag antibody were sequentially added and incubated for 2 h at room temperature for complex formation. The immobilization supports were washed with PBST to remove excess unbound proteins after each incubation step. To evaluate the formation and stability of the model protein complexes, we conjugated EGFR–DDK, anti-DDK antibody and anti-EGFR antibody with Alexa Fluor 405, 488 and 647 N-hydroxysuccinimide (Invitrogen), respectively, following the manufacturer’s protocol. We then assembled the three- protein complex (anti-EGFR–EGFR-DDK–anti-DDK) on a plate surface using these fluorescently- labeled proteins. Fluorescence signals were measured at different time points after complex formation, with PBST washing, to evaluate interaction stability. Respective protein concentrations were determined by intrapolating standard curves.

[0145] TETRIS assay workflow on model protein complexes. Assembled model protein complexes were incubated with TETRIS units (3 μg / ml) for 2 h at room temperature to facilitate antibody binding to distinct protein components. After washing with PBST, DNA complements (500 nM, Integrated DNA Technologies; Table 2) were added and incubated for 1 h at room temperature. Following washing with annealing buffer (50 mM Tris–HCl, 10 mM MgCl2, pH 7.0), the barcodes were generated enzymatically (T4 DNA ligase, 20,000 U / ml), liberated by adding sodium hydroxide solution (100 mM), purified with desalting columns and analyzed by TaqMan assay as previously described. Reaction conditions for TETRIS complement hybridization and barcode liberation were optimized for improved signal-to-noise ratio. Finally, we converted the TETRIS barcode distribution to resolve the composition and abundance of protein complexes. By taking into account the barcode formation probability, we performed combinatorics analysis and transformation to convert barcode 57 KILPATRICK TOWNSEND 797562351signals to protein signals (see Figure 20 and Figure 8 for details). Control TETRIS units bearing IgG isotype control antibodies and control identifiers were included in all experiments.

[0146] Spatial activity characterization. To evaluate the spatial activity of the TETRIS assay, we prepared TETRIS units with PEG linkers of different length (9.6 nm and 66.3 nm; Quanta BioDesign, Nanocs) and applied them to assay interacting proteins and differentially-separated non-interacting proteins. For interacting proteins, two-protein complex (anti-EGFR–EGFR-DDK) was assembled on a high-binding plate (MaxiSorp, Thermo Scientific) as previously described, using EGFR-DDK (OriGene) as the adsorbed bait protein and humanized anti-EGFR antibody (Cetuximab) as the prey protein. Protein separation between the interacting pair was determined from published value67. For non-interacting proteins, we used EGFR-DDK and humanized anti-HER2 antibody (Trastuzumab, Roche) as a pair of size-matched but non-interacting protein targets. To vary the distance between the non-interacting proteins, we prepared mixtures containing different amounts of BSA spacer molecules and used the protein hydrodynamic sizes to estimate average separation between the non- interacting protein pair. Specifically, equimolar amounts (20 nM) of the non-interacting proteins were mixed with varying amounts of BSA spacer molecules (0 to 0.75 mg / ml). The mixtures were adsorbed onto different wells in a high-binding plate. The immobilization plates were further blocked with 2% w / v BSA in PBS for 2 h at room temperature. Following washing with PBST, TETRIS units (anti- human IgG1 Fc and anti-DDK, 3 μg / ml) were added to the plated proteins and incubated for 2 h at room temperature. After washing with PBST, DNA complements (500 nM, Integrated DNA Technologies; Table 2) were introduced and incubated for 1 h at room temperature. As described above, the barcodes were generated by enzymatic ligation, liberated and purified with desalting columns, before analysis by TaqMan assay.

[0147] Evaluation of TETRIS performance. To evaluate the assay sensitivity, two-protein complexes (humanized anti-EGFR–EGFR-DDK) assembled on polystyrene beads (1 μm, Spherotech) were serially diluted and analyzed with the TETRIS assay as previously described (TETRIS units: anti-human IgG1 Fc and anti-DDK). For comparison, sandwich enzyme-linked immunosorbent assay (ELISA) and proximity extension assay (PEA) were performed on the prepared protein complexes with a pair of antibodies targeting distinct protein components in the protein complex. To perform the ELISA, capture antibodies (anti-human IgG1 Fc, 5 μg / ml) were adsorbed onto MaxiSorp ELISA plates and blocked with 2% w / v BSA in PBS for 2 h at room temperature. Protein complexes assembled on polystyrene beads were then added and incubated for 2 h at room temperature. After washing with PBST, detection antibodies modified with biotin (anti-DDK, 3 μg / ml) were added and incubated for 2 h at room temperature. Following incubation with streptavidin-conjugated horseradish peroxidase (HRP, Thermo Scientific) and chemiluminescent 58 KILPATRICK TOWNSEND 797562351substrate (Thermo Scientific), the resultant chemiluminescence intensity was measured by a Spark 10M microplate reader. To perform PEA, we prepared the proximity probe conjugates by tagging anti-human IgG1 Fc and anti-DDK antibodies with thiol-modified DNA strands (Integrated DNA Technologies; Table 3) as described in our previous publication28. PEA was then performed using these antibody–DNA conjugates following published protocol51. Briefly, protein complexes assembled on polystyrene beads were mixed with the proximity probe conjugates (pre-hybridized with the extension primer at 2:1 primer-to-antibody ratio) to a final concentration of 50 pM in 25 mM Tris–HCl, 4 mM EDTA, 0.016 mg / ml sheared salmon sperm DNA (Invitrogen), and incubated for at 37 °C for 1 h. The mixture was then mixed with 40 μM of each deoxynucleotide triphosphate and 62.5 U / ml T4 DNA polymerase (NEB) to initiate the extension reaction. The reaction was incubated at 37 °C for 20 min and heat-inactivated at 80 °C for 10 min. The extension product was analyzed by TaqMan assay. To evaluate the accuracy of the TETRIS assay in distinguishing mixtures of protein complexes, protein complexes were individually assembled (pure two- and three-protein complexes on polystyrene beads) before ratiometric mixing to generate complex compositions (mixtures of two- and three-protein complexes). These preparations were then characterized by the TETRIS assay (TETRIS units: anti-human IgG1 Fc, anti-rabbit IgG, anti-EGFR).

[0148] TETRIS assay workflow in cells. Cells were washed with cold Dulbecco’s phosphate- buffered saline (DPBS, HyClone), fixed with 4% formaldehyde (Pierce) for 15 min at room temperature and permeabilized with 0.1% v / v Triton X-100 (Sigma-Aldrich) for 10 min at room temperature to allow TETRIS units to access subcellular compartments. The permeabilized cells were then incubated with TETRIS units (3 μg / ml) to label distinct proteins for 1 h at 4 °C, and washed with 0.5% w / v BSA in PBS. The labeled cells were then incubated with DNA complements bearing a universal overhang and phosphorylated 5’-end for bidirectional hybridization (500 nM, Integrated DNA Technologies; Table 2). Cellular mixtures were incubated for 1 h at room temperature. Following washing with the annealing buffer, enzymatic ligation was performed and the resultant barcodes were recovered and analyzed by TaqMan assay or next-generation sequencing (see details below). Finally, we converted the TETRIS barcode signals to resolve the composition and abundance of protein complexes (see FIG.20 and FIG.8 for details). Control TETRIS units bearing IgG isotype control antibodies and control identifiers were included in all experiments.

[0149] Next-generation sequencing workflow. Cell samples were incubated with TETRIS units and DNA complements bearing unique molecular identifiers (UMIs) (Table 5). After TETRIS barcode generation, the samples were treated with sodium hydroxide solution (100 mM) to liberate the formed barcodes and the product mixture was purified through desalting columns. To enrich for long sequences, the purified barcodes were size-selected using SPRIselect beads (Beckman Coulter). 59 KILPATRICK TOWNSEND 797562351The bead-to-sample ratio was independently optimized. We then used Accel-NGS 1S Plus DNA Library Kit (Swift Biosciences) to prepare sequencing libraries following the manufacturer’s instructions. After library preparation, PCR amplification was performed with Q5 high-fidelity DNA polymerase (NEB). Thermal cycling condition consisted of 1 cycle of 95 °C for 3 min, 10 cycles of 95 °C for 15 s, 66 °C for 30 s and 72 °C for 15 s, 1 cycle of 72 °C for 5 min. The final libraries were quantified by qPCR with KAPA Library Quantification Kits (Roche). To characterize the amplified products and validate size enrichment, samples were run on a Bioanalyzer (Agilent 2100 Expert Software, version B.02.08.SI648) (Agilent) with High Sensitivity DNA Kit (Agilent). After passing quality control, the libraries were sequenced with HiSeq 4000 (HiSeq Software Suite, version 3.4.0) (Illumina), using 2 × 151 bp paired end sequencing runs. To account for size-enrichment of different combination barcodes, equal amounts of size-matched control oligonucleotides for 2-mer, 3-mer, 4- mer and 5-mer combination barcodes, respectively, were spiked into the sample mixtures and subjected to identical treatment. The enrichment efficiencies determined from these size-control sequences were used to normalize the combination barcode counts.

[0150] Sequencing data analysis. Each sequenced read comprised a set of sequencing adaptors, a sample label indicative of cell type, universal repetitive overhangs, a permutation of identifier sequences and their UMIs. Sequencing data were processed by Geneious Prime (version 2020) and data analysis was performed using SPSS Statistics (version 28). Adaptor sequences were first trimmed, paired-end reads merged and low-quality reads removed. Duplicated sequence reads with identical UMIs were also removed to account for PCR amplification. The identity and permutation of identifiers of each read was then mapped against a custom library comprising all identifier sequences. The reads were then grouped by their sample label and their counts were normalized against that of length-matched control oligonucleotides to account for different enrichment efficiencies during the library preparation process. The normalized combination barcode distribution was then converted to protein complex signals using the mathematical model described in FIG.20. Scripts used for the combination barcode analysis are available as Supplementary Code. The protein complex signals were further normalized against that of sample-matched IgG isotype controls to account for non-specific antibody binding. Subsequently, different-ordered protein interaction maps (based on the 2-mer, 3-mer, 4-mer and 5-mer protein complexes, respectively) were constructed for each sample to illustrate the diversity and abundance of different protein interactions. The map was constructed by defining each protein marker and each protein complex as a node, and each co- occurrence of protein marker in a protein complex as a line. The network graph was visualized using Gephi (version 0.9.2). Protein marker nodes were positioned in a circular layout, and protein complex nodes were positioned using force-directed layout. 60 KILPATRICK TOWNSEND 797562351

[0151] Comparisons with annotated protein interaction data. To perform comparative analyses with previously-annotated protein interaction data, we constructed positive reference sets (PRS) and negative reference sets (NRS) against known interactions available in five public databases (PINA68, STRING69, BioGRID70, IntAct71and CancerNET72) and a recent interaction database developed specifically for breast cancer by Kim et al.12under different selection stringency (FIG.34). The high- confidence PRS includes only congruent protein interactions reported across all of the databases; the general PRS includes protein interactions reported in at least 3 of the annotated databases. We then performed bootstrap sampling of the TETRIS barcodes to generate six replicates; the first three replicates were used as the training study and the remaining three for the validation study. Thresholding in the training study was performed to maximize detection of interactions in the PRS pairs while minimizing the number of positive-scoring NRS pairs73. This was applied consistently to the validation study.

[0152] Cross-link co-immunoprecipitation. Cross-link co-immunoprecipitation was performed using Pierce Crosslink IP Kit (Thermo Scientific) following the manufacturer’s protocol. Briefly, following protein cross-linking, cells were lysed in 25 mM Tris pH 7.4, 150 mM sodium chloride, 1 mM EDTA, 1% v / v NP-40, 5% v / v glycerol and protease inhibitors (Thermo Scientific). The lysates were centrifuged at 4 °C and the supernatants were collected and quantified by BCA assay. The lysates were incubated overnight at 4 °C with polystyrene beads (3 μm, Spherotech) pre-coated with capture antibodies. After washing, the mixtures were incubated with 5 μg / ml sandwich detection antibodies (Table 4) for 1 h at 4 °C. Following washing, FITC-conjugated secondary antibody (BD Biosciences) was added and incubated for 30 min at 4 °C. The labeled beads were analyzed using a CytoFLEX flow cytometer (CytExpert, version 2.0) (Beckman Coulter).

[0153] Western blotting. Cells were lysed in radioimmunoprecipitation assay (RIPA) buffer containing 50 mM Tris–HCl, 150 mM sodium chloride, 1% v / v Triton X-100, 0.5% w / v sodium deoxycholate, 0.1% w / v sodium dodecyl sulfate and protease inhibitors. The lysates were centrifuged for 10 min at 4 °C and the supernatants were collected and quantified by BCA assay. Lysate proteins were separated by sodium dodecyl sulfate polyacrylamide gel electrophoresis (SDS–PAGE), transferred onto polyvinylidene fluoride membrane (PVDF, Invitrogen), and immunoblotted with primary antibodies against protein markers (phospho-EGFR (Y1068, Cell Signaling, 1:1000); phospho-Gab1 (Y627, Cell Signaling, 1:1000); phospho-PLCγ1 (Y783, Cell Signaling, 1:1000); phospho-Akt (S473, Cell Signaling, 1:2000); phospho-SHC (Y239 / 240, Cell Signaling, 1:1000); phospho-p44 / 42 MAPK (Erk1 / 2) (T202 / Y204, Cell Signaling, 1:2000); GAPDH (Cell Signaling, 1:1000)). Following incubation with horseradish peroxidase-conjugated secondary antibody (Cell 61 KILPATRICK TOWNSEND 797562351Signaling) and chemiluminescent substrate (Thermo Scientific), the resultant chemiluminescence was detected by a ChemiDoc Touch imaging system.

[0154] Flow cytometry. Following fixation and permeabilization, as previously described, cell suspensions were incubated with 5 μg / ml primary antibodies (Table 4) for 1 h at 4 °C. The cells were then washed with 0.5% w / v BSA in PBS and labeled with 2 μg / ml FITC-conjugated secondary antibody (BD Biosciences) for 30 min at 4 °C. After washing, the cells were analyzed using a CytoFLEX flow cytometer. Mean fluorescence intensity of all cells, excluding debris, was determined using FlowJo (version 10.6.2), and marker expression signals were normalized against that of IgG isotype control antibodies.

[0155] Immunofluorescence. Cells for immunofluorescence were cultured on an 8-well chamber slide (Nunc). The cells were fixed, permeabilized and blocked with 5% w / v BSA in PBS for 1 h at room temperature. The cells were then incubated with 5 μg / ml primary antibodies (EGFR, Abcam; SHC, BD Biosciences; GRB2, BD Biosciences; p53, BD Biosciences; MDM2, Invitrogen; MDMX, Merck) overnight at 4 °C, washed with PBS, before being incubated with 2 μg / ml secondary antibodies (Molecular Probes) for 1 h at room temperature. After washing, the cells were stained with nuclear dye Hoechst 33342 (Molecular Probes) and mounted (Vector Laboratories). Fluorescence images were acquired using a Leica DMi8 microscope (Leica Application Suite X, version 3.4.2) (Leica Microsystems) at 20× magnification. Image analysis was performed using ImageJ (version 1.52f).

[0156] Clinical samples. The study was approved by the National Healthcare Group Domain Specific Review Board (Ref: 2014 / 01088). All subjects were recruited according to IRB-approved protocols after obtaining informed consent. Breast fine needle aspiration (FNA) biopsy samples (n = 72) were collected and analyzed (Table 8). Specifically, tissue specimens were collected from patients with breast cancer during surgery; patient-matched samples were collected from the same patients and categorized into tumor, border and normal tissues by a board-certified pathologist. On these pathology-verified samples, we performed FNA biopsy, by repeat insertion and aspiration with a 21- gauge needle, to retrieve cells from the bulk tissue74. This procedure was performed three times for each tissue specimen (three biological replicates). Individual collected aspirates were suspended in DPBS, fixed, permeabilized and assayed through the TETRIS workflow, as previously described. Clinical diagnoses (including molecular subtyping) were established from gold-standard pathology reports. Cancer aggressiveness was determined through histologic grading of tumor tissues according to the Nottingham grading system75. Specifically, tissue specimens were stained with haematoxylin and eosin stain (H&E stain) and graded by a board-certified pathologist based on three key morphological features: 1) degree of tubule or gland formation, 2) nuclear pleomorphism and 3) 62 KILPATRICK TOWNSEND 797562351mitotic count. All samples were anonymized and TETRIS measurements were conducted blinded from the clinical results.

[0157] Statistical analyses. Unless otherwise stated, all measurements were performed in triplicate, and the data are presented as mean ± standard deviation. We performed Shapiro–Wilk tests to evaluate data normality and one-way ANOVA to detect differences for three or more groups of measurements. For inter-sample comparisons, multiple pairs of samples were each tested via two-tailed unpaired Student’s t-test or Mann–Whitney U test, and the resulting P values were adjusted for multiple hypothesis testing using Bonferroni correction. P < 0.05 was determined as significant. Distinct models / scores were developed for individual protein abundances, 2-mer and 3-mer complexes, respectively, through logistic regression with leave-one-out cross validation, for classifying cancer vs. control samples (Fig. 5b) as well as more aggressive cancer vs. less aggressive cancer samples (Figs. 5d–e). The receiver operating characteristic (ROC) curve was generated from the patient profiling data and constructed by plotting sensitivity versus (1 − specificity), and the value of area under the curve (AUC) was computed using the trapezoidal rule. Detection sensitivity, specificity and accuracy were calculated using standard formulas. Statistical analysis and data presentation were performed using Python (version 3.6), SPSS Statistics (version 28) and GraphPad Prism (version 9.1.2). 2. The TETRIS technology

[0158] The TETRIS technology is designed to measure complex protein interactions directly in whole cells. It consists of three functional steps: protein labeling, TETRIS barcode generation and barcode distribution analysis (Fig. 1a). During protein labeling, TETRIS units – antibodies conjugated with single-stranded DNA through a polyethylene glycol (PEG) linker – are used to tag proteins in cells. Each conjugated DNA comprises a unique sequence (identifier), which encodes the protein identity, and a universal region (overlap), which enables hybridization with neighboring TETRIS units in close proximity. The DNA strand is conjugated to the PEG linker at its mid-internal position to facilitate hybridization on both ends of the DNA (FIG.6a). To generate TETRIS barcodes, single-stranded complementary DNA (complements) are subsequently added to enable bidirectional barcode growth. These complements bind and chain adjacent TETRIS units: each contains a unique complementary identifier, which hybridizes with the protein-encoding DNA sequence on matched TETRIS units, and a repetitive complementary overhang, which binds to the universal overlap of adjacent TETRIS units (FIG.6b). All complements have enzymatically- or chemically-ligatable ends to facilitate covalent linkages during barcode formation (FIG. 6c and Table 1). Through these independent interactions, the technology generates TETRIS barcodes that explicitly linearize both the molecular constituents (via the barcodes) and spatial interactions (via the repetitive overlaps) within 63 KILPATRICK TOWNSEND 797562351individual protein complexes. Only in interacting proteins, closely-spaced TETRIS units connect bilaterally; through in situ ligation of adjacent complements, identifying DNA barcodes elongate bidirectionally to encode comprehensive protein interactions. Non-interacting proteins, however, do not experience any barcode growth (FIG. 6d). Finally, DNA barcodes are liberated from protein- bound TETRIS units (FIG.6e) and analyzed in a high-throughput manner to decode complex protein interactions. As identical protein complexes can generate a mixture of fully- and partially-formed TETRIS barcodes and the same barcodes are produced from different-ordered protein complexes (i.e., heterogeneous and degenerate barcoding), we perform large-scale combinatorics transformation of the full distribution of generated DNA barcodes, to decode the composition and abundance of originating protein complexes and map different-ordered protein interactions (Fig. 1b). Detailed TETRIS experimental methodology (molecular structures and sequential workflow) and analytical methodology (combinatorics data transformation) can be found in FIG. 7. To validate its clinical utility, we applied the technology for in situ profiling of protein interactions across various subcellular compartments in scarce patient samples. We found that as compared to control samples, cancer cells show similar relative amounts of individual proteins and lower-order protein complexes (e.g., 2-mer protein complexes) but more diverse and more abundant higher-order protein complexes (e.g., 4-mer protein complexes), and that these higher-order protein complexes can be used to more effectively distinguish patient disease aggressiveness (FIG.1c–d). 3. Assessment of TETRIS performance

[0159] We designed the TETRIS unit as a templated antibody–PEG–DNA hybrid to achieve interaction-driven bidirectional barcode growth. In developing the technology, we first optimized the design of the DNA system for barcoding. Using blunt-end and sticky-end duplex DNA units that bear different-sized universal overlaps, we performed ligation reactions at different DNA concentrations, to evaluate barcode ligation at varying inter-unit spacing and optimize the universal overlap length (FIG. 14a and Table 1). We determined that a universal DNA overlap of 6 nucleotides had the optimal balance (i.e., high barcode signal with closely-spaced units and low background signal with far-spaced units). Importantly, while the blunt-end units formed only dimer DNA products, the optimized overlap design generated much longer barcodes, confirming its ability for bidirectional barcode formation (Fig.2a and FIG.14b).

[0160] With the optimized overlap design, we next developed the hybrid TETRIS unit to achieve precise spatial control. Specifically, we conjugated antibodies with single-stranded DNA, via a PEG linker to the mid-internal position of the DNA (FIG.6a). The conjugation was both density- and size- tuned. For density optimization, we prepared TETRIS units with different DNA copies per antibody (Table 1). In these reactions, we evaluated their resultant antibody performance (FIG. 15) and 64 KILPATRICK TOWNSEND 797562351barcode formation when the TETRIS units were differentially separated (FIG. 16). The optimized DNA density not only preserved antibody performance but also maximized barcoding performance. For size optimization, we first prepared TETRIS units with different-sized PEG linkers (FIG. 17a and Table 2). When applied to measure different protein separation (i.e., interacting proteins as well as non-interacting proteins separated by increasing amounts of spacer molecules), the TETRIS units showed distinct spatial activity profiles (Fig.2b). Upon the addition of DNA complements, the short- PEG TETRIS units readily formed barcodes when the target proteins were interacting, and demonstrated a sharp signal decrease with increasing protein separation, while the long-PEG units showed consistent barcode signals, even with increasing protein separation. Subsequent size optimization through antibody conjugation with different-length DNA further confirmed that the selected TETRIS unit design – antibody bearing a single-stranded DNA of 46 nucleotides via a short PEG linker – showed the best signal-to-noise ratio (FIG. 17b). This design was thus used for subsequent development of TETRIS assay conditions (FIG.18).

[0161] Employing the optimized conditions, we evaluated the TETRIS assay performance. To assess its specific detection, we applied the TETRIS assay to various control configurations, including no-target control, non-interacting control and blocked-interaction control; the assay generated strong barcode signals only when specific target proteins were interacting, and showed negligible signals for the control configurations (FIG. 19). In a titration experiment, the TETRIS assay achieved ~1000- fold and ~150-fold improvement in its limit of detection (LOD = 1.1×10–18mol) compared to the gold-standard enzyme-linked immunosorbent assay and proximity extension assay, respectively (Fig. 2c and Table 3), and showed a good correlation (R2= 0.9875) (Fig. 2d). We further assessed the TETRIS assay’s capacity for measuring more complex protein interactions. Specifically, by considering the barcode formation efficiency, we developed combinatorics analysis (FIG.20) to 1) derive TETRIS barcode distributions from different-ordered protein interactions and 2) decode constituent protein interactions from an ensemble barcode distribution measured (FIG.8). To develop and evaluate the process, we assembled a two-protein model complex (anti-EGFR–EGFR-DDK) and a three-protein model complex (anti-EGFR–EGFR-DDK–anti-DDK). We chose these protein components (EGFR-DDK, anti-EGFR and anti-DDK) to assemble the model complexes because of their known strong binding affinities with each other36,37and stable interactions (FIG. 21a). Using known amounts of these model protein complexes, we numerically predicted and experimentally validated their respective DNA barcode distributions (Fig. 2e, FIG. 21b, FIG. 22 and Table 2). Likewise, by analyzing the DNA barcode distribution generated from complex mixtures (i.e., mixtures of different-ordered model protein complexes), we evaluated the composition and 65 KILPATRICK TOWNSEND 797562351abundance of different protein interactions (Fig. 2f). Upon data transformation, the TETRIS assay could accurately reveal constituent protein interactions of different orders in heterogeneous mixtures.

[0162] TETRIS measurement of cellular protein complexes. To evaluate the technology for cellular studies, we applied the TETRIS assay to measure biological protein interactions directly in whole cells. We chose to measure the cellular protein complex EGFR–GRB2–SOS1 (Fig. 3a), due to its well-established unique interactions and known controllability upon epidermal growth factor (EGF) stimulation38,39. Specifically, we employed the TETRIS assay to measure the three-protein complex, corresponding two-protein complexes and individual protein abundances (Fig. 3b, FIG. 23a and Table 4). The TETRIS experiments were performed in situ in different cell lines treated with EGF (enhanced protein interactions through stimulation) and in untreated control cells (weakened protein interactions at baseline); the measured TETRIS DNA barcode distribution (before data transformation) can be found in FIG.23b. As a gold-standard comparison, we performed cross-link co-immunoprecipitation (co-IP) to measure the protein complexes and individual proteins in treated and control cell lysates. Both the TETRIS and co-IP measurements showed that while the EGF stimulation did not alter the abundance of individual proteins (FIG.23a), it increased the extent of protein interactions (Fig.3b). Importantly, the TETRIS assay demonstrated a good correlation (R2= 0.9851) with the gold-standard analysis, across different cell lines and treatments, indicating reliable measurements of protein complexes and individual protein abundances (Fig.3c).

[0163] Next, we expanded the TETRIS evaluation to assess biological protein interactions (found in various subcellular locations) as well as their time-dependent changes in different cell lines undergoing targeted treatments. Using EGFR-based interactions38and tumor protein P53 (p53)-based interactions40as representative cytoplasmic and nuclear protein complexes, respectively (FIG. 24), we induced dynamic protein interaction changes through targeted cellular treatments (i.e., EGF to activate EGFR signaling41, FIG.25a–b and e; and oxaliplatin to initiate DNA damage repair by p53 cascade42, FIG.25c–d). Specifically, we applied the TETRIS assay to measure dynamic interaction changes in cells upon signaling activation (FIG.26a). In these studies, the TETRIS assay continued to demonstrate reliable measurements against gold-standard analysis (FIG.27a). Even in cases where there were no apparent changes in individual protein abundances upon cellular stimulation, the TETRIS assay could effectively captured dynamic protein interaction changes at different time points (e.g., increase in amount of protein complexes and decrease to baseline) (FIG.27b). These measured protein interaction changes are consistent with published studies on respective signaling activation and recovery dynamics43. More interestingly, in both the EGFR- and p53-based interaction systems, the higher-order three-protein complexes showed not only the most significant fold increases but also the earliest changes, demonstrating the dynamics of multiprotein interactions in orchestrating 66 KILPATRICK TOWNSEND 797562351different cellular activities (FIG. 26b). Indeed, by evaluating compositional changes in protein interactions upon cellular activation (Fig.3d), we observed that across all tested cell lines the three- protein complexes showed the most significant relative fold increases upon stimulation, as compared to individual proteins or lower-order two-protein complexes (Fig.3e). These cell line measurements of biological protein complexes thus not only validate the accuracy of the TETRIS assay, but also illustrate the responsive nature of higher-order protein interactions as indicative biomarkers of cellular activities.

[0164] Massively-parallel analysis of different-ordered protein interactions. We further advanced the technology to achieve highly-multiplexed interaction analysis through barcode sequencing. The workflow is presented in FIG. 9. To amplify DNA barcodes and eliminate duplicates, we incorporated unique molecular identifiers (UMIs) into the TETRIS workflow; complements bearing 10-nucleotide UMIs were added to identify individual protein complexes (FIG. 28 and Table 5). After ligation, UMI-tagged TETRIS barcodes were liberated from respective protein complexes and size-enriched for long barcodes to facilitate subsequent sequencing analysis (FIG.s 29-30). To account for this enrichment, we spiked all barcode samples with control oligonucleotides (i.e., size-matched to respective barcodes) (FIG. 31 and Table 6) to determine the enrichment efficiencies of different-sized barcodes and decipher the target barcode abundance. Following barcode sequencing, reads were de-duplicated through UMI matching and sorted according to their unit length; all barcode counts were normalized against their respective size-matched controls. Finally, we used the normalized barcode distribution for data transformation, so as to identify and quantify different-ordered protein interactions. All protein interactions are presented as network graphs to illustrate the abundance and composition of protein complexes, by defining protein markers as peripheral nodes, protein complexes as central nodes, and each confirmed TETRIS interaction in a protein complex as a line (FIG.9).

[0165] Employing the massively-parallel TETRIS approach, we evaluated complex interactions among 34 distinct proteins in various human breast cancer and control cells (Table 5). For each cell line tested, we performed barcode combinatorics transformation (FIG.32), so as to characterize the diversity and abundance of target proteins and protein interactions (i.e., individual proteins and different-ordered protein interactions such as 2-mer, 3-mer, 4-mer and 5-mer protein complexes). For diversity comparison, while the cancer and control cells showed a similar diversity of 2-mer protein complexes, the cancer cells had much more diverse populations of higher-order protein complexes (3-mer and above) (Fig.4a and FIG.10a). For abundance comparison, while the cancer cells showed similar amounts of individual proteins or low-order protein complexes (2-mer protein complexes) with respect to the control cells, they consistently showed more abundant higher-order protein 67 KILPATRICK TOWNSEND 797562351complexes (3-mer and above) across different complexes identified (Fig.4b and FIG. 10b). Using these multiplexed interaction measurements, we evaluated the performance of the TETRIS assay with known binary protein interactions found in public databases. The TETRIS measurements showed a good agreement (FIG. 33 and Table 7); 38–44% of the TETRIS high-abundance two-protein complexes are represented in the public databases (vs. ~22% as reported by other large-scale protein interaction studies12). Importantly, leveraging previously-annotated protein interaction data, we constructed positive and negative reference interaction sets (FIG. 34) and established the TETRIS performance characteristics using aggregated data from all five cell lines used in this study (FIG.11). As shown in these analyses, when comparing against published large-scale interaction assays for detecting two-protein complexes, the TETRIS technology showed superior performance. Likewise, it demonstrated consistent performance for cell line-specific comparisons; by evaluating protein interactions found in BT474 vs. MCF10A cells, the TETRIS assay identified more overlapping than different pairwise protein interactions (FIG. 35), consistent with established features that both cell lines are derived from breast tissues but have varying protein expression profiles44. Beyond its accuracy in measuring two-protein interactions, the TETRIS assay further illustrates the distinguishing properties of higher-order protein complexes (Fig. 4c). The large diversity and high relative abundance of higher-order protein complexes, both indicative of strong cellular signaling activities45, potentiate their use as reflective disease biomarkers. Indeed, by profiling ER, PR and HER2-based protein complexes, we further validated the accuracy of the TETRIS technology to molecularly subtype different breast cancer cell lines (Fig. 4d). Using the TETRIS technology, we also detected in cancer cells, certain protein complexes were present in significantly lower abundance than in control cells, for example, PR–SOS1, LMO4–CYFIP1, TIF2–UBS3B.

[0166] Multiplexed profiling of clinical samples. To evaluate the clinical utility of the TETRIS technology in detecting complex protein interactions, we conducted a feasibility study. We aimed at addressing the following questions: (1) if TETRIS can be applied to measure complex protein interactions directly in scarce clinical samples, (2) to validate the accuracy of TETRIS profiling in detecting cancer, and (3) if different-ordered protein complexes can distinguish other clinical features such as disease aggressiveness.

[0167] We obtained fine needle aspiration biopsies from patients with breast cancer; for each patient, we evaluated patient-matched cancer, border and control samples (n = 48, Table 8). Based on our measured protein interactions (Fig.4a and FIG.10a) and published biomarker studies46,47, we applied the TETRIS technology to profile five key protein markers (HER2, ER, PR, SCRIB and WASF3) and their respective protein complexes (Fig. 5a). Interestingly, even for protein markers (e.g., WASF3, NCKAP1 and CYFIP1) that showed similar individual protein abundances across 68 KILPATRICK TOWNSEND 797562351samples, their protein complexes (e.g., WASF3–NCKAP1, WASF3–CYFIP1, WASF3–NCKAP1– CYFIP1) could effectively distinguish the cancer samples from the matched controls (FIG. 12), indicating that complex protein interactions could be used as independent disease biomarkers. We thus constructed respective logistic regression models based on different protein targets (i.e., individual proteins, 2-mer and 3-mer protein complexes) with leave-one-out cross-validation. As compared with gold-standard pathology, the TETRIS analyses demonstrated a high accuracy in cancer diagnosis (Fig.5b, area under the curve (AUC) = 0.9401 for total proteins; 0.9557 for 2-mer complexes; 0.9453 for 3-mer complexes). Furthermore, the TETRIS profiling of ER, PR and HER2 complexes could accurately classify the cancer samples into distinct molecular subtypes (luminal, non-luminal HER2-positive or triple-negative) (Fig. 5c, total proteins: 95.83%; 2-mer complexes: 91.67%; 3-mer complexes: 91.67%), thereby validating the TETRIS assay against clinical protein profiling.

[0168] Beyond technology validation, motivated by the responsive nature of higher-order protein interactions during cellular activities, we next applied the TETRIS technology to cancer specimens to characterize patient disease aggressiveness (n = 24, Table 8). As cancer aggressiveness was clinically determined through cellular morphology characterization, we investigated if different- ordered protein complexes could distinguish this independent disease feature. The TETRIS measurements showed that while the protein markers (individual protein abundances) universally failed to differentiate disease aggressiveness, the interaction markers (protein complexes formed by the same proteins) demonstrated better abilities to classify the clinical groups (FIG.13). Specifically, through regression models with leave-one-out cross-validation, we computed distinct composite scores using individual proteins, 2-mer and 3-mer protein complexes, respectively (Fig. 5d). The higher-order protein interactions (3-mer complexes) showed the best performance to distinguish disease aggressiveness (p = 4.458×10–5) while the individual proteins could not (p = 1.000) (Fig.5e). DISCUSSION

[0169] Motivated by the programmability of DNA nanotechnology, through nanostructure design and analytical advances, we developed the TETRIS technology to achieve in situ barcoding of complex protein interactions. As compared to current approaches to analyze protein interactions, TETRIS offers unique advantages with respect to accurate and informative measurements directly in cells. For example, while high-throughput complementation assays (e.g., sequencing-based yeast two-hybrid assays48–50) serve as gold-standard platforms for screening binary protein interactions, they require protein engineering, are performed on model organisms and cannot be readily administered on clinical samples. In comparison, TETRIS requires no genetic modification and enables explicit measurements of different-ordered protein interactions in scarce clinical specimens. 69 KILPATRICK TOWNSEND 797562351Likewise, as compared to conventional mass spectrometry-based proteomics12, which measures in native cell lysates, requires extensive processing on large sample amounts (e.g., cell lysis and affinity purification) and tends to suffer from false results (e.g., loss of weak protein interactions during processing), TETRIS is performed in situ in fixed whole cells, requires minimal processing and generates accurate results comparable to gold-standard assays. Finally, as compared to published proximity-based barcoding approaches, which use univalent barcodes (e.g., antibody–DNA) and thus measure only binary protein interactions51,52, TETRIS leverages a complex hybrid structure (antibody–PEG–DNA) that enables bidirectional fusion of DNA barcodes. Through integrated molecular design, assay workflow and data transformation, the TETRIS assay achieves not only reliable measurements against gold-standard binary assays, but also informative assessment of complex protein interactions of different orders.

[0170] The technology has the potential to be further expanded. As the current study focuses on protein–protein interactions and uses off-the-shelf antibodies to construct TETRIS units, choosing suitable antibodies (epitopes) is an important consideration; with well-validated and carefully selected commercial antibodies, we demonstrate that the TETRIS explicit measurements of higher- order protein interactions could capture interesting signaling dynamics and enhance sample classification accuracy, beyond the performance of individual proteins and two-protein complexes. To extend the scope of the technology and overcome potential challenges in antibody selection, the TETRIS assay could be readily expanded, through incorporating other affinity molecules such as aptamers and nucleic acids53,54, thereby enabling comprehensive and informative analysis of other macromolecular interaction networks (e.g., protein–RNA and protein–DNA interactions). Likewise, the use of small molecule drug probes as affinity elements could advance the platform for high- throughput screening of therapeutic agents, providing direct measurement of drug-induced complex (dis)assembly55. Motivated by recent technology developments to examine protein–protein interactions in live cells (e.g., proximity labeling approaches56and yeast two-hybrid assays48,49), TETRIS has the potential to be further developed to investigate dynamic interactions in live cells, while bypassing the need for genetic and / or protein engineering used in published studies. Further technical developments through the integration of enhanced intracellular delivery systems57, new chemistries for barcode labeling and ligation58,59, and advanced molecular assays60,61are likely to expand the technology’s capabilities to measure dynamic interactions, reveal spatiotemporal heterogeneity and establish comprehensive multi-omics interactomes32,62.

[0171] Clinically, the technology can be applied to achieve various biomedical innovations. Unlike conventional technologies which often require protein engineering and / or extensive processing, TETRIS measures protein interactions directly in patient samples, is highly scalable, and generates 70 KILPATRICK TOWNSEND 797562351robust and informative readouts. These complementary properties make the technology well-suited for clinical workflow, thereby expanding the clinical reach of protein interaction biomarkers. In particular, TETRIS supports versatile assay configurations to accommodate diverse clinical needs. In its scalable sequencing format, the technology presents extensive multiplexing capabilities to enable simultaneous interrogation of massive interaction networks; this not only captures comprehensive information, but also can expedite the development of new responsive clinical signatures (e.g., higher- order multiprotein interactions to treatment effects). In its user-friendly PCR format, TETRIS generates quick results, uses existing infrastructure of clinical laboratories, and can be readily adapted with current clinical workflow. With this versatility, we anticipate applying the technology to various patient samples (e.g., ascites and blood), in a spectrum of diseases (e.g., other cancers and neurodegenerative diseases), to evaluate composite signatures63. Further technical developments through the integration of advanced fluidics64,65and analytics66could accelerate large-scale clinical validation. REFERENCES 1. Huttlin, E. L. et al. Architecture of the human interactome defines protein communities and disease networks. Nature 545, 505–509 (2017). 2. Luck, K. et al. A reference map of the human binary protein interactome. Nature 580, 402–408 (2020). 3. Irish, J. M., Kotecha, N. & Nolan, G. P. Mapping normal and cancer cell signalling networks: towards single-cell proteomics. Nat Rev Cancer 6, 146–155 (2006). 4. Wells, J. A. & McClendon, C. L. Reaching for high-hanging fruit in drug discovery at protein- protein interfaces. Nature 450, 1001–1009 (2007). 5. Scott, D. E., Bayly, A. R., Abell, C. & Skidmore, J. Small molecules, big targets: drug discovery faces the protein-protein interaction challenge. Nat Rev Drug Discov 15, 533–550 (2016). 6. Keskin, O., Tuncbag, N. & Gursoy, A. Predicting protein-protein interactions from the molecular to the proteome level. Chem Rev 116, 4884–4909 (2016). 7. Maurel, D. et al. Cell-surface protein-protein interaction analysis with time-resolved FRET and snap-tag technologies: application to GPCR oligomerization. Nat Methods 5, 561–567 (2008). 8. Kobayashi, H., Picard, L. P., Schönegge, A. M. & Bouvier, M. Bioluminescence resonance energy transfer-based imaging of protein-protein interactions in living cells. Nat Protoc 14, 1084–1107 (2019). 71 KILPATRICK TOWNSEND 7975623519. Galarneau, A., Primeau, M., Trudeau, L. E. & Michnick, S. W. Beta-lactamase protein fragment complementation assays as in vivo and in vitro sensors of protein protein interactions. Nat Biotechnol 20, 619–622 (2002). 10. Stelzl, U. et al. A human protein-protein interaction network: a resource for annotating the proteome. Cell 122, 957–968 (2005). 11. Sharma, K. et al. Proteomics strategy for quantitative protein interaction profiling in cell extracts. Nat Methods 6, 741–744 (2009). 12. Kim, M. et al. A protein interaction landscape of breast cancer. Science 374, eabf3066 (2021). 13. Skinnider, M. A. et al. An atlas of protein-protein interactions across mouse tissues. Cell 184, 4073–4089.e17 (2021). 14. Qin, W., Cho, K. F., Cavanagh, P. E. & Ting, A. Y. Deciphering molecular interactions by proximity labeling. Nat Methods 18, 133–143 (2021). 15. Dunham, W. H., Mullin, M. & Gingras, A. C. Affinity-purification coupled to mass spectrometry: basic principles and strategies. Proteomics 12, 1576–1590 (2012). 16. Seeman, N. C. & Sleiman, H. F. DNA nanotechnology. Nat Rev Mater 3, 1–23 (2017). 17. Liu, N. & Liedl, T. DNA-assembled advanced plasmonic architectures. Chem Rev 118, 3032– 3053 (2018). 18. Jones, M. R., Seeman, N. C. & Mirkin, C. A. Programmable materials and the nature of the DNA bond. Science 347, 1260901 (2015). 19. Sobczak, J. P., Martin, T. G., Gerling, T. & Dietz, H. Rapid folding of DNA into nanoscale shapes at constant temperature. Science 338, 1458–1461 (2012). 20. Lin, C. et al. Submicrometre geometrically encoded fluorescent barcodes self-assembled from DNA. Nat Chem 4, 832–839 (2012). 21. Song, P. et al. Programming bulk enzyme heterojunctions for biosensor development with tetrahedral DNA framework. Nat Commun 11, 838 (2020). 22. Zakeri, B. & Lu, T. K. DNA nanotechnology: new adventures for an old warhorse. Curr Opin Chem Biol 28, 9–14 (2015). 23. Jani, M. S., Veetil, A. T. & Krishnan, Y. Precision immunomodulation with synthetic nucleic acid technologies. Nat Rev Mater 4, 451–458 (2019). 24. Fredriksson, S. et al. Protein detection using proximity-dependent DNA ligation assays. Nat Biotechnol 20, 473–477 (2002). 25. Tavallaie, R. et al. Nucleic acid hybridization on an electrically reconfigurable network of gold- coated magnetic nanoparticles enables microRNA detection in blood. Nat Nanotechnol 13, 1066– 1071 (2018). 72 KILPATRICK TOWNSEND 79756235126. Liu, X. et al. Complex silica composite nanomaterials templated with DNA origami. Nature 559, 593–598 (2018). 27. Ho, N. R. Y. et al. Visual and modular detection of pathogen nucleic acids with enzyme-DNA molecular complexes. Nat Commun 9, 3238 (2018). 28. Sundah, N. R. et al. Barcoded DNA nanostructures for the multiplexed profiling of subcellular protein distribution. Nat Biomed Eng 3, 684–694 (2019). 29. Sundah, N. R. et al. Catalytic amplification by transition-state molecular switches for direct and sensitive detection of SARS-CoV-2. Sci Adv 7, (2021). 30. Stuart, T. et al. Nanobody-tethered transposition enables multifactorial chromatin profiling at single-cell resolution. Nat Biotechnol 41, 806–812 (2023). 31. Chen, F. et al. Cellular macromolecules-tethered DNA walking indexing to explore nanoenvironments of chromatin modifications. Nat Commun 12, 1965 (2021). 32. Ambrosetti, E. et al. A DNA-nanoassembly-based approach to map membrane protein nanoenvironments. Nat Nanotechnol 16, 85–95 (2021). 33. Deng, Y. et al. Spatial-CUT&Tag: spatially resolved chromatin modification profiling at the cellular level. Science 375, 681–686 (2022). 34. Klein, A. M. et al. Droplet barcoding for single-cell transcriptomics applied to embryonic stem cells. Cell 161, 1187–1201 (2015). 35. Travaglini, K. J. et al. A molecular cell atlas of the human lung from single-cell RNA sequencing. Nature 587, 619–625 (2020). 36. Patel, D. et al. Monoclonal antibody cetuximab binds to and down-regulates constitutively activated epidermal growth factor receptor vIII on the cell surface. Anticancer Res 27, 3355–3366 (2007). 37. Zhao, H., Shen, A., Xiang, Y. K. & Corey, D. P. Three recombinant engineered antibodies against recombinant tags with high affinity and specificity. PLoS One 11, e0150125 (2016). 38. Avraham, R. & Yarden, Y. Feedback regulation of EGFR signalling: decision making by early and delayed loops. Nat Rev Mol Cell Biol 12, 104–117 (2011). 39. De, S., Dermawan, J. K. & Stark, G. R. EGF receptor uses SOS1 to drive constitutive activation of NFκB in cancer cells. Proc Natl Acad Sci U S A 111, 11721–11726 (2014). 40. Wade, M., Li, Y. C. & Wahl, G. M. MDM2, MDMX and p53 in oncogenesis and cancer therapy. Nat Rev Cancer 13, 83–96 (2013). 41. Citri, A. & Yarden, Y. EGF-ERBB signalling: towards the systems level. Nat Rev Mol Cell Biol 7, 505–516 (2006). 42. Wang, D. & Lippard, S. J. Cellular processing of platinum anticancer drugs. Nat Rev Drug Discov 4, 307–320 (2005). 73 KILPATRICK TOWNSEND 79756235143. Fey, D., Aksamitiene, E., Kiyatkin, A. & Kholodenko, B. N. Modeling of receptor tyrosine kinase signaling: computational and experimental protocols. Methods Mol Biol 1636, 417–453 (2017). 44. Kao, J. et al. Molecular profiling of breast cancer cell lines defines relevant tumor models and provides a resource for cancer gene discovery. PLoS One 4, e6146 (2009). 45. Vinayagam, A. et al. A directed protein interaction network for investigating intracellular signal transduction. Sci Signal 4, rs8 (2011). 46. Vieira, A. F. & Schmitt, F. An update on breast cancer multigene prognostic tests-emergent clinical biomarkers. Front Med (Lausanne) 5, 248 (2018). 47. Jafari, S. H. et al. Breast cancer diagnosis: Imaging techniques and biochemical markers. J Cell Physiol 233, 5200–5213 (2018). 48. Yachie, N. et al. Pooled-matrix protein interaction screens using Barcode Fusion Genetics. Mol Syst Biol 12, 863 (2016). 49. Trigg, S. A. et al. CrY2H-seq: a massively multiplexed assay for deep-coverage interactome mapping. Nat Methods 14, 819–825 (2017). 50. Schlecht, U., Liu, Z., Blundell, J. R., St Onge, R. P. & Levy, S. F. A scalable double-barcode sequencing platform for characterization of dynamic protein-protein interactions. Nat Commun 8, 15586 (2017). 51. Lundberg, M., Eriksson, A., Tran, B., Assarsson, E. & Fredriksson, S. Homogeneous antibody- based proximity extension assays provide sensitive and specific detection of low-abundant proteins in human blood. Nucleic Acids Res 39, e102 (2011). 52. Agasti, S. S., Liong, M., Peterson, V. M., Lee, H. & Weissleder, R. Photocleavable DNA barcode- antibody conjugates allow sensitive and multiplexed protein analysis in single cells. J Am Chem Soc 134, 18499–18502 (2012). 53. You, M. et al. Engineering DNA aptamers for novel analytical and biomedical applications. Chem Sci 2, 1003–1010 (2011). 54. Wu, L. R. et al. Continuously tunable nucleic acid hybridization probes. Nat Methods 12, 1191– 1196 (2015). 55. Pan, S. et al. Extracellular vesicle drug occupancy enables real-time monitoring of targeted cancer therapy. Nat Nanotechnol 16, 734–742 (2021). 56. Lam, S. S. et al. Directed evolution of APEX2 for electron microscopy and proximity labeling. Nat Methods 12, 51–54 (2015). 57. Mitchell, M. J. et al. Engineering precision nanoparticles for drug delivery. Nat Rev Drug Discov 20, 101–124 (2021). 58. Parker, C. G. & Pratt, M. R. Click chemistry in proteomic investigations. Cell 180, 605–632 (2020). 74 KILPATRICK TOWNSEND 79756235159. Marks, K. M. & Nolan, G. P. Chemical labeling strategies for cell biology. Nat Methods 3, 591– 596 (2006). 60. Gu, L. et al. Multiplex single-molecule interaction profiling of DNA-barcoded proteins. Nature 515, 554–557 (2014). 61. Dan, K., Veetil, A. T., Chakraborty, K. & Krishnan, Y. DNA nanodevices map enzymatic activity in organelles. Nat Nanotechnol 14, 252–259 (2019). 62. Liu, Y. et al. High-spatial-resolution multi-omics sequencing via deterministic barcoding in tissue. Cell 183, 1665–1681.e18 (2020). 63. Lim, C. Z. J., Zhang, L., Zhang, Y., Sundah, N. R. & Shao, H. New sensors for extracellular vesicles: insights on constituent and associated biomarkers. ACS Sens 5, 4–12 (2020). 64. Duncombe, T. A., Tentori, A. M. & Herr, A. E. Microfluidics: reframing biological enquiry. Nat Rev Mol Cell Biol 16, 554–567 (2015). 65. Fordyce, P. M. et al. De novo identification and biophysical characterization of transcription- factor binding sites with microfluidic affinity analysis. Nat Biotechnol 28, 970–975 (2010). 66. Chen, Y. et al. Collaborative equilibrium coupling of catalytic DNA nanostructures enables programmable detection of SARS-CoV-2. Adv Sci (Weinh) 8, e2101155 (2021). 67. Li, S. et al. Structural basis for inhibition of the epidermal growth factor receptor by cetuximab. Cancer Cell 7, 301–311 (2005). 68. Du, Y. et al. PINA 3.0: mining cancer interactome. Nucleic Acids Res 49, D1351–D1357 (2021). 69. Szklarczyk, D. et al. The STRING database in 2021: customizable protein-protein networks, and functional characterization of user-uploaded gene / measurement sets. Nucleic Acids Res 49, D605–D612 (2021). 70. Oughtred, R. et al. The BioGRID database: a comprehensive biomedical resource of curated protein, genetic, and chemical interactions. Protein Sci 30, 187–200 (2021). 71. Orchard, S. et al. The MIntAct project-IntAct as a common curation platform for 11 molecular interaction databases. Nucleic Acids Res 42, D358–63 (2014). 72. Meng, X. et al. CancerNet: a database for decoding multilevel molecular interactions across diverse cancer types. Oncogenesis 4, e177 (2015). 73. Braun, P. et al. An experimentally derived confidence score for binary protein-protein interactions. Nat Methods 6, 91–97 (2009). 74. Casaubon, J. T., Tomlinson-Hansen, S. & Regan, J. P. in StatPearls (StatPearls Publishing, Treasure Island (FL), 2023). 75. Dalton, L. W., Page, D. L. & Dupont, W. D. Histologic grading of breast carcinoma. A reproducibility study. Cancer 73, 2765–2770 (1994). 75 KILPATRICK TOWNSEND 79756235176. Shatsky, M. et al. Quantitative tagless copurification: a method to validate and identify protein- protein interactions. Mol. Cell Proteomics 15, 2186–2202 (2016). 77. Gavin, A. C., Maeda, K. & Kühner, S. Recent advances in charting protein-protein interaction: mass spectrometry-based approaches. Curr. Opin. Biotechnol.22, 42–49 (2011). 78. Armean, I. M., Lilley, K. S. & Trotter, M. W. Popular computational methods to assess multiprotein complexes derived from label-free affinity purification and mass spectrometry (AP- MS) experiments. Mol. Cell Proteomics 12, 1–13 (2013). 76 KILPATRICK TOWNSEND 797562351Table 1. Sequences for TETRIS characterization. Ligation characterization S IFAM / * For enzymatic ligation, the complements were modified with a phosphate group at the 5’-end ( / 5Phos / ). For chemical ligation, the complements were modified with an alkyne group at the 5’-end ( / 5Hexynyl / ) and an azide group at the 3’-end ( / 3AzideN / ). Note: Underlined bases indicate the overhang. 77 KILPATRICK TOWNSEND 797562351Table 2. TETRIS barcodes and their TaqMan qPCR probes and primer sets.78 KILPATRICK TOWNSEND 79756235179 KILPATRICK TOWNSEND 797562351Table 3. Sequences for proximity extension assay.80 KILPATRICK TOWNSEND 797562351Table 4. Protein markers and antibodies used. Protein markers Descriptions Antibodies 7S 3 a role in the regulation of cell morphology and cytoskeletal3B7D10 81 KILPATRICK TOWNSEND 797562351organization. It is associated with the movement, invasion and c- 3 e -involved in cell signaling and protein ubiquitination.82 KILPATRICK TOWNSEND 79756235183 KILPATRICK TOWNSEND 797562351Table 5. Sequences for decoding individual proteins and protein complexes. Identifiers for sequencingIdentifier 31 RAF1 ACTCGTATCCTCT / iUniAmM / TTAATGTCGGGTT 84 KILPATRICK TOWNSEND 797562351Identifier 32 PTPN11 ACTCGTAGTATCT / iUniAmM / GAAGGACGACATCComplement 29 / 5Phos / ACGAGTNNNNNNNNNNTGGGTTGCAGCTTTACCGTA 85 KILPATRICK TOWNSEND 797562351Note: N denotes a random base. 86 KILPATRICK TOWNSEND 797562351Table 6. Size controls and their respective qPCR primer sets.Note: N denotes a random base. 87 KILPATRICK TOWNSEND 797562351KITLPaACbTBleRCTBI TK47447T4OWNSEND7U97B5 S P63R2 B;35; c1c-- MMeettQ 8PT 0F 64 42 0; 1P ;0P8 05 88 51 81N o N voelvelN o N voelvelN o N voelveld d d d d d d d d d d DtopuN R R R R R RSoN o N oepN o N o NepN o N N N N N Nep epNepN NepTblRicvelvelveolrtevev oeveor v ov ov ov ov ov ov o o ov o ov ov o Idad l l l te e e e e e e erdl l l l l l l terdte erdl te e erdl l te NdGtabaseR R R R R R R R R RsN o Nev oelv p eeo pN Nep epN N Nep ep epNep epNepN N R ep Plrtoer ov ov or o ov ov ov o o o ov o o ov o ov ov oINd te e ed l l terdte e e erdl l l terdterdte erdl terdte erdl te e erdl l teA d R R R R R R R RSN o Nev oelv p eeo pN Neo olrtertev oelv p eeo pNeo olrtertev pNeeo olrtev p eeo p eo pN R e R e N R e N N R euoomlrtertertev peo po olrtertev peo ortev oev peorm tea8d d d d d d d d d d l d l l dry8KILPATRB BICTK4 T747T4 4OWNSEND V N7A9N O S75G16 LA231 P5;E;E1 RR RFRI F1I1Q 8 O TA75A0952;Q ;Q9U9UJMJM33N o N voelvelN o N voelvelN o N voelvelN N N N R N N N N N N N R N N N N R N R ov o o o epelvelvelveo olrtev o o o o o o ep o o o o ep o epN o N o N o N o N o elvelvelvelvelvelveolrtevelvelveveorteveorteveveveveved d l l d l d l l l l lN o N N R R N N N N R R N R N N N N R R R N N R N voelv o eelv p eeo plrtoer ov o o o ep ep o ep o o o o ep ep ep o o ep oN o ev v v or or v or v v v v or or or v v or v vdte e e ed l l l l ted te ed l te e e e ed l l l l ted ted te e ed l l te e ed l lN N N R R N R R R R R R R oo o ep ep oN o N Nep epNepN N N Nep ep epN NepN N vv v o o v v ov ov o o ov o ov ov ov ov o o o ov ov o ov oel el elrterdte e e e erdl l l l terte elrte el el el elrterterte el elrte elvel8d d d d d d d9KILPATRSIKSCBKBK R RT 3 3OWNSEND79 EH7GE56 FR2 R335;c ;-PL1 MeCtG 1P0 P0 25 13 83 6;P00;P8 15 98 11 74R e R peo prtoerdtedR e R peo prtoerdtedR e R peo prtoerdtedR epN o N o N o N o N o N o N o N o N R oepN o N R R R R o N o N oep epN oepN o N o N o NepN ortevelvelvelvelvelvelvelvelveolrtevelveveveortoerteveortevevev oeveor oteved d l l l d d l d l l l l d lR e R pepN R oepN R oepN o N o N o N R oepN o N o N R R R R R oep ep ep ep epN o N o N R o N oepN ortoerteveolrteveorteveveveveorteveveveortoertoertoertoerteveveveveor oteved d d l d l l l l d l l l d d d d d l l l l d lR e R pepN R oepN R R R R R R R R R oepN N N Nep epN Nep ep ep ep epN N N NepN o to o o o o o o o o o o o o o o o orertvertvertvevevevertrtvevertortortortortveveveveortved ed l ed l ed l l l l e e l l e e e e e l l l l e l9d d d d d d d d0KILPATRS SIKCBKBK R RT 3 3OWNSEND GP79L7RC5B G6232;15 H;1E SRT2AT3P6 P2 19 99 13 7;P40;P4 46 02 76 63R epN oorteved lR epN or oteved lR epN oorteved lR epN o N o N o N o N o N o N o N o N o N o N o N o N R oepN R o N o N o N oepN o N o N o N o N ortevelvelvelvelvelvelvelvelvelvelvelvelveolrteveveveveortevevevev oeved d l l l l d l l l l lR epN N N N N N N N R N N R R R N N N N R N R R R or otev o o o o o o o ep o o ep ep ep o o o o ep o ep ep epN o elvelveveveveveveorteveveortoertoerteveveveveorteveortoertoerteved l l l l l l d l l d d d l l l l d l d d d lR epN o N o N o N R R R R R R R R o N o N N NepN Nep ep epN N N NepNep ep epN ov v v v v ov ov ov o ov ov o o o ov ov ov ov o ov o o o orte e e e e e e e erdl l l l l l l l te e erdl l terterte el el el elrte elrtertertevel9d d d d d d d1KILPM MATD DRA AIC-M-MKB BTO-2 -2W31 31NSENDE P77RT9P5RN62F3I 11;E5 ;1 C RtIR PFI1Q 9PU18JM0331; ;Q Q999U70JM83N o N voelvelN o N voelvelN o N voelvelN N N N N N N N R N N N R R N N N R R R ov o o o o o o o epelvelvelvelvelvelvelveo olrtev oelv o eelv p eeo plrtoer otev oelv o eelv p eeo pN rtoer oN tev oN ev o eev pN eo oN rtev oeved d d l d d l l l d l lN o N N R R R R N R N N N R R R N N R R N N N R N N vo o ep eelvelveo p ep eplrtoertoertoer o etev peor otev oev o eev p eeo p eprtoertoer otev o eev p eeo prtoer otev oev o eev peor otev oeved d d d l d l l l d d d l l d d l l l d l lN N N R R R R R R R R R R R R oo o ep ep ep ep ep epN N Nep ep epN Nep epN N NepN N velvelveolrtoer o o o o ov ov ov o o o ov ov o o ov ov ov o ov ovdterdterdterdterte el el elrterterte el elrterte el el elrte el el9d d d d d d d d2KILPM MATD DRAIC- A M -MKB BTO-2 -W32131NSEND79 T L7 IM5 F6 2;O23 P5 L41 C;SGO1S1 Q 1P5 65 19 966;8P;1 Q901778489N o N voelvelN o N voelvelN o N voelvelN N N N N N N R N N N R R R oo o o o o o ep o o oN oep epN o N o N o NepN o N N N N N velvelvelvelvelvelveolrteveveveveor or v v v ov o v ov ov ov ov ovdl l l l ted te e e e erdl l l l te e e e e e ed l l l l l lN N N N N N N R N R R R R R ov oelv oelv oelv oelv oelv o eelv peolr oN tev oN ev oN ev o eev p eeo pN N N Ner o ov ov ov ov pN Neo ov ov pN Neo ov ov podl l l l terdte e e e erdl l l l te e erdl l te e erdl l tedN N R R R R R R ov oN elv oN elv oN elv oN Neelv oelv oelv pN N N Neeo olrtev oev oev oev p eeo pN N N Ner o ov ov ov o pN Neo o o pN Neo o o podl l l l terdte e e everdl l l l tevelvelrtevelvelrte9d d d3KILPATRM MICC CKFT7F7OWNSEND N79C KS756AHC23P51 ;; cc-- M1 MeettQ 9YP22 9A357 3;P ;0P8085 58 81 1N R oev peolrtedN R oev peolrtedN R oev peolrtedN R oeN R e R e R N N N N N N R R N N N N N R N N N N N vpo o po p eo po o o o o o o ep eo p eo o o o o o po o o o o oelrtever r r vevevevevever r vevevevever vev v v vdl ted ted ted l l l l l l ted ted l l l l l ted l el el el elN R e N R e R e R e R e N R e R e N N R R N N R N N R N N N R R ov peolr otev peo pr o pr o pr or ov peo pr or ov o ep ep o o ep o o ep o o o ep epeveor or vev or v v or v v v or ordl ted ted ted ted l ted ted l l ted te ed l l te e ed l l te e e ed l l l ted tedN R N R R R R N R R R N R R R R R R oev peo o ep ep ep ep ep ep ep ep epN NepN NepN N Nep eplrteveor or or o ov o o o ov o o ov ov o ov ov o ov ov ov o odl ted ted terdte erdl terdterte elrterte el elrte el elrte el el elrterte9d d d d d d d d4KILPATRM MICC CKF7FT7OWNSENDE79RT7RI5 F F62I31 2;P5 ;S1 T LA CT G 31Q 9 Q U1JM553 9; 6P ;P4017 96 13 74N o N voelvelN o N voelvelN o N voelvelN o N N N N N N N N R e N N N N N R e N N N R R R N N N volv o o o o o o o po o o o o o p eo o o o p eo p eo po o o oeelvelvelvelvelvelvelvelrtevelvelvelvelvelrtevelvelvelrtertertevelveved d d d d l lN N R e N N N R e N N R e R e N N N N R N N R R R R N N R ov o pelveolr otev o o pevev or ov ov po p er or ov ov ov ov por ov o ev p eo p ep epr or or or ov o ev pordl l el te e ed l l ted te e e e ed l l l l te e ed l l ted ted ted te e ed l l tedN N R N N N R R R R R R R R R ov o eelv peo o o o epN Nep epN N N NepN Nep ep ep epN Neplrteveveveo or v ov o o ov ov ov ov o ov ov o o o o ov ov odl l l te e erdl l terdte el el el elrte el elrterterterte el elrte9d d d d d d d5KILPATM MRC F CICK1 F1T0A 0AOWNSEND9 L F77M OX56O A234 15;S;S1 OSCR 1IBP6 P1 59 56381;Q 7;0Q718 481960N o N voelvelN o N voelvelN o N voelvelN o N N N N N R N N R N N N N N R N R N R N N N N vo o o o o epelvelvelvelvelveo olrtev o epelveolr otev o o o o ep eelvelvelvelveolr otev p eeo o plrteveor otev oev oev oN ev oeved d d d l d l l l l lN o N N R N N R N N R R R N N N R N R N R N N N N N vo o ep eelvelveolr otev o p eelveolr otev oelv p eeo p eprtoertoer otev oev o eev peor o etev peor o etev peor otev oev oev oev oeved d l d d d l l l d l d l d l l l l lN N N R R R R N N R R R N R R R oo o ep ep ep ep o o ep ep ep oN NepNepNepN N N N N velvelveolrto o oer v v o o o v ov ov o ov o ov o ov ov ov ov ovdterdterdte el elrterterte el el elrte elrte elrte el el el el el9d d d d d d d6KILPATM MRC F CICK1 F1T0A 0AOWNSEND79 S7HP56C23;RS;H5 CE1R IRB3P2 P9 035634;0Q1;1P42116 80 60N o N voelvelN o N voelvelN o N voelvelN o N N N R e N N N N N N R e R e N N N N N N N R e N R N N volv olv olv po o o o o o o po po o o o o o o o p eo o po o oee e elrtevelvelvelvelvelvelrtertevelvelvelvelvelvelvelrtevelrtevelved d d d d lN N N N R e R e N R e N N N R e R e N N N N N N R R R R N R ov o o o pelvelvelveo plrtoer ov por ov ov ov po p er or ov ov ov o o o p ep ep ep o epev v v or or or or v ordte ed l te e e ed l l l ted te e e e e ed l l l l l l ted ted ted te ed l tedN N N N R R N R N R R R R R R R ov oelv oelv o eelv p eeo p o ep oN Nep epN N N N N Nep ep ep epNeplrtoer v o v ov ov o o ov ov ov ov ov ov o o o o ov odte erdl te e e erdl l l terte el el el el el elrterterterte elrte9d d d d d d d7KILPATRICKTOWNSENDF79OG RT7XI5BF6A231 22;E5;P ;E1 L R RCR R GF FI I11 1PP6Q 55 2 13 9 5179;3596P ;Q ;Q1991U9U74JMJ3M3 N R oev pN eo olrteved lN R oepN vo oelrteved lN R oev pN eo olrteved lN R oeN vpo oelrteved lN R oeN vpeolr oteved lN R oepN vo oelrteved l98KILOPAt ch a I c I cD C R M B R M AtT TT e r nv ar nv a u a aMa g is o abR rcsinas cina rs cinctanc ng edi I ng ediesuta leIoivoivo lee an e an el b8C mKa elmoedmrtsy amre . CTbaO u ul cainpepalstlinWa tr al si stesicNualSEiN1nfD6Cor71()12 9 1(8 (.1 ( 07 ( 236.0 0.65.0 .08– 2 2 23 3 L6 (5u.–8 4. 0. m3% 7)% 0)% 0% 34 1 9 2 0i) ).4 .5 0%na)lS0 1 4( 0 2 5Nu0 (.2 (8 (00 0 1.Ho bt4 2 5 (2 E n-Nyp0.0. .0– 5 2– 6 0. R luumi%00 00 0 41 .6 714832mb ng)%)% %)).6 %+i) naelr(( F%i)g.T5c1)(0 6 0 11 5 7 rip4(.0 (82.9 05(.0 9.7.07– 2 2 6 (29le35..9–0. .1-ne% 0% 1 0))% % 0 1)). 7 5 7 g2 8 %)ative01 ( a(0 ( 5 0 2 2 6GLog.1 (06 8 (0.0 0.%637.3 .3 020– 25 3–%34 .98 6 (2 r91.55.a0dwg0e1g rerD ass i)%)%)).4 .5 %)–2 d) eNiu vmefn feb eree ssnt2 1 r(Fia(21 ( 4 0 1 12 8 (GH i(%igtio11 (.7 (0 9.111.1 7 .07– 23 9.8 6 (3 75ra gd h ).g5nd of% 1 .% 78 0% % 4)1 .3 –. 7 .5 .06 8 0%e3r –) ad ee )9) )9) )

Claims

WHAT IS CLAIMED IS:

1. A composition of N complexes identifying N targets, each of the N complexes identifying a different target, wherein each complex comprises: 1) a binding domain, wherein the binding domain identifies a target, 2) a flexible linker, and 3) an encoding polynucleotide, wherein the encoding polynucleotide comprises a modified unit at an internal position, wherein the encoding polynucleotide is conjugated to the flexible linker via the modified unit, wherein the encoding polynucleotide comprises a unique identifier that identifies the binding domain and wherein the unique identifier identifies the target, wherein the binding domains of different complexes identify different targets, wherein the unique identifiers of different complexes identify different targets, thereby the unique identifiers of the N complexes identify the N targets, wherein optionally, the N targets are proteins, lipids, nucleic acids, small molecule compounds, and derivatives thereof, and wherein optionally, N is in a range of 1 to 100.

2. The composition of claim 1, wherein the modified unit is a phosphoramidite with an amine group.

3. The composition of claim 1, wherein the flexible linker is a polyethylene glycol (PEG), wherein optionally, the PEG linker has 2-200 units.

4. The composition of claim 1, wherein the encoding polynucleotide comprises an overlap oligonucleotide, and optionally, the overlap oligonucleotides in all the complexes have the same nucleic acid sequence.

5. The composition of claim 1, wherein the binding domain is an antibody, a peptide, a nucleic acid, a small molecule, or an aptamer.

6. The composition of claim 1, wherein the encoding polynucleotide is unligatable or ligatable. 100 KILPATRICK TOWNSEND 7975623517. The composition of claim 1, wherein: - the overlap oligonucleotide has a length of 0 to 100 nucleotides, - the unique identifier has a length of 5 to 100 nucleotides, - the linker has a length in a range from 2 nm to 100 nm, for example, in a range from 9.6 nm to 66.3 nm, and / or - the internal nucleotide of each encoding polynucleotide is located at least 5 nucleotides away from the closer of the two ends of the encoding polynucleotide.

8. A method of measuring one or more of N targets in a sample, wherein N is an integer, the method comprising: 1) contacting the sample with the N complexes of claim 1, 2) contacting the N complexes with N complementary polynucleotides, each of the N complexes identifying a different target, wherein each complementary polynucleotide comprises a barcode being complementary to the unique identifier of one of the N complexes, and wherein the N complexes comprise N barcodes, wherein optionally, each complementary polynucleotide comprises a UMI, quantifying the one or more of the N barcodes, thereby measuring the abundance of the one or more of the N targets, wherein optionally, measuring the amount of the one or more of the N barcodes is by one or more of sequencing and quantitative PCR, wherein optionally, the encoding polynucleotide in each complex comprises an overlap oligonucleotide, and wherein each complementary polynucleotide has a universal complementary overlap, wherein the universal complementary overlap is complementary to the overlap oligonucleotide in each complex.

9. A method of quantifying interactions and / or abundance of N targets in a sample, wherein N is an integer, the method comprising: contacting the sample with the N complexes of claim 1, 101 KILPATRICK TOWNSEND 797562351contacting the N complexes with N complementary polynucleotides, wherein each complementary polynucleotide comprises a barcode being complementary to the unique identifier in one of the N complexes, wherein the barcode identifies the target the one of the unique identifier identifies, wherein each complementary polynucleotide has ligatable ends, wherein optionally, each complementary polynucleotide comprises a UMI, generating link nucleic acid molecules in a reaction mixture, wherein at least one of the link nucleic acid molecules comprises the sequences of M of N complementary polynucleotides and thus comprises M of the N barcodes, optionally separating the link nucleic acid molecules from the complexes, quantifying the amount of the link nucleic acid molecules comprising the M barcodes, wherein M is an integer no greater than N, wherein optionally, each complementary polynucleotide has a universal complementary overlap wherein the universal complementary overlap is complementary to the overlap oligonucleotide in each complex.

10. The method of claim 9, wherein the link nucleic acid molecules comprise one or more copies of at least one of the M barcodes, and / or wherein generating link nucleic acid molecules in situ comprises ligating the complementary polynucleotides, or extension using a DNA polymerase.

11. The method of claim 9, wherein the method further comprises quantifying one or more of the N targets, wherein for each target, quantifying the target by summing the barcode identifying the target in individual form and combination form.

12. A method of disease diagnosis or prognosis, the method comprising: quantifying interaction and / or abundance of N targets within a cell from a patient sample with the method of claim 9, diagnosing the patient as having a disease or having increased risk of developing the disease if the amount of the link nucleic acid molecules comprising the M barcodes is greater than a control, wherein the targets promote the disease, or diagnosing the patient as having a disease or having increased risk of developing the disease if 102 KILPATRICK TOWNSEND 797562351the amount of the link nucleic acid molecules comprising the M barcodes is lower than a control, wherein the targets are suppressors of the disease, wherein optionally, the control is the amount of link nucleic acid molecules comprising the M barcodes in a normal sample.

13. A method of monitoring disease progression in a patient or monitoring a patient’s response to a therapy, the method comprising: quantifying interaction and / or abundance of N targets within a cell from a patient sample with the method of claim 9 at a first time point and a second time point during a time period, optionally, the time period is when the patient receives a therapy, wherein the second time point is later than the first time point, determining that the disease has progressed if the amount of the link nucleic acid molecules comprising the M barcodes at the second time point is different from the amount of the link nucleic acid molecules at the first time point, and / or determining that the patient is responsive to the therapy if the amount of the link nucleic acid molecules at the second time point is different from the amount of the link nucleic acid molecules at the first time point.

14. The method of claim 12, wherein the disease is a tumor, and wherein the method detects the interaction of two or more targets, optionally the N targets comprise one or more targets in FIG.

4.

15. A kit for quantifying interactions and / or abundance of two or more of N targets, wherein the kit comprises N complexes of claim 1 and N complementary polynucleotides, wherein N is an integer, wherein each complementary polynucleotide comprises a complementary identifier being complementary to the unique identifier of one complex, wherein each complementary polynucleotide has ligatable ends, wherein optionally, the polynucleotide of each complex comprises an overlap oligonucleotide, and wherein optionally, each complementary polynucleotide has a universal complementary overlap 103 KILPATRICK TOWNSEND 797562351that is complementary to the overlap oligonucleotide.

16. A method of producing a complex of the composition of claim 1 by conjugating the binding domain to the encoding polynucleotide through a flexible linker, the method comprising: providing the binding domain and the encoding polynucleotide in a molar ratio in a range from 0.5 to 6, modifying the internal nucleotide of the encoding polynucleotide with an amino functional group, treating the encoding polynucleotide with modified internal nucleotide with an SPDP–PEG–NHS ester to produce a reactive encoding polynucleotide, treating the binding domain with sulfo-SMCC to produce a reactive binding domain, contacting the reactive encoding polynucleotide with the reactive encoding binding domain, thereby forming the complex.

17. A system comprising: a processor; and a memory device including instructions executable by the processor for causing the processor to perform operations comprising: determining a distribution of barcode combinations in a sample, wherein each barcode combination of the barcode combinations includes at least one barcode of a corresponding target; determining a contribution of a target complex to the distribution of the barcode combinations; and determining, based on a conversion that uses the contribution, an abundance of the target complex in the sample, wherein optionally, the number of constituents in the target complex is in a range of 1 to 100, wherein optionally, one or more of targets are proteins, nucleic acids, lipids, small molecule compounds, and derivatives thereof, and wherein optionally the at least one barcode is a nucleic acid.

18. The system of claim 17, wherein the memory device further includes instructions executable 104 KILPATRICK TOWNSEND 797562351by the processor for causing the processor to determine the abundance of the target complex by performing operations comprising: determining a longest combination barcode of the barcode combinations, wherein the longest combination barcode corresponds to a largest target complex; determining, based on an amount of the longest combination barcode in the distribution and a probability of barcode formation, a first amount of each of multiple barcode sub-combinations in the longest combination barcode contributed by the largest target complex; and determining an abundance of the largest target complex in the sample based on the first amount of the largest combination barcode and the probability of barcode formation.

19. The system of claim 18, wherein the memory device further includes instructions executable by the processor for causing the processor to determine the abundance of the target complex by performing operations comprising: removing the first amount of each of the multiple barcode sub-combinations in the longest combination barcode contributed by the largest target complex from the distribution to generate an updated distribution; determining a next longest combination barcode in the updated distribution, wherein the next longest combination barcode corresponds to a next largest target complex; determining, based on an amount of the next longest combination barcode in the updated distribution and the probability of barcode formation, a second amount of each of the multiple barcode sub-combinations in the next longest combination barcode contributed by the next largest target complex; and determining an abundance of the next largest target complex in the sample based on the second amount of the next largest combination barcode and the probability of barcode formation.

20. The system of claim 19, wherein the operations further comprise: iteratively determining abundances of target complexes until a remaining longest combination barcode includes only a single barcode.

21. The system of claim 18, wherein the probability of barcode formation corresponds to a likelihood of a single linkage forming in a link barcode molecule, wherein optionally the link 105 KILPATRICK TOWNSEND 797562351barcode molecule incorporates a unique molecular identifier (UMI).

22. The system of claim 18, wherein the operations further comprise: determining the abundance of the largest target complex using a combinatorics equation, and / or determining an interaction map of constituents in target complexes based on the abundance of the target complex.

23. The system of claim 17, wherein an amount of a link barcode molecule that is generated is a function of a barcode combination length, an abundance and a length of an originating target complex, and a probability of barcode formation.

24. A method of disease diagnosis or prognosis, the method comprising: determining an abundance of a target complex in a sample with the system of claim 17; determining that the abundance of the target complex is different than a control greater than a threshold difference; and diagnosing a patient as having a disease or having an increased risk of developing the disease based on the abundance of the target complex being different than the control greater than the threshold difference, wherein optionally, the control is an amount of the target complex in a normal sample.

25. One or more non-transitory computer-readable storage media storing instructions that, upon execution executable by one or more processors of a system, cause the system to perform operations comprising: determining a distribution of barcode combinations in a sample, wherein each barcode combination of the barcode combinations includes at least one barcode of a corresponding target; determining a contribution of a target complex to the distribution of the barcode combinations; and determining, based on a conversion that uses the contribution, an abundance of the target complex in the sample, wherein optionally the number of constituents in the target complex is in a range of 1 to 100, 106 KILPATRICK TOWNSEND 797562351and wherein optionally one or more of targets are proteins, nucleic acids, lipids, small molecule compounds, and derivatives thereof.

26. The one or more non-transitory computer-readable storage media of claim 25, wherein the operations further comprise determining the abundance of the target complex by: determining a longest combination barcode of the barcode combinations, wherein the longest combination barcode corresponds to a largest target complex; determining, based on an amount of the longest combination barcode in the distribution and a probability of barcode formation, a first amount of each of multiple barcode sub-combinations in the longest combination barcode contributed by the largest target complex; and determining an abundance of the largest target complex in the sample based on the first amount of the largest combination barcode and the probability of barcode formation.

27. The one or more non-transitory computer-readable storage media of claim 26, wherein the operations further comprise determining the abundance of the target complex by: removing the first amount of each of the multiple barcode sub-combinations in the longest combination barcode contributed by the largest target complex from the distribution to generate an updated distribution; determining a next longest combination barcode in the updated distribution, wherein the next longest combination barcode corresponds to a next largest target complex; determining, based on an amount of the next longest combination barcode in the updated distribution and the probability of barcode formation, a second amount of each of the multiple barcode sub-combinations in the next longest combination barcode contributed by the next largest target complex; and determining an abundance of the next largest target complex in the sample based on the second amount of the next largest combination barcode and the probability of barcode formation.

28. The one or more non-transitory computer-readable storage media of claim 27, wherein the operations further comprise: iteratively determining abundances of target complexes until a remaining longest combination 107 KILPATRICK TOWNSEND 797562351barcode includes only a single barcode.

29. The one or more non-transitory computer-readable storage media of claim 26, wherein the probability of barcode formation corresponds to a likelihood of a single linkage forming in a link barcode molecule, wherein optionally, the link barcode molecule incorporates a unique molecular identifier (UMI).

30. The one or more non-transitory computer-readable storage media of claim 26, wherein the operations further comprise: determining the abundance of the largest target complex using a combinatorics equation.

31. The one or more non-transitory computer-readable storage media of claim 25, wherein the operations further comprise: determining an interaction map of constituents in target complexes based on the abundance of the target complex.

32. The one or more non-transitory computer-readable storage media of claim 25, wherein an amount of a link barcode molecule that is generated is a function of a barcode combination length, an abundance and a length of an originating target complex, and a probability of barcode formation.

33. The one or more non-transitory computer-readable storage media of claim 25, wherein the at least one barcode is a nucleic acid. 108 KILPATRICK TOWNSEND 797562351

Citation Information

Patent Citations

  • Spatial detection of biomolecule interactions

    US20230416809A1

  • Multiplex analysis of single cell constituents

    WO2017075265A1

  • Chromatin profiling compositions and methods

    WO2024112948A1