Determination of single cell and spatial protein information by recoding amino acid polymers into DNA polymers

By employing PUMA tags and TFS to reverse translate protein sequences into DNA polymers, the method addresses the limitations of current proteomic analysis, providing high-throughput and spatial proteomic capabilities for sensitive and accurate protein characterization, leading to early disease detection and therapeutic insights.

WO2025155747A1PCT designated stage expired Publication Date: 2025-07-24ABRUS BIO INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/011915
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-17
Filing Date
2025-01-16
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Current tools and technologies are inadequate for sensitive, accurate, and economical characterization of single-cell and spatial proteomes, lacking the ability to de novo discover biomarkers and provide high-throughput, unbiased analysis of protein sequences and concentrations.

Method used

The use of peptide unique molecular association (PUMA) tags and trifunctional supports (TFS) to reverse translate protein sequences into DNA polymers, preserving metadata for cell and spatial origin, enabling high-throughput and spatial proteomic analysis.

Benefits of technology

Enables accurate determination of protein sequences and concentrations with spatial context, facilitating the discovery of novel biomarkers and improving healthcare by allowing early detection of diseases like cancer and enhancing therapeutic discovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025011915_24072025_PF_FP_ABST
    Figure US2025011915_24072025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to compositions of matter, methods, and systems for analyzing polymeric macromolecules, including polymeric macromolecules such as peptides, polypeptides, and proteins from single cells and from spatially-orientated cells.
Need to check novelty before this filing date? Find Prior Art

Description

DETERMINATION OF SINGLE CELL AND SPATIAL PROTEIN INFORMATION BY RECODING AMINO ACID POLYMERS INTO DNA POLYMERSCROSS-REFERENCE

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 622,004, filed January 17, 2024, which application is incorporated herein by reference.INCORPORATION BY REFERENCE OF SEQUENCE LISTING

[0002] The preset application is being filed along with a Sequence Listing in electron format. The Sequence Listing is provided as a file entitled 062954-503001 WO.xml, created January 15, 2025, which is 314,501 bytes in size. The information in the electronic format of the Sequence Listing is incorporated by reference in its entirety.FIELD

[0003] The present disclosure relates to compositions of matter, methods, and systems for analyzing macromolecules. Some aspects include analyzing macromolecules from single cells or spatially- orientated cells.BACKGROUND

[0004] Proteins are fundamental to cellular function. Accordingly, the sequences of the thousands of proteins within each cell, as well as their concentrations, are useful indicators of cell health. Aberrant sequences or concentrations of proteins may signal a disease state. However, tools and technologies are currently lacking for sensitive, accurate, economical, and unbiased characterization of proteomes, especially the proteomes of single cells.

[0005] Spatial biology adds another dimension to single-cell analysis. Spatial context influences cellular functions, developmental processes, and disease mechanisms. Spatial biology is starting to unravel the complexity of the tumor microenvironment, and being used to evaluate disease progression, and predicting immune- therapeutic efficacy and mechanism. However, current immunofluorescent technologies are incapable of de novo discovery of biomarkers. For these and other reasons, better tools to evaluate protein and peptide sequence and concentration should be developed.

[0006] There is thus a need in the art for compositions of matter, methods, and systems for highly- parallelized, accurate, sensitive, and high-throughput single-cell and spatial proteomic analysis. The present disclosure addresses this and other needs.SUMMARY

[0007] Disclosed herein are compositions of matter, methods, and systems for analyzing polymeric macromolecules, including peptides, polypeptides, and proteins. These aspects may be useful in the context of single cells and spatial orientations and in a highly-parallel and high-throughput manner via recoding the protein sequences into DNA polymers while preserving and assembling metadata information, such as cell of origin, and spatial orientation of origin.

[0008] Disclosed herein, in some embodiments, area single cell reverse translation methods, comprising (a) partitioning cells of a sample; (b) contacting the partitioned cells with reagents that permeabilize cell membranes of the cells, and release proteins from the cells; (c) providing a trifunctional support (TFS) to the cells, the TFS comprising: (x) a solid support surface, (y) an immobilized peptide unique molecular association (PUMA) tag, and (z) a reactive moiety for binding an amino acid residue of the peptide; (d) coupling a protein of the cells to the TFS; (e) reverse translating the PUMA tag and protein coupled to the TFS to generate a memory oligonucleotide comprising a sequence corresponding to identity and positional information of amino acids of the protein and PUMA tag; (f) obtaining sequence information for the memory oligonucleotide; and (g) based on the obtained sequence information, determining the identity and positional information of the amino acid residue of the protein and assigning metadata to the sequence information that includes information of a cell of origin of the protein.

[0009] Disclosed herein, in some embodiments, are spatial assay methods, comprising; (a) providing a plurality of amino acid identifier conjugates (PUMA-conjugates) each comprising: (x) a peptide, and (y) a reactive moiety capable of coupling to a protein, and optionally (z) an orthogonal reactive moiety capable of joining the PUMA-tagged protein conjugate to a solid support; (b) depositing the plurality of PUMA-conjugates on the solid support; (c) depositing a sample of cells on a surface of the solid support; (d) capturing proteins of the cells with the reactive moiety of the PUMA-conjugates, thereby joining the proteins of the cells with the PUMA-conjugates and generating PUMA-tagged protein conjugates; and (e) reverse translating the PUMA-tagged protein conjugates to obtain a memory oligonucleotide comprising sequences corresponding to identity and positional information of amino acids of the proteins, and corresponding with the PUMA tag;(f) obtaining sequence information for the memory oligonucleotide; and (g) based on the obtained sequence information, determining the identity and positional information of the amino acids of the proteins and assigning metadata that includes its original spatial location within the sample of cells.

[0010] Described herein, in some embodiments, is a peptide unique molecular association (PUMA)- conjugate, comprising: (a) a PUMA tag; (b) a reactive moiety for coupling to an amino acid; and (c) optionally, an orthogonal reactive moiety that binds a solid support. In some embodiments, use of the PUMA tag provides metadata relating: a molecule, spatial location, relative spatial location, condition, cell, or sample of origin, for molecule(s) that associate with the PUMA tag. In some embodiments, thePUMA tag comprises repeats of a sequence of amino acids. In some embodiments, the PUMA tag comprises a constant portion and a variable portion. In some embodiments, the PUMA tag may comprise 1, 2, or more portions each with the same or a different function. In some embodiments, the PUMA tag comprises a natural amino acid. In some embodiments, the PUMA tag comprises a nonnatural amino acid.

[0011] Described herein, in some embodiments, is a trifunctional support (TFS), comprising: (x) a solid support surface; (y) immobilized peptide unique molecular association tags (PUMA tags); and (z) reactive moieties associated with the PUMA tags, wherein the reactive moieties each bind an amino acid. In some embodiments, the solid support comprises a hydrogel, a functionalization molecule, or a passivation molecule. In some embodiments, the solid support comprises a bead, in some embodiments, the bead comprises a bead type corresponding with its PUMA tag sequence. In some embodiments, the TFS further comprises a moiety that facilitates solubility, a moiety that reduces non-specific interactions, a sequence or moiety that facilitates removal of a PUMA-tagged protein from the solid support surface, such as a cleavable linker, or a combination thereof. In some embodiments, the pool of TFS beads comprises a first TFS of a first bead type, and a second TFS of a second bead type. In some embodiments, the amounts of the first and second TFS are mixed in about equal proportions. In some embodiments, the bead types of the first and second TFS have PUMA tag sequences that differ from one another. In some embodiments, the TFS may be used in a single-cell assay, a spatial assay, or both. In some embodiments, the TFS may be used to capture proteins from a tissue section for spatial analysis, wherein the TFS lacks a functional group that supports reverse translation and provides a higher density of capture site or larger size appropriate for protein capture. A method herein may include any aspect or combination of aspects in this section.BRIEF DESCRIPTION OF THE DRAWINGS

[0012] So that the manner in which the above recited features of the present disclosure can be understood in detail, a more particular description of the disclosure, briefly summarized above, may be had by reference to embodiments, some of which are illustrated in the appended drawings. It is to be noted, however, that the appended drawings illustrate only exemplary embodiments and are therefore not to be considered limiting of its scope and may admit to other equally effective embodiments.

[0013] Accordingly, the foregoing and other features and advantages of the present disclosure will be more fully understood from the following detailed description of illustrative embodiments taken in conjunction with the accompanying drawings in which:

[0014] FIG. 1 schematically illustrates a simplified block diagram of an exemplary workflow for analyzing polymeric macromolecules, including polymeric macromolecules such as peptides, and proteins, from single cells according to embodiments of the present disclosure.

[0015] FIG. 2 schematically illustrates various operations of process 1200, according to the workflow of FIG. 1 and embodiments of the present disclosure, for single cell analysis of a sample of cells. Operations of a reverse translation process to create a memory oligo convolve PUMA and analyte information, requiring in silico deconvolution.

[0016] FIG. 3 schematically illustrates various operations of process 1300 according to the workflow of FIG. 1 and the embodiments of the present disclosure. Within this alternate embodiment the transfer of PUMA-tagged protein conjugates to a surface that supports reverse translation is eliminated

[0017] FIG. 4 schematically illustrates components of a TFS bead assembled with: 1) a protein analyte,2) an isothiocyanate-conjugate, and 3) a Peptide Unique Molecular Association (PUMA) tag.

[0018] FIG. 5 schematically illustrates a simplified block diagram of an exemplary workflow for spatially mapping polymeric macromolecules, including polymeric macromolecules such as peptides, and proteins, of a spatially-oriented plurality of cells according to embodiments of the present disclosure.

[0019] FIG. 6A-6C schematically illustrate various operations of a process 2100, according to the workflow of FIG. 5 and embodiments of the present disclosure, for spatial mapping proteins of a microsectioned tissue sample. Note that TFS beads may be ordered or randomly arrayed.

[0020] FIG. 7A-7B shows PCR data of an assembled recode block comprising an amino acid with a non-natural R-group. Signal is clearly discernable over background.

[0021] FIG. 8 shows an electrophoresis gel image that indicates formation, via reverse translation methods, of recode blocks comprising amino acids having non-natural R-groups and a memory oligo comprising said recode blocks.

[0022] FIG. 9 provides a block diagram of an automated system for production of spatial information.

[0023] FIG. 10 illustrates a simplified block diagram of an exemplary workflow for analyzing polymeric macromolecules, including polymeric macromolecules such as peptides, and proteins, according to embodiments of the present disclosure.

[0024] Fig 11 illustrates a scheme for synthesis of an example PUMA referred to as PUMA-1.

[0025] Fig 12 shows the mass spectrum of the Pep20-24507 intermediate (+ mode).

[0026] Fig. 13 shows the mass spectrum of PUMA-1 (positive mode).

[0027] FIG. 14 schematically illustrates joining oligonucleotides associated with a CRC and a non- covalent biomolecular recognition molecule through their interaction with elements of an immobilized biomolecule.

[0028] FIG. 15 shows q-PCR data of the product created from the configuration depicted in FIG. 14.

[0029] It should be understood that the drawings are not necessarily to scale, and that like reference numbers refer to like features. It is contemplated that elements and features of one embodiment may be beneficially incorporated in other embodiments without further recitation.DETAILED DESCRIPTION

[0030] Methods and compositions described herein provide assay systems: 1) for efficient highly- multiplexed assay of proteins within single cells. These may be useful for determining the proteome within a cell, and the proteomes across cells of homogeneous and heterogeneous cell samples, tissues, and experimental conditions, and 2) for efficient highly-multiplexed assay of proteins within spatially- oriented cells. These methods allow correlation of protein information including: sequence or partial sequence, isoform, post-translational modification status, and abundance, with spatial information of spatially- oriented cells to provide high-depth high resolution maps of proteins in tissues.

[0031] Once hypothesis-free tools are accessible, the discovery in single cell and spatial context of novel biomarkers, accurate determination of concentrations for even the lowest-abundance proteins, discovery of important post-translational modifications, and monitoring of the dynamics of the proteome in response to their environment and in influence of their environment, are some of the first steps toward improving healthcare. These tools will create deeper understanding and earlier detection of important signatures of cancer and other health conditions will allow diagnosis at the earliest stages, facilitate therapeutic discovery, and create beneficial impact on patient care by informing the course of treatment.

[0032] Spatial multi-omics includes the study of gene expression and protein abundance with spatial context to elucidate functional biology. The present disclosure is compatible with spatial transcriptomic technologies and provides the capability to determine spatial locations of protein molecules within histological tissue sections, thereby expanding the repertoire of spatial multi-omic tools. Spatial multi- omics can facilitate an improved understanding of tissue and cellular microenvironments, and thereby improve our understanding of disease etiology, progression, treatment. Introduction of efficient, high- depth, high-throughput, hypothesis-free, spatial proteomics will accelerate the rate of scientific discovery. The present disclosure provides the capability to discover novel proteins, and protein expression patterns across cell populations with cellular and spatial resolution, all while enjoying the cost and throughput advantages of NGS sequencing.

[0033] The disclosed PUMA-tagging system can be seamlessly integrated into multi-omics workflows, combining spatial proteomic data with transcriptomic and epigenomic datasets. By assigning metadata to proteins based on cell origin and spatial context, researchers can correlate these findings with gene expression profiles obtained through single-cell RNA sequencing or spatial transcriptomics. This integration would provide a holistic view of cellular function, uncovering intricate relationships between protein activity, gene regulation, and epigenetic modifications. For instance, combining spatially resolved proteomic data with epigenetic marks could elucidate the role of histone modifications in regulating local protein expression during tumor progression or immune response.

[0034] The disclosure includes a novel composition of matter that we call a Peptide Unique Molecular Association tag (PUMA tag), and a secondary novel composition called a Tri-Functional Surface (TFS) that comprises a PUMA conjugate.

[0035] All of the functionalities described in connection with one embodiment are intended to be applicable to the additional embodiments described herein except where expressly stated or where the feature or function is incompatible with the additional embodiments. For example, where a given feature or function is expressly described in connection with one embodiment but not expressly mentioned in connection with an alternative embodiment, it should be understood that the feature or function may be deployed, utilized, or implemented in connection with the alternative embodiment unless the feature or function is incompatible with the alternative embodiment.

[0036] The practice of the techniques described herein may employ, unless otherwise indicated, conventional techniques and descriptions of organic chemistry, polymer technology, molecular biology (including recombinant techniques), cell biology, proteomics, biochemistry and sequencing technology, which are within the skill of those who practice in the art. Such conventional techniques include polymer array synthesis, hybridization and ligation of polynucleotides and other polymers, and detection of hybridization using a label. Specific illustrations of suitable techniques can be had by reference to the examples herein. However, other equivalent conventional procedures can, of course, also be used. Such conventional techniques and descriptions can be found in standard laboratory manuals such as Green et al., Eds. (1999), Genome Analysis: A Laboratory Manual Series (Vols. I-IV); Weiner, Gabriel, Stephens, Eds. (2007), Genetic Variation: A Laboratory Manual; Dieffenbach, Dveksler, Eds. (2003), PCR Primer: A Laboratory Manual; Mount (2004), Bioinformatics: Sequence and Genome Analysis; Sambrook and Russell (2006), Condensed Protocols from Molecular Cloning: A Laboratory Manual; and Sambrook and Russell (2002), Molecular Cloning: A Laboratory Manual (all from Cold Spring Harbor Laboratory Press); Stryer, L. (1995) Biochemistry (4th Ed.) W.H. Freeman, New York N.Y.; Gait, “Oligonucleotide Synthesis: A Practical Approach” 1984, IRL Press, London; Nelson and Cox (2000), Lehninger, Principles of Biochemistry 3rd Ed., W. H. Freeman Pub., New York, N.Y.; Berg et al. (2002) Biochemistry, 5th Ed., W.H. Freeman Pub., New York, N.Y.; all of which are herein incorporated in their entirety by reference for all purposes.

[0037] In the following description, numerous specific details are set forth to provide a more thorough understanding of the present disclosure. However, it will be apparent to one of ordinary skill in the art that the present disclosure may be practiced without one or more of these specific details. In other instances, well-known features and procedures well known to those skilled in the art have not been described in order to avoid obscuring the disclosure.

[0038] Methods herein may include aspects of WO2024015875 or W02024040236, which are incorporated herein in their entirety. For example, some methods herein may include aspects of a chemically reactive conjugate, or of a protein recoding method or a reverse translation method of WO2024015875 or WG2024040236.Single Cell Proteomics

[0039] Single-cell studies have historically focused on nucleic acids. [Mincarelli, et al, Proteomics, 2018, 18, 1700312] However, the genome and transcriptome only provide partial information. Proteins, which do most of the actual work in cells, are not always correlated with transcription levels due to several factors. [Franks A, Airoldi E, Slavov N (2017) PLoS Comput Biol 13(5): el005535] In addition, post-translational modifications are masked in single-cell nucleic acid assays. An effective single-cell proteomics assay will allow exploration of the proteome at the individual cell level to gain insights into cellular heterogeneity and functional differences. Single-cell proteomics has the potential to reveal how individual cells differ in protein expression, post-translational modifications, and signaling pathways. Applications include inter-cellular heterogeneity crucial for understanding developmental biology; cancer research; neurodegenerative diseases research; immune system profiling; autoimmune disease research; drug discovery and toxicology studies, and more.

[0040] In addition to experimental applications, advanced computational approaches, including machine learning models, may be applied to analyze PUMA-derived proteomic signatures. Supervised learning algorithms can be trained on labeled datasets to identify proteomic patterns predictive of disease states, treatment responses, or patient outcomes. Furthermore, techniques such as t-SNE or UMAP may be used to uncover cellular heterogeneity within tissue samples. Expanding this approach to incorporate multi-omics data, such as transcriptomics or metabolomics, may be useful in generating predictive models with unparalleled depth, and may be useful for informing precision medicine strategies.

[0041] A generalized method of the present disclosure is presented in FIG.l. Novel concepts within the method include the disclosure of a peptide unique molecular association tag, (PUMA tag), PUMA code space, PUMA-conjugates, and PUMA surfaces. Each of these are explained in the Descriptions, Figures, and Examples herein.

[0042] Disclosed herein, in some embodiments, are single-cell reverse translation methods, comprising: (a) partitioning cells of a sample; (b) contacting the partitioned cells with reagents that permeabilize cell membranes of the cells, and release proteins from the cells; (c) providing a trifunctional support (TFS) to the cells, the TFS comprising: (x) a solid support surface, (y) an immobilized peptide unique molecular association (PUMA) tag, and (z) a reactive moiety for binding an amino acid residue of the peptide; (d) coupling a protein of the cells to the TFS; (e) reverse translating the PUMA tag and protein coupled to the TFS to generate a memory oligonucleotide comprising a sequence corresponding to identity and positional information of amino acids of the protein and PUMA tag; (f) obtaining sequence information for the memory oligonucleotide; and (g) based on the obtained sequence information, determining the identity and positional information of the amino acid residue of the protein and assigning metadata to the sequence information that includes information of a cell of origin of the protein.

[0043] Disclosed herein, in some embodiments, are single-cell reverse translation methods. A method may include partitioning cells. The cells may be cells of a sample. A method may include contacting the partitioned cells with a reagent or reagents. The reagent or reagents may permeabilize cell membranes of the cells. The reagent or reagents may release proteins from the cells. A method may include providing a multifunctional support such as a trifunctional support (TFS) to the cells. The multifunctional support may include a solid support surface. The multifunctional support may include a tag such as an immobilized peptide unique molecular association (PUMA) tag. The multifunctional support may include a reactive moiety. The reactive moiety may be used for binding an amino acid residue of the peptide. A method may include coupling a protein of the cells to the multifunctional support. A method may include reverse translating the tag (e.g. PUMA tag). A method may include reverse translating the protein coupled to the multifunctional support. Reverse translating may generate an oligonucleotide such as a memory oligonucleotide. The oligonucleotide or memory oligonucleotide may include a sequence corresponding to identity of amino acids of the protein and tag (e.g. PUMA tag). The oligonucleotide or memory oligonucleotide may include a sequence corresponding to positional information of amino acids of the protein and tag (e.g. PUMA tag). The oligonucleotide or memory oligonucleotide may include a sequence corresponding to identity or positional information of amino acids of the protein. The oligonucleotide or memory oligonucleotide may include a sequence corresponding to identity or positional information of amino acids of the tag (e.g. PUMA tag). The oligonucleotide or memory oligonucleotide may include a sequence corresponding to identity and positional information of amino acids of the protein and tag (e.g. PUMA tag). A method may include obtaining sequence information for the oligonucleotide (e.g. memory oligonucleotide). A method may include determining identity information of an amino acid residue of the protein (e.g. the amino acid residue bound by the reactive moiety). A method may include determining positional information of an amino acid residue of the protein(e.g. the amino acid residue bound by the reactive moiety). A method may include determining the identity and positional information of the amino acid residue of the protein. Determining identity or positional information an amino acid may be based on sequence information (e.g. the obtained sequence information). A method may include assigning metadata to the sequence information. Metadata may include information of a cell of origin of the protein.Process 1000 - Single Cell

[0044] Disclosed herein, in some embodiments, are methods for associating the metadata that includes cell of origin together with identity and positional information of an amino acid residue of a protein coupled to a solid support, the method comprising: (a) partitioning cells of a sample; (b) contacting the partitioned cells with reagents that permeabilize the cell membrane; (c) contacting the contents of the cell with processing reagents; (d) providing a trifunctional support (“TFS”) comprising (x) a solid support surface, (y) an immobilized peptide unique molecular association (PUMA) tag, and (z) a reactive moiety for binding an amino acid residue of the peptide; (e) coupling a protein to the TFS; (f)reverse translating the immobilized protein information; (g) obtaining sequence information for the memory oligo; and (h) based on the obtained sequence information, determining identity and positional information of an amino acid residue of the peptide and assigning metadata that includes information of the cell of origin.

[0045] Disclosed herein, in some embodiments, are methods. A method may include associating metadata together with identity information of an amino acid residue of a protein. A method may include associating metadata together with positional information of an amino acid residue of a protein. A method may include associating metadata together with identity and positional information of an amino acid residue of a protein. The protein may be coupled to a solid support. The metadata may include cell of origin information. The method may include partitioning cells. The cells may be cells of a sample. The method may include contacting the partitioned cells with a reagent or reagents. The reagent or reagents may permeabilize cell membranes of the cells. The method may include contacting contents of the cells with processing reagents. The reagent or reagents may permeabilize a cell membrane of a cell of the cells. The method may include contacting contents of the cell with processing reagents. The method may include providing a multifunctional support. The multifunctional support may include a solid support surface. The multifunctional support may include an immobilized peptide unique molecular association (PUMA) tag. The multifunctional support may include a reactive moiety for binding an amino acid residue. The amino acid residue may be an amino acid residue of the peptide, such as a terminal (e.g. N-terminal) amino acid residue of the peptide. The method may include coupling a protein to the multifunctional support. The method may include providing a trifunctional support (“TFS”). The multifunctional support may include a TFS. The TFS may include a solid support surface. The TFS may include an immobilized peptide unique molecular association (PUMA) tag. The TFS may include a reactive moiety for binding an amino acid residue. The amino acid residue may be an amino acid residue of the peptide, such as a terminal (e.g. N-terminal) amino acid residue of the peptide. The method may include coupling a protein to the TFS. The method may include reverse translating protein information or immobilized protein information. The method may include reverse translating the immobilized protein information. Reverse translating may produce an oligonucleotide. The oligonucleotide may be or include a memory oligonucleotide. The reverse translating may produce a memory oligonucleotide. The memory oligonucleotide may include a sequence corresponding to identity or positional information of amino acids of the protein and PUMA tag. The method may include obtaining sequence information for a memory oligo. The method may include determining identity information of an amino acid residue of the peptide (e.g. based on the obtained sequence information). The method may include determining positional information of an amino acid residue of the peptide (e.g. based on the obtained sequence information). The method may include determining identity and positional information of an amino acid residue of the peptide (e.g. based on the obtained sequence information). The method may include assigning metadata. The assigned metadata may include information of a cell or the cell of origin.Process 1100 - Single Cell, Plurality of Proteins

[0046] Disclosed herein, in some embodiments, are methods for associating the metadata that includes cell of origin with a plurality of proteins derived from a single cell, the method comprising: (a) partitioning cells of a sample; (b) contacting the partitioned cells with reagents that permeabilize the cell membrane; (c) contacting the contents of the partitioned cells with processing reagents; (d) providing a trifunctional support (“TFS”) comprising (x) a solid support surface, (y) an immobilized peptide unique molecular association (PUMA) tag, and (z) a reactive moiety for binding an amino acid residue of the peptide ; (e) coupling proteins of one cell to one TFS (f) reverse translating the immobilized protein information; (g) obtaining sequence DNA information for the memory oligos; and (h) based on the obtained DNA sequence information, assigning metadata that includes information of the cell of origin to the plurality of protein sequences obtained.

[0047] Disclosed herein, in some embodiments, are methods for associating the metadata that includes cell of origin with a plurality of proteins derived from a single cell. The method may include partitioning cells of a sample. The method may include contacting the partitioned cells with reagents that permeabilize the cell membrane. The method may include contacting the contents of the partitioned cells with processing reagents. The method may include providing a multifunctional support such as a trifunctional support (“TFS”) comprising (x) a solid support surface, (y) an immobilized peptide unique molecular association (PUMA) tag, and (z) a reactive moiety for binding an amino acid residue of the peptide. The method may include coupling proteins of one cell to one TFS. The method may include reverse translating the immobilized protein information. The method may include obtaining sequence DNA information for the memory oligos. The method may include: based on the obtained DNA sequence information, assigning metadata that includes information of the cell of origin to the plurality of protein sequences obtained.Process 1200 - Single Cell, Plurality of Proteins Across a Plurality of Cells

[0048] Disclosed herein, in some embodiments, are methods for associating the metadata that includes cell of origin with a plurality of proteins derived from each of a plurality of cells, the method comprising: (a) partitioning cells of a sample; (b) contacting the partitioned cells with reagents that permeabilize the cell membrane; (c) contacting the contents of the each partitioned cell with processing reagents; (d) providing a trifunctional support (“TFS”) comprising (x) a solid support surface, (y) an immobilized peptide unique molecular association (PUMA) tag, and (z) a reactive moiety for binding an amino acid residue of the peptide to each partition; (e) coupling proteins of each cell to one TFS across a plurality of partitions and pooling the contents of each partition; (f) reverse translating the immobilized protein information; (g) obtaining DNA sequence information for the memory oligos; and(h) based on the obtained DNA sequence information, assigning metadata that includes information of the cell of origin to the plurality of proteins across the plurality of cells.

[0049] Disclosed herein, in some embodiments, are methods for associating the metadata that includes cell of origin with a plurality of proteins derived from each of a plurality of cells, the method comprising. The method may include partitioning cells of a sample. The method may include contacting the partitioned cells with reagents that permeabilize the cell membrane. The method may include contacting the contents of the each partitioned cell with processing reagents. The method may include providing a multifunctional support such as a trifunctional support (“TFS”) comprising (x) a solid support surface, (y) an immobilized peptide unique molecular association (PUMA) tag, and (z) a reactive moiety for binding an amino acid residue of the peptide to each partition. The method may include coupling proteins of each cell to one TFS across a plurality of partitions and pooling the contents of each partition. The method may include reverse translating the immobilized protein information. The method may include obtaining DNA sequence information for the memory oligos. The method may include based on the obtained DNA sequence information, assigning metadata that includes information of the cell of origin to the plurality of proteins across the plurality of cells.Process 1300 - Single Cell, Plurality of Proteins, Cells Kept Separate Through Reverse Translation

[0050] Disclosed herein, in some embodiments, are methods for associating the metadata that includes cell of origin with a plurality of proteins derived from a single cell, the method comprising: (a) partitioning cells of a sample; (b) contacting the partitioned cells with reagents that permeabilize the cell membrane; (c) contacting the contents of the partitioned cells with processing reagents; (d) providing a solid support for immobilization of analyte proteins; (e) coupling the proteins of one cell to the support (f) reverse translating the immobilized protein information; (g) indexing the memory oligos of a partition (e.g. using commercially-available sample indexing kits, or techniques known to those skilled in the art), then pooling the indexed memory oligos and obtaining DNA sequence information for the pool of memory oligos; and (h) based on the obtained DNA sequence information, assigning metadata that includes information of the cell of origin to the plurality of protein sequences obtained.

[0051] Disclosed herein, in some embodiments, are methods for associating the metadata that includes cell of origin with a plurality of proteins derived from a single cell. The method may include partitioning cells of a sample. The method may include contacting the partitioned cells with reagents that permeabilize the cell membrane. The method may include contacting the contents of the partitioned cells with processing reagents. The method may include providing a solid support for immobilization of analyte proteins. The method may include coupling the proteins of one cell to the support. The method may include reverse translating the immobilized protein information. The method may include indexing the memory oligos of a partition. The method may include indexing the memory oligos of a partition using commercially-available sample indexing kits, or techniques known to those skilled in the art, then pooling the indexed memory oligos. The method may include indexing obtaining DNAsequence information for the pool of memory oligos. The method may include based on the obtained DNA sequence information, assigning metadata that includes information of the cell of origin to the plurality of protein sequences obtained.Trifunctional Supports

[0052] Disclosed herein, in some embodiments, is a trifunctional support (TFS). A TFS may include a solid support surface. A TFS may include an immobilized peptide unique molecular association tag (PUMA tag). A TFS may include immobilized peptide unique molecular association tags (PUMA tags). A TFS may include reactive a moiety associated with a PUMA tag. A TFS may include reactive moieties associated with PUMA tags. A reactive moiety may bind an amino acid. Reactive moieties may bind amino acids. A TFS may include (x) a solid support surface; (y) immobilized peptide unique molecular association tags (PUMA tags); and (z) reactive moieties associated with the PUMA tags, wherein the reactive moieties each bind an amino acid.

[0053] In some embodiments, the solid support comprises a hydrogel, a functionalization molecule, or a passivation molecule. In some embodiments, the solid support comprises a bead. In some embodiments, the bead comprises a bead type corresponding with its PUMA tag sequence.

[0054] A TFS may include a moiety that facilitates solubility. A TFS may include a moiety that reduces non-specific interactions. A TFS may include a sequence or moiety that facilitates removal of a PUMA- tagged protein from the solid support surface, such as a cleavable linker. A TFS may include a moiety that facilitates solubility, a moiety that reduces non-specific interactions, a sequence or moiety that facilitates removal of a PUMA-tagged protein from the solid support surface, such as a cleavable linker, or a combination thereof.

[0055] Some embodiments relate to or include a pool of TFS beads comprising a first TFS of a first bead type, and a second TFS of a second bead type. In some embodiments, amounts of the first and second TFS are mixed in about equal proportions In some embodiments, the bead types of the first and second TFS have PUMA tag sequences that differ from one another.

[0056] Some embodiments relate to or include use of a TFS. Use of a TFS may be in a single-cell assay. Use of a TFS may be in a spatial assay. Some embodiments relate to or include use of the TFS of any one of the previous claims, in a single-cell assay, a spatial assay, or both.

[0057] Some embodiments relate to or include use of a TFS to capture proteins from a tissue section for spatial analysis. In some embodiments, the TFS lacks a functional group that supports reverse translation, or provides a higher density of capture site or larger size appropriate for protein capture. Some embodiments relate to or include use of a TFS to capture proteins from a tissue section for spatial analysis, wherein the TFS lacks a functional group that supports reverse translation, and provides a higher density of capture site or larger size appropriate for protein capture.Single Cell Aspects

[0058] In some embodiments the solid support, or TFS, comprises chemical moieties, which may support reverse translation as described herein, such as: amine, carboxyl, NHS, aldehyde, azide, alkyne, maleimide, thiol, hydrazine, oxyamine, tetrazine, an alkene, acryl, trans-cyclooctene (TCO), a DBCO, a bicyclononyne, a norbornene, a strained alkyne, strained alkene, and / or a derivative thereof.

[0059] In some aspects, cells are partitioned via mechanical methods. For example, by using limiting dilution into a microwell plate or device, or by using methods similar to those employed in the Qiagen QIAcuity dPCR system.

[0060] In some aspects, a sample is a biological sample. The biological sample may comprise a bacterial cell, a mammalian cell, a whole blood sample, a tumor sample, a tissue sample, or any engineered cell or cultured cell sample.

[0061] In some aspects the nuclei preps, exosome preps or other subcellular preps may be used to enrich for analysis of proteins associated with corresponding subcellular entities.

[0062] In some aspects, a sample is processed with processing reagents or purification methods, such as filtration, normalization, or sorting prior to permeation.

[0063] In some aspects, cells are partitioned via commercially available cell sorting instrumentation. For example, partitioned using automation such as a Cellumation cv.SINGULATE instrument, Namocell Benchtop Cell Sorter, Molecular Devices DispenCell, or Cytena F.Sight Omics, which can isolate 384 cells into a microplate in under 8 minutes.

[0064] In some aspects sorting may be employed to enrich or deplete cells based on parameters evident from optical microscopy, cytometric evaluation, a protein markers identification study, or other analytical technique.

[0065] In some aspects, cells are partitioned via single-emulsion or double-emulsion partition methods, [ Brower, et. al., Anal Chem. 2020 Oct 6;92(19):13262-13270], or otherwise entrapped in liposomes, polymeric micelles or nanoparticles. And in some aspects, cells are partitioned via advanced picoliter droplet technologies, such as those provided by 10X Genomics and BioRad. For example, reagent delivery and partitioning may be accomplished using Gel Bead in Emulsion GEM technology.

[0066] In some aspects the permeabilizing reagents are chosen from a list of permeabilizing reagents and / or kits that include: organic solvents, such as methanol and acetone, detergents, such as saponin, Triton X-100, Tween-20, SDS, deoxycholate, and lytic enzymes, such as lysozyme, zymolase, salts, antibiotics, reducing agents, and reagents formulated and available commercially through BD Biosciences, ThermoFisher, Fisher Scientific, Sigma and others. [Jamur, et.al., Methods Mol Biol. (2010) 588:63-6 ; Naglak et.al., Bioprocess Technol. (1990) 9:177-205]

[0067] In some aspects, physical disruption, for example using acoustic or pressure waves, and / or heat or freeze thaw cycles and / or osmotic pressure are employed to permeabilize the cell.

[0068] In some aspects, processing reagents include: detergent(s), surfactant(s), chaotropic agent(s), reducing agent(s), alkylating agent(s), derivatizing agents, protease inhibitors, and / or enzymes usedmodify the cellular components, such as collagenase and hyaluronidase, DNase, RNase, CRISPR, and chemicals for manipulation of protein samples described previously. [Previero, A., et al. (1973) Febs Letters, 33(1), 135-138; Pesavento, J.J., et al. Mol Cell Proteomics. 2007 Sep;6(9): 1510-26; Suttapitugsakul, S., et al. Mol Biosyst. 2017;13(12):2574-2582; Rusnak, F., et al. J Biomol Tech. 2004; 15(4) :296-304; Honegger, A., et al. Biochem J. 1981 ;199(l):53-59; Jay, D.G. J Biol Chem. 1984;259(24):15572-15578; Tu, et.al., J Proteome Res. (2010), 9(10):4982-91; Friedman, et. al., J Protein Chem. (2001), 20(6):431-53]

[0069] In some aspects, a sample is processed with agents that target highly-abundant proteins for depletion from the analysis. For example, IgM antibodies are provided for agglutination of highly- abundant proteins.

[0070] In some aspects, a sample is processed with agents, such as IgG conjugates target selected proteins for enrichment in the analysis.

[0071] In some aspects, enzymes used for permeabilization and / or processing are synthetically engineered to render them spectators in reverse translation operations.

[0072] In some aspects, step (b) is optional. Non-permeabilized cells may be particularly useful when targeting cell surface receptors, membrane proteins or proteins otherwise associated with the exterior cell membrane.

[0073] In some aspects step (d) comprises a pool of TFS beads, each TFS bead type differentiated by the PUMA sequence immobilized on its surface. TFS beads randomly partitioned provide a different PUMA sequence to each of the partitions. In some aspects, on average approximately one bead occupies one partition, and the number of beads per partition is distributed according to Poisson statistics. In some aspects, beads and partitions are manipulated to achieve nominally one, two, or any value between 0 and 1000000 beads per partition using methods that do not result in a Poisson distribution.

[0074] In some aspects the PUMA tag is joined to the TFS via a linker.

[0075] In some aspects, the reactive moiety of the TFS may bind either directly or indirectly to an amino acid residue of the peptide.

[0076] In some aspects, reagents are removed and / or exchanged between any appropriate step(s) of the method disclosed herein. Methods to exchange reagents may include: iterative serial dilutions of the TFS, centrifugation, filtration, or another suitable means.

[0077] In some aspects, Poisson loading of partitions, or randomized or inaccurate processing of cells, is employed such that in step (e) 0, 1, or more cells may interact with 0, 1, or more TFS. In these cases, deconvolution of signals during analysis in step(h) provides identity and positional information of an amino acid residue of the peptide and metadata that includes information of the cell of origin.

[0078] In some aspects, the metadata that includes cell of origin provides a probabilistic estimate of the possible subset of cells from which the protein may have originated.

[0079] In some aspects, one or more operations of the method are repeated one or more times to increase a step yield of the method. For example, in specific aspects, operations (b), (c), (e) and / or (f) are repeated one or more times to increase the step yield.

[0080] In some aspects the PUMA tag is joined to protein analytes via action of an enzyme or chemical reagent. For example, sortase may be employed to add various amino acid sequences to the N-terminus of a protein. Exopeptidase may be used to prune incorporated amino acids sequences to remove amino acids superfluous to the function of the PUMA tag.

[0081] In some aspects, a protein may be joined at multiple locations with a PUMA tag. Thus, an individual protein may be labeled at multiple sites with a PUMA tag that is related to the same cell but comprises a different unique molecule identifier sequence.

[0082] In some aspects, the PUMA tag joined to the TFS is an amino acid sequence between 4 and 16 monomers long with a free amino terminus.

[0083] In some aspects the TFS comprises a bead having diameter between 0.1 and lOOuM.

[0084] In some aspects, the PUMA tag is joined to the TFS using a trifunctional molecule.

[0085] In some aspects the PUMA-conjugate to the TFS beads or to other solid support may be covalent or non-covalent. Attachment mechanisms may involve a nucleic acid, hybridization of nucleic acids, a Tn5 sequence, a poly(d)T sequence, a random hexamer sequence, protease cleavable sequence, an antibody, aptamer, a scFv, nanobody, or any suitable affinity binding molecule, or a combination thereof.

[0086] In some aspects the complexity of the pool of TFS beads is any integer between 0 and 1E9.

[0087] In some aspects, an ITC-conjugate is employed to functionalize proteins for immobilization to a TFS.

[0088] In some aspects, it may be advantageous to couple the ITC-conjugate to the protein, and subsequently join the ITC-conjugated protein to the TFS. This allows high concentration to afford fast kinetics and high yield of ITC-conjugate-functionalized protein with reduced expense and steric hinderance. Thus, a scheme is envisioned comprising: 1) an ITC-conjugate comprising a moiety (RG) for conjugating to the protein and a functional group (X) and 2) a complementary chemically-reactive moiety of the TFS, that can react to immobilize the protein analyte to the TFS. Examples of pairs of functional groups that may be suitable for joining the ITC-conjugate and TFS have been shown in PCT / US23 / 70077. The joining operation may be spontaneous, triggered (e.g., by light, temperature, solution composition change or other environmental change), or catalyzed.

[0089] In some aspects the ITC-conjugate comprises a linker. Suitable chemical linkers include but are not limited to alkyl, aryl, ether, amide, cycloalkyl, branched or linear structures or combinations thereof. The linker may include a -C(O)-, -O-, -S-, -S(O)-, -NH-, -C(O)O-, -C(O)Cl-C10 alkyl, -C(O)Cl-C10 alkyl-O-, -C(O)Cl-C10 alkyl-CO2-, -C(O)Cl-C10 alkyl-NH-, -C(O)Cl-C10 alkyl-S-, -C(O)Cl-C10 alkyl-C(O)-NH-, -C(O)Cl-10 alkyl-NH-C(O)-, -C1-C10 alkyl-, -C1-C10 alkyl-O-, -C1-C10 alkyl- CO2-, -C1-C10 alkyl-NH-, -C1-C10 alkyl-S-, -C1-C10 alkyl-C(O)-NH-, -C1-C10 alkyl-NH-C(O)-, -CH2CH2SO2-C1-C10 alkyl-, CH2C(0)-C1-C1-10 alkyl-, =N-(0 or N)-Cl-C10 alkyl-O-, =N-(O or N)-C1-C10 alkyl-NH-, =N-(O or N)-Cl-C10 alkyl-CO2-, =N-(O or N)-Cl-C10 alkyl-S-,, or

[0090] In some aspects the isothiocyanate reactive group of an ITC-conjugate may be substituted by any appropriate reactive group to formulate a more generalized conjugate for immobilization of proteins to solid supports, such as TFS supports.

[0091] In some aspects, one or more of the operations of the method are performed in any suitable sequential order or are simultaneously performed.

[0092] In some aspects, the method comprises (a) fragmenting peptides, protein, and / or protein complexes at any suitable step; (b) activating zero, 1, 2, or more moieties of each fragmented peptide, protein, and / or protein complex; and (c) joining the peptides to a TFS or solid support.

[0093] In some aspects, subunits of a given protein are co-immobilized directly or through their interaction with native subunits to a TFS. Subsequently, the one or more subunits may be simultaneously reverse translated and associated with metadata that includes information of the cell of origin.

[0094] In certain aspects, methods described herein comprise cross-linking peptides, protein, and / or protein complexes to facilitate discovery, identification, and investigation of protein interactomes.

[0095] In some aspects the uniformity of attachment sites across the plurality of beads of each TFS bead type provides normalization of analytes.

[0096] In some aspects, one or more of the steps in the herein described processes are automated, including mechanical, fluidic, pneumatic, electrical, optical, and / or control systems.Methods for Creating a PUMA Pool When the PUMAs are Read Separately

[0097] The present disclosure provides methods for generating a pool of Peptide Unique Molecular Association (PUMA) sequences that are distinguishable both from a known set of peptides or proteins (referred to as the "protein library") and from each other. This may be achieved, among other ways, by ensuring a minimum Hamming distance between each PUMA and any peptide or protein in the protein library, as well as between each PUMA within the pool.

[0098] The primary objective is to create a set of PUMAs that can be easily distinguished from a known set of peptides or proteins (the protein library) and from each other. This may be achieved by ensuring that each PUMA has a minimum Hamming distance of 'N' from any peptide or protein in the protein library and a minimum Hamming distance of 'M' from any other PUMA in the pool. The Hammingdistance between two strings of equal length is the number of positions at which the corresponding symbols are different.

[0099] By doing so, the identity and composition of the pool of PUMAs itself is innovative and novel and can be used across multiple scenarios where a unique set of peptide sequences may be beneficial for labeling, such as labeling proteins in a cell, in a spatial location, in applications like tracking proteinprotein interactions, monitoring enzymatic activity, identifying biomarkers in diagnostic assays, studying protein dynamics in live cell imaging, and facilitating targeted drug delivery systems. A Protein Library (P) and a PUMA Library (B) were utilized in the analysis process disclosed herein.Example Analysis Process

[0100] Hamming Distance Calculation: For each potential PUMA, calculate its Hamming distance from each peptide or protein in P and from each already selected PUMA in B.

[0101] Validation: A PUMA is considered valid if its Hamming distance from any peptide or protein in P is at least 'N', and its Hamming distance from any other PUMA in B is at least 'M'.

[0102] Iterative Generation: Generate potential PUMAs and validate them. If a PUMA meets the criteria, add it to B; otherwise, discard it.

[0103] Termination: The process continues until the desired number of PUMAs is generated or no more valid PUMAs can be found.Example Pseudocode for Generating Pool of PUMA Sequences function generatePUMAs(proteinLibrary, desiredPUMACount, hammingDistanceProtein, hammingDistancePUMA ):PUMALibrary = [] while len( PUMALibrary) < desiredPUMACount: potentialPUMA = generateRandomPUMA( ) if isValidPUMA(potentialPUMA, proteinLibrary, PUMALibrary, hammingDistanceProtein, hammingDistancePUMA ):P UMALibrary. append( potentialP UM A ) return PUMALibrary function isValidPUMA(PUMA, proteinLibrary, PUMALibrary, hammingDistanceProtein, hammingDistancePUMA ): for protein in proteinLibrary: if calculateHammingDistance(PUMA, protein) < hammingDistanceProtein: return false for existingPUMA in PUMALibrary:if calculateHammingDistance(PUMA, existingPUMA) < hammingDistancePUMA: return false return true function calculateHammingDistance( string 1, string2): distance = 0 for i in range(len( string 1 ) ): if string / [i ] != string2[i]: distance += 1 return distance

[0104] In a simplified scenario, we consider a small protein library and the process of generating a PUMA library. The protein library contains proteins "CAT" and "DOG". The PUMA library is generated with a minimum Hamming distance 'N' from P = 2 and 'M' from each other = 2. Through this process, valid PUMAs such as "ABC" and "CBD" are identified according to the specified Hamming distance criteria. In this simplified example, the PUMAs ABC and CBD are both valid according to the specified Hamming distance criteria. The process of generating and validating PUMAs continues until the desired number or a saturation point (where no more valid PUMAs can be generated) is reached.Example Bl: Small Protein and PUMA Library

[0105] In a model application, a subset of the human proteome was used as a model protein library to generate PUMAs (Table 7). The approach outlined above was employed for this purpose, and a list of valid PUMA sequences was generated as outlined in the table provided.

[0106] The list of valid PUMA sequences for this example was generated as follows and is shown inTable 1Table 1: Example List of Valid PUMA Sequences Generated from Model Protein Library

[0107] In an alternate embodiment, the disclosure may incorporate doubled PUMA sequences to add error tolerance without the constraint of the initial code having a hamming distance >1. This embodiment is particularly useful in cases where one or more amino acids might be missed during sequencing or processing. The doubled sequences for the PUMAs in the example are provided in the table.Example B2: Small Protein and PUMA Library with Error Tolerance

[0108] Building on the methods described in the previous examples, Example B2 illustrates the generation of PUMA libraries with enhanced error tolerance. This is achieved by creating PUMA sequences with various minimum Hamming distances, which can be used either as standalone sequences or in conjunction with doubled sequences for added reliability.Increasing error toleranceVaried Hamming Distances: In this example, PUMA sequences may be generated with increased minimum Hamming distances. This approach allows for greater error tolerance in sequence identification and analysis, reducing the likelihood of misidentification or sequence overlap.Single vs. Doubled Sequences: The method involves the generation of both single and doubled PUMA sequences. While single sequences benefit from increased Hamming distances, doubled sequences offer additional error tolerance, particularly useful in scenarios where amino acid sequencing might be prone to omissions or errors.PUMA Library Generation: The process of generating the PUMA library follows the same iterative and validation steps as described previously, but with the adjusted criteria for Hamming distances.

[0109] In this example, a PUMA library was generated with the specified criteria for increased Hamming distances of >3 from existing peptide or protein sequences or other PUMA sequences. The libraries generated included both single and doubled PUMA sequences. The Table 2 lists some of the PUMAs generated in this process, demonstrating the application of both the increased Hamming distance criteria and the concept of sequence doubling for error tolerance.Table 2: Example PUMA Library with Error Tolerance

[0110] By manipulating the Hamming distance criteria and incorporating sequence doubling, one can tailor the PUMA library to specific requirements of error tolerance and sequence differentiation. This example expands the utility of the PUMA library generation method by introducing variable Hamming distances and the option for sequence doubling. This allows for a more robust and error-tolerant approach to PUMA sequence generation, suitable for a wide range of applications where sequence accuracy and reliability are paramount.

[0111] In the design of PUMA libraries, the principles of error detection and error correction as defined by the Hamming distance take on significant relevance. Just as in coding theory, where the minimum Hamming distance between codewords determines their error detecting and correcting capabilities, PUMA sequences can similarly benefit from this concept to enhance their robustness against sequencing errors.

[0112] For example, in a PUMA library where each sequence has a minimum Hamming distance of 4 from any other sequence, we can interpret this as the library being capable of detecting up to 3 errors and correcting up to 1 error. This is derived from the coding theory principle that a code with a minimum Hamming distance 'd' can detect up to 'd-1' errors and correct [(d-l) / 2J errors.

[0113] In the context of PUMA sequences, this means that if a sequencing error alters up to three amino acids in a PUMA sequence, this error can still be detected because no other valid PUMA sequence will be within a Hamming distance of 3 of the original sequence. Furthermore, if there is only a single amino acid error in the sequence, the original PUMA can be accurately identified and corrected, as it will be the only valid PUMA within a Hamming distance of 1.

[0114] This error-correcting capability becomes especially powerful when considering doubled PUMA sequences. The duplication of sequences in PUMAs not only enhances the error tolerance but also provides a redundancy that can be exploited for error correction. For instance, if a sequencing error occurs in one half of the doubled sequence, the correct sequence can still be inferred from the other half, provided the error falls within the correctable limit as defined by the Hamming distance.

[0115] Thus, when designing PUMA libraries, especially for applications demanding high accuracy and reliability, integrating the principles of error detecting and correcting codes from coding theory offers a robust framework. This approach ensures not only the uniqueness and specificity of PUMAs but also their resilience to sequencing errors, thereby enhancing the overall integrity and utility of the PUMA-based applications.Example B3: Small Protein and PUMA Library Using Probability Transition Matrix

[0116] The present disclosure describes a novel method for generating a library of PUMAs that are specifically optimized to minimize conflict both with a predefined protein library and amongst the PUMAs themselves. This optimization is achieved through the application of a probability-based fuzzy matching algorithm, which represents a significant advancement over traditional methods for sequence comparison and uniqueness determination.

[0117] Central to this method is the use of a probability transition matrix. This matrix quantitatively expresses the likelihood of erroneous substitutions between different amino acids, a common occurrence in sequencing processes. By quantifying these probabilities, the matrix becomes a pivotal tool for calculating the transformation cost from one amino acid to another. This cost calculation may employ, as an example, the negative logarithm of the substitution probability, inversely correlating the likelihood of a substitution with its impact on sequence similarity. Consequently, rarer substitutions are assigned higher costs, effectively reflecting their greater significance in the overall measure of sequence dissimilarity. Other cost functions may provide similar results or may be preferred in certain scenarios.Probability transition matrix details1. Probability Transition Matrix: Let's denote this matrix as T, where each element Tzj represents the probability of substituting amino acid i with amino acidj during sequencing.2. Cost Function: The cost Cij of transforming amino acid i into amino acidj is defined in this example as the negative logarithm of the substitution probability Pij. Mathematically, this can be expressed as: Cij = -log(Pzj)Substitution Probability: Tzj is the probability that amino acid i is erroneously read or interpreted as amino acidj during sequencing. For instance, if Tij is high, it means there's a high likelihood of confusing amino acid i withj.Cost Calculation: The use of the negative logarithm inverts this probability, such that rare substitutions (low Tzj) are assigned higher costs. This is because a lower probability Tzj results in a higher negative logarithm value.Significance of Rare Substitutions: By assigning higher costs to rarer substitutions, the method effectively amplifies the importance of these unlikely events in the overall measure of sequence dissimilarity. This is crucial in sequencing, where the rarity of certain errors can be significant in identifying and characterizing sequences.

[0118] Example: If the probability of substituting a measurement of amino acid A (Alanine) with G (Glycine) is 0.01 (or 1%), and the probability of substituting a measurement of amino acid A with C (Cysteine) is 0.1 (or 10%) then the cost of these substitution would be calculated as:Cag= -log(Tag) = -log(O.Ol) = 2Cac= -log(Tac) = -log(0.1) = l

[0119] The likelihood of a a->g transition is lower, resulting in a higher cost in the cost function, thus resulting in a higher string distance for use in the fuzzy matching algorithm.

[0120] The fuzzy matching algorithm may utilize dynamic programming techniques to calculate a modified distance metric between peptide sequences similar to the Smith- Waterman [Smith, Temple F. & Waterman, Michael S. (1981). "Identification of Common Molecular Subsequences" (PDF). Journal of Molecular Biology. 147 (1): 195-197] or Needleman-Wunsch [ Needleman, Saul B. & Wunsch, Christian D. (1970). "A general method applicable to the search for similarities in the amino acid sequence of two proteins". Journal of Molecular Biology. 48 (3): 443-53] algorithms, which are foundational in bioinformatics for sequence alignment, but with an added layer to handle probabilistic costs of substitutions and indels. This metric is not limited to mere substitutions; it also encompasses indels (insertions and deletions), thus offering a more comprehensive and realistic assessment of sequence similarity. The algorithm is designed to ensure that each newly proposed PUMA maintains a certain threshold of uniqueness when compared to existing sequences, including both peptides or proteins from the library and previously generated PUMAs. The dynamic programming fuzzy matching algorithm, as applied in the context of peptide sequencing, works by calculating a modified distance metric between peptide sequences. This approach is more comprehensive than traditional methods as it accounts for substitutions and indels (insertions and deletions), thus offering a more realistic assessment of sequence similarity. The algorithm ensures that each newly proposed PUMA maintains a certain threshold of uniqueness when compared to existing sequences. These dynamic programming techniques systematically break down the complex problem into simpler sub-problems, calculating and storing the solutions of these sub-problems to avoid redundant computation.

[0121] One method of implementing such a fuzzy matching algorithm may be outlined as function fuzzyMatch(seql, seq2, probMatrix) initialize matrix dp with dimensions [len(seql)+l ][len(seq2)+l ] for ifrom 1 to len(seql ) dp[i][0] = dp[i-l][0] + costOflndel(seql[i-l]) for j from 1 to len(seq2) dp[0][j] = dp[0][j-l] + costOflndel(seq2[j-l]) for ifrom 1 to len(seql ) for j from 1 to len(seq2) costMatch = dp[i-l ][j-l ] + costOfSubstitution(seql [i-1 ], seq2[j-l ], probMatrix) costDelete = dp[i-l][j] + costOfIndel(seql[i-l])costinsert = dp[i][j-I] + costOflndel(seq2[j-l]) dp[i][j] = min(costMatch, costDelete, costinsert) return dp[len( seql ) ][len( seq2 ) ]

[0122] This aspect of the disclosure thus innovates by employing a probabilistic approach to Hamming distance calculation. Instead of the traditional method, where each substitution contributes equally to the distance, this approach utilizes the probability matrix to apply weights to each substitution. The weighted or probabilistic Hamming distance is calculated by summing the costs associated with each amino acid transformation in the sequences being compared. This method adjusts the distance metric to reflect the biological reality of sequencing errors, where certain amino acid substitutions are more probable than others.

[0123] Given the increased complexity introduced by the probabilistic approach, the disclosure may utilize various heuristic methods or even machine learning models to efficiently navigate the combinatorial space of potential PUMAs. These advanced computational techniques are particularly useful for managing the complexity of large datasets and for ensuring the generation of an optimal set of PUMAs that adhere to predefined similarity thresholds. Heuristic algorithms such as genetic algorithms or simulated annealing might be employed to intelligently explore the sequence space, striking a balance between exploring new potential sequences and refining known promising candidates.

[0124] The disclosure also contemplates alternative methods to further enhance the PUMA generation process. For instance, parallel processing techniques could be implemented to improve computational efficiency, especially beneficial when dealing with extensive libraries of proteins and PUMAs. Additionally, the disclosure could integrate other forms of sequence similarity assessments or machine learning models trained to recognize and predict the most suitable PUMAs based on learned patterns from the probability matrix and sequence data. The disclosure, while primarily focused on the use of Hamming Distance or a probability-based fuzzy matching algorithm for generating PUMAs, also considers a range of alternative approaches aimed at enhancing the PUMA generation process. These alternatives encompass various methodologies for weighting sequence comparisons, different distance calculation techniques, and advanced computational strategies to create unique, error-tolerant code spaces.

[0125] Machine learning models could be trained on large datasets to predict the suitability of potential PUMAs. These models would learn from the complex patterns in the probability matrix and sequence data, enabling them to make informed predictions about the uniqueness and error tolerance of new PUMAs.

[0126] The integration of deep learning algorithms, particularly those specialized in pattern recognition, could further enhance the PUMA generation process. These algorithms could analyze vastdatasets to identify optimal PUMA candidates, considering factors like sequence complexity, error probabilities, and similarity thresholds.

[0127] The disclosure may include customized distance metrics tailored to specific biological contexts or sequencing technologies. These metrics might consider factors like structural properties of amino acids, physicochemical characteristics, or other relationships. Dynamic weighting schemes could be implemented, where the weights assigned to amino acid substitutions are not static but adapt based on contextual factors such as sequence position, neighboring amino acids, specific biological functions, instrument variability, reagent lot variation, or environmental variation, among other methods. This adaptive approach could provide a more accurate reflection of the biological significance of each substitution.

[0128] A hybrid approach combining various methodologies could be employed for optimal PUMA generation. For instance, a combination of machine learning models for initial prediction of potential PUMAs, followed by a graph-based analysis for final selection, could be an effective strategy.Scaling to Large Proteomes and PUMA Sequence Spaces

[0129] The examples provided earlier focus on a small sample space with 88 proteins and 10 PUMA sequences for illustrative purposes. However, the methods outlined are scalable to much larger datasets. This scalability is a cornerstone of the approach, vital for applications in comprehensive proteomic analyses. The methods presented herein may scale to any number of proteins and PUMA sequences as follows.

[0130] The combinatorial basis of sequence generation is rooted in the diversity of the 20 canonical amino acids. For a peptide sequence of length L there are 20AL potential combinations. This vast sequence space is crucial for generating unique PUMA sequences on a large scale.

[0131] To illustrate, consider generating 1,000,000 unique PUMA sequences against 1,000,000 proteoforms resulting in 2,000,000 unique sequences. The minimum peptide length can be 5 as 20A5 = 3,200,000, exceeding 2,000,000, as indicated in the accompanying table. This table also reflects how the number of viable codes increases exponentially with the sequence length and varies according to different Hamming distances.

[0132] The methods can be adapted to analyze various proteomic complexities. For instance, in drug discovery or biomarker research, where numerous unique peptides are required, the approach can efficiently generate sequence diversity. Methods described above may be scaled as described to encompass as many proteins or peptides as needed with as many PUMA sequences as needed for a given application. In addition, the examples above are searching for uniqueness across the N-terminal end of the protein in the FASTA file assuming sequencing starts there. This could be amended to look for unique strings across entire proteins, or starting at expected cleavage sites such as expected upon trypsinization or any other degradation methods.Methods for Creating a PUMA Pool When Sequencing Results in Convolved Results

[0133] In some embodiments, the PUMA sequence and protein sequence may be measured in a convolved manner, in that the n-terminal (or other) amino acids from both the protein and PUMA are read simultaneously. This introduces other constraints into the design of the PUMA sequences as follows.Definition of Protein Library (P)

[0134] A predefined set of proteins is established, with each protein represented as a sequence of amino acids. For simplicity, let's consider a small protein library with sequences consisting of only 3 amino acids each.

[0135] Example Protein Library (P):

[0136] Protein 1: CAT

[0137] Protein 2: COT

[0138] In this example, each protein is represented using a three-letter amino acid code.Generation of PUMA Library (B)

[0139] A set of PUMAs is generated, matching the length of the peptides or proteins in the library (P). The PUMAs should be designed such that when convolved with the peptides or proteins, they result in a unique sequence that can be deconvolved back to its original components.

[0140] Initial PUMA Pool:

[0141] PUMA 1: DAB

[0142] PUMA 2: DOB

[0143] Each PUMA is a unique combination of three amino acids, distinct from those in the protein library.Convolution Process

[0144] Each PUMA from the library (B) is convolved with each peptides or protein in (P). Convolution here means combining a PUMA and a peptides or protein to generate a sequence where the order of amino acids is known, but their origin (PUMA or protein) is not.

[0145] Example Convolved Sequences:

[0146] Convolution of CAT (Protein 1) and DAB (PUMA 1): [C / D]-[A / A]-[T / B]

[0147] Convolution of COT (Protein 2) and DAB (PUMA 1): [C / D]-[O / A]-[T / B]

[0148] Convolution of CAT (Protein 1) and DOB (PUMA 2): [C / D]-[A / O]-[T / B]

[0149] Convolution of COT (Protein 2) and DOB (PUMA 2): [C / D]-[O / O]-[T / B]

[0150] Upon closer inspection, it becomes evident that the convolved sequence [C / D]-[A / O]-[T / B] is ambiguous. This sequence can result from the convolution of either CAT with DOB or COT with DAB. Thus, it does not meet the unique deconvolution requirement.Analysis of the Flawed Convolution:

[0151] The convolved sequence [C / D]-[A / O]-[T / B] does not uniquely map back to a single protein- PUMA pair.

[0152] This ambiguity arises due to the overlapping elements between the peptides or proteins and PUMAs.

[0153] This example serves as a crucial illustration of a PUMA sequence space that may need optimization in this particular use case. A randomized pool may result in errors or mismatches when aligning and identifying the peptides or proteins in the mixture. It highlights the necessity of careful PUMA design to ensure that every convolved sequence can be uniquely deconvolved to its original components for these applications.

[0154] This disclosure’s method accounts for such potential overlaps and ambiguities in the design phase, ensuring that the generated PUMAs, when convolved with the peptides or proteins, always produce distinct and uniquely identifiable sequences. This requirement is vital for the reliability of applications like protein sequencing, where the accuracy of sequence identification is paramount.Solution Approach

[0155] To overcome this issue, the method for generating PUMAs would include additional steps to check for and eliminate such ambiguities. This could involve more sophisticated algorithms that analyze all possible convolutions of peptides or proteins and PUMAs and ensure unique deconvolvability for each combination. The algorithm might need to iteratively adjust the PUMAs, test for unique deconvolution again, and repeat this process until a satisfactory PUMA set is obtained.

[0156] The present disclosure provides a method and system for generating a PUMA pool, wherein each PUMA can be uniquely deconvolved from a convolved peptide-protein sequence, even in thepresence of sequencing errors. The method involves generating a set of PUMAs and validating each against a predefined library of protein sequences to ensure unique deconvolution.Example B4: Generation PUMA Pool Assuming Error Free Measurement

[0157] This embodiment describes a method for generating a pool of Peptide Unique Molecular Identifiers (PUMAs) under the assumption of error-free measurement. The method ensures that each PUMA, when convolved with a protein sequence, results in a combined sequence that is uniquely identifiable and can be deconvolved back to its original constituents.Detailed Steps and Methodology1. Definition of Protein Library (P) :• A predefined set of proteins is established, with each protein represented as a sequence of amino acids.• For the purpose of convolution, protein sequences in the library are truncated to match the length of the PUMAs to be generated.2. Generation of PUMA Library (B):• A set of PUMAs is generated, where each PUMA is of the same length as the truncated proteins in the library (P).• The generation process involves creating PUMAs that are unique both individually and in their convolved forms with the proteins.3. Convolution Process:• Each PUMA is convolved with each truncated protein sequence.• The convolution of a PUMA and a protein involves generating all possible combinations of character swaps between the two sequences.• This process results in a set of convolved pairs for each PUMA-protein combination.4. Unique Deconvolution Requirement:• For each potential PUMA, the generated convolved pairs are analyzed to ensure uniqueness.• A PUMA is considered unique if none of its convolved pairs with any protein sequence maps to a pair consisting of an existing PUMA and a protein. function generatePUMAs(protein_sequences, num_pumas, puma_length): truncated_proteins = truncateProteinSequences(protein_sequences, puma_length) puma_library = [] while length of puma_library < num_pumas:potential_puma = generateRandomPUMA(puma_length) if isUniquePUMA(potential junta, truncated_pro terns, puma_library): append potential_puma to puma_library return puma_library function truncateProteinSequences(protein_sequences, puma_length): return list of sequences with each sequence truncated to puma_length function generateRandomPUMA(puma_length): return a randomly generated sequence of amino acids of length puma_length function isUniquePUMA(puma, truncated _protein_sequences, puma library): for each protein in truncated _protein_sequences: convolved joairs = convolve(puma, protein) for each pair in convolved _pairs: if pair is in a non-unique mapping: return False return True function convolve (puma, protein): return all possible combinations of character swaps between puma and proteinThe output of such a method can generate 100 or more unique PUMA sequences of length 5 (Table 3) against the sequences listed in Table 7.Table 3: Example PUMA Sequences

[0158] The embodiment outlined here presents a comprehensive method for generating a pool of PUMAs, each uniquely deconvolvable from its convolved sequences with proteins. This approach ensures that PUMAs are distinct and identifiable, even when measured in conjunction with protein sequences, thereby enhancing the accuracy and reliability of applications such as protein sequencing.Example B5: Generation of Error-Tolerant PUMA Pool

[0159] This embodiment outlines a method for generating a pool of PUMAs that not only ensures unique convolutions with protein sequences but also adheres to specific Hamming distance constraints. This method combines the convolution uniqueness as described in Example B4 with the Hammingdistance criteria from Example B2, thereby enhancing the distinctiveness and error tolerance of the generated PUMAs.

[0160] Detailed Steps and Methodology1. Definition of Protein Library (P) :• A predefined set of proteins is established, each represented as a sequence of amino acids.• Protein sequences are truncated to match the length of the PUMAs to be generated, facilitating direct comparison and convolution.2. Generation of PUMA Library (B):• PUMAs are generated to be of the same length as the truncated proteins in the library (P).• The generation process entails creating PUMAs that are unique both individually, in their convolved forms with proteins, and in compliance with Hamming distance requirements.3. Convolution and Hamming Distance Process:• Each PUMA undergoes convolution with each truncated protein sequence, generating all possible combinations of character swaps.• The convolved pairs are subjected to a uniqueness analysis, ensuring no pair maps to an existing PUMA-protein pair.• Additionally, each convolved pair is checked to ensure it maintains a minimum Hamming distance of N from any protein and M from any PUMA in the library.4. Unique Deconvolution and Hamming Distance Requirement:• For each potential PUMA, generated convolved pairs are scrutinized for both unique deconvolution and adherence to specified Hamming distance thresholds.• A PUMA is deemed acceptable if it satisfies both the convolution uniqueness criterion and the Hamming distance constraints relative to the proteins and other PUMAs.

[0161] Similar processes as above result in the generation of an exemplary 10 PUMA sequences (Table 4) that satisfy these criteria.Table 4: Example PUMA SequencesExample B6: Generation of Error-Tolerant PUMA Pool with probability transition matrices

[0162] This embodiment extends the method for generating Peptide Unique Molecular Identifiers (PUMAs) by incorporating a probabilistic approach to sequence comparison, enhancing the capacity to generate error-tolerant PUMAs. The method integrates a probability-based fuzzy matching algorithm with convolution uniqueness and Hamming distance criteria, as outlined in previous examples.Incorporation of Probability Matrix• Probability Transition Matrix Usage: Central to this method is the application of a probability transition matrix. This matrix quantifies the likelihood of erroneous substitutions between different amino acids. By incorporating these probabilities into the PUMA generation process, the method assigns varying weights to different amino acid substitutions, reflecting their likelihood and impact on sequence similarity.• Probabilistic Hamming Distance Calculation: The traditional Hamming distance calculation is augmented with this probability matrix. The weighted or probabilistic Hamming distance is computed by summing the transformation costs, derived from the negative logarithm of substitution probabilities, for each amino acid pair in the sequences being compared. This calculation results in a distance metric that is more representative of the biological reality of sequencing errors.Methodology and Unique Deconvolution Requirement• Generation and Uniqueness Check: PUMAs are generated and convolved with truncated protein sequences. The method assesses each PUMA for uniqueness not only based on convolution results but also considering the probabilistic Hamming distance from existing proteins and PUMAs.• Error-Tolerant Design: By integrating the probability matrix into the uniqueness checks, the generated PUMAs are optimized for error tolerance, ensuring that even in the presence of likely substitution errors, the PUMAs remain distinct and identifiable.

[0163] This embodiment, through the integration of a probability matrix with convolution and Hamming distance criteria, offers a more sophisticated approach to generating PUMAs. It ensures that the PUMAs are not only unique in their individual and convolved forms but also exhibit a high degreeof error tolerance, making them ideal for applications requiring precise molecular identification in potentially error-prone environments. This method exemplifies a robust, adaptable approach to PUMA generation, catering to the complex requirements of molecular identification and analysis in modern biological and biomedical research.Scaling to large proteomes and PUMA sequence spaces

[0164] In the standard scenario without convolution, the number of possible unique sequences for PUMAs of length L is 20AL (due to the 20 canonical amino acids). When considering convolution, the complexity increases because each position in the sequence has two possible sources (protein or PUMA). For each PUMA of length L, there are 2AL possible convolutions between a measured protein and a PUMA, given that each position in the sequence can come from either the protein or the PUMA.

[0165] When considering convolution, the complexity increases because each position in the sequence has two possible sources (protein or PUMA), but this doesn't directly halve the number of unique codes. Instead, it adds a layer of complexity in determining the uniqueness of each convolved sequence. The key is not just the number of sequences but the uniqueness of the convolution outcomes. Using 20AL / 2AL as an upper bound for the sequence space in the context of convolved PUMA and protein sequences can provide a theoretical upper limit. However, this calculation assumes that each convolved combination results in a unique sequence, which may not always be the case. In the case where some convolved combinations. In reality, the uniqueness of convolved sequences will likely be less than this upper bound, due to potential overlaps and similarities in the convolved sequences. The actual number of unique sequences will depend on the specific amino acid compositions of the proteins and PUMAs and how they interact when convolved.

[0166] Therefore, while 20AL / 2AL gives a broad upper limit, the effective sequence space for unique convolutions is likely to be smaller. This means that in practice, shorter peptides might suffice, but the exact length would need to be determined through more detailed analysis and possibly empirical testing. This is sufficient in this work for showing that there exists a length L of PUMAs to expand methods described to an arbitrarily large set of proteoforms and PUMAs

[0167] Thus, the methods described above can scale as necessary to arbitrary sizes of protein sequence sets and PUMA sequence sets.Incorporation of Alternative Amino Acids into PUMA Pools

[0168] The inclusion of non-natural amino acids in the generation of PUMAs expands the combinatorial possibilities and enhances the uniqueness and specificity of PUMAs. This embodiment describes modifications to the existing methodologies to incorporate non-natural amino acids and discusses their impact on the PUMA generation process.

[0169] This methodology expansion leads to several key impacts on PUMA generation. Firstly, the sequence space for PUMA generation is significantly broadened, which is particularly advantageous when dealing with large proteomes or in applications where high specificity is critical. The inclusion of non-natural amino acids not only enhances the specificity of PUMAs but also potentially improves the error tolerance of the system. The unique chemical properties of these amino acids can facilitate clearer differentiation in sequencing processes, thereby reducing the likelihood of erroneous sequence matches. Non natural amino acids may comprise modified or derivatized natural amino acids.

[0170] Extended Amino Acid Library: The set of amino acids used for generating PUMAs is expanded beyond the 20 canonical amino acids to include non-natural amino acids. This expansion increases the diversity of potential PUMAs, enhancing the ability to generate unique sequences.

[0171] Modification of Probability Matrix: The probability transition matrix is updated to include the likelihood of substitutions involving non-natural amino acids. The cost function between natural and alternative amino acids may be very high, resulting in high tolerance to substitution errors.

[0172] Adjustment of Hamming Distance Calculations: The algorithm for calculating Hamming distances and probabilistic Hamming distances is adapted to account for the increased diversity in amino acid substitutions. This may involve redefining the distance metrics to appropriately weight substitutions involving non-natural amino acids.

[0173] Increased Sequence Space: Incorporating non-natural amino acids significantly expands the sequence space, allowing for a much larger number of potential unique PUMAs. This is particularly beneficial when scaling to large proteomes or when high specificity is required.

[0174] Enhanced Specificity and Error Tolerance: The use of non-natural amino acids can enhance the specificity of PUMAs, as these amino acids are less likely to occur naturally in protein sequences. This reduces the likelihood of accidental matches with natural proteins. The error tolerance of the system may also be improved, as the distinct chemical properties of non-natural amino acids can lead to clearer differentiation in sequencing processes.

[0175] Incorporating non-natural amino acids into PUMA sequences offers a significant advancement in creating highly specific and diverse molecular identifiers.

[0176] In the case of convolved sequences, the incorporation of non-natural amino acids simplifies the deconvolution process. Since these amino acids are not present in the natural protein sequence space, their presence in a convolved sequence directly indicates their origin from the PUMA pool. This simplification is significant, especially in scenarios where the protein library is large or not fully characterized.

[0177] When considering the probability transition matrix method, the inclusion of alternative amino acids can lead to larger string distances between sequences. This is due to the modified side groups of these amino acids, which are likely to be quite distinct from the canonical amino acids, reducing the probability of erroneous substitutions. Consequently, this can enhance the accuracy of sequence differentiation and error tolerance in the PUMA generation process.

[0178] The guaranteed increase in Hamming distance with the inclusion of each non-natural amino acid simplifies the design for error tolerance and enhances the robustness of the PUMAs. This aspect is particularly beneficial in applications where complete knowledge of the protein library is not available, ensuring that the PUMAs remain unique and identifiable even in the presence of unknown proteoforms.

[0179] These considerations highlight the utility and versatility of incorporating alternative amino acids into PUMA pools, particularly in enhancing error tolerance and simplifying sequence analysis in complex proteomic environments.Example B7: Error Tolerant PUMA Pools Using Artificial Amino Acids

[0180] In the case of a PUMA library incorporating alternative amino acids, the PUMA sequences are guaranteed to be unique compared to the protein sequence space, simplifying the discovery process and resulting in more possible choices for PUMA sequences. In addition, for every alternative amino acid incorporated into a PUMA sequence, a hamming distance of 1 is guaranteed against the protein sequence space. This both simplifies design for error tolerance and can guarantee error tolerance in the case where the protein library is unknown or where all potential proteoforms in the library have not been characterized.

[0181] As an example, a library of PUMA sequences was generated incorporating at least 2 alternative amino acids from a set { 1 ,2, 3, 4, 5 } against the protein library sequences listed in Table 7 with Hamming Distance of at least 2 from known proteins and other PUMA codes.Table 5: Examples of PUMA sequences incorporating two alternative amino acidsAlternative Algorithmic Approaches: Dealing with Computational Complexity and Using Heuristic ApproachesScaling of Computational Complexity1. Hamming Distance Calculations: For each potential PUMA, the Hamming distance can be calculated with respect to each protein in the library and each already generated PUMA. If there are P proteins in the library and B PUMAs to be generated, the number of Hamming distance calculations for each new PUMA is proportional to P + B.2. PUMA Generation: Generating a PUMA involves checking through a large combinatorial space, especially if the PUMAs are long. The number of potential PUMAs is exponential in the length of the PUMAs.3. Validation Steps: Each potential PUMA can be validated against all proteins and previously selected PUMAs, which involves multiple comparisons for each new PUMA.4. Iterative Nature: The process is iterative. As the number of already selected PUMAs increases, the time to validate each new PUMA also increases, because it may be compared against an increasing number of existing PUMAs.

[0182] For small datasets (a small number of short proteins and a modest number of PUMAs), a straightforward algorithmic approach might be feasible. However, as the size of the dataset increases, so does the computational load, potentially making the problem computationally infeasible to solve exactly. In such cases, heuristic approaches can be beneficial:Reduced Search Space: Heuristic methods can intelligently reduce the search space by focusing on more promising PUMA candidates, thus avoiding the need to exhaustively generate and test every possible PUMA.Faster Convergence: Heuristics can lead to faster convergence on a suitable set of PUMAs that meet the criteria, especially in larger and more complex datasets.Balancing Exploration and Exploitation: They can balance between exploring new potential PUMAs and exploiting the most promising ones found so far, thus efficiently navigating the combinatorial space.Handling Real-World Data Variability: Heuristics are adept at dealing with the complexities and variabilities inherent in real-world biological data.

[0183] While heuristic approaches might not always be necessary for smaller datasets, they become increasingly relevant as the scale of the problem grows. In large-scale applications, where the number of proteins is high and a large PUMA library is required, heuristics provide a practical solution to generate a PUMA library that satisfies the specified Hamming distance constraints. They offer a balance between computational feasibility and the quality of the solutions, making them an essential tool in complex bioinformatics problems.

[0184] The task of generating a unique PUMA library for protein sequencing, especially with error tolerance, presents significant computational challenges. The complexity arises from several factors:Combinatorial Explosion: The number of possible convolutions between PUMAs and proteins grows exponentially with the size of the PUMA and protein libraries and the length of the sequences. This exponential growth leads to a combinatorial explosion in the search space. Error Tolerance: Accounting for errors (substitutions, insertions, deletions) further multiplies the complexity. The system may ensure unique deconvolution not only for error-free sequences but also for sequences with up to E errors. This requirement dramatically increases the number of possible scenarios to consider.Constraints Satisfaction: The problem is essentially one of constraint satisfaction, where each PUMA may meet the stringent condition of unique deconvolution in the presence of errors. Checking this condition for each potential PUMA is computationally intensive.Optimization Problem: Identifying the optimal set of PUMAs that maximizes the number of unique deconvolutions is an optimization problem, often with a large number of local optima that complicate the search for the global optimum.

[0185] Given these challenges, straightforward or brute-force approaches are often computationally infeasible, especially for large datasets. This is where heuristic approaches become valuable. Heuristic methods provide ways to intelligently navigate the search space, offering practical solutions to otherwise intractable problems. These methods do not guarantee a perfect solution but aim to find a good enough solution within a reasonable amount of time.How Heuristic Approaches HelpReducing Search Space: Heuristics can significantly reduce the search space by focusing on more promising areas, thus avoiding exhaustive enumeration of all possibilities.Handling Complexity: These methods are particularly adept at handling the complexity and nuances of real-world data, such as the variability in protein and peptide sequences.Iterative Refinement: Many heuristic algorithms work iteratively, progressively refining the solutions. This approach allows for continuous improvement and adaptation, often leading to better solutions over time.Balancing Exploration and Exploitation: Heuristic algorithms are designed to balance exploration (searching new areas) and exploitation (refining known good areas). This balance is crucial in avoiding local optima and finding near-optimal solutions efficiently.Leveraging Domain Knowledge: Heuristic methods can incorporate domain-specific knowledge, which can guide the search process more effectively. For instance, knowledge about typical error patterns in protein sequencing can be used to design more effective PUMA sequences.

[0186] By employing heuristics, the disclosure offers a robust, efficient, and adaptable solution suitable for the demanding needs of protein sequencing and related bioinformatics applications. The following examples describe some potential heuristic approaches but are not limiting.Greedy Algorithm

[0187] A greedy algorithm can be used to iteratively construct the PUMA library. This approach involves adding one PUMA at a time to the library, each time choosing a PUMA that maximizes the number of unique deconvolutions possible with the current set of proteins.Process:1. Initialize an empty PUMA library.2. Sequentially generate potential PUMAs.3. For each potential PUMA, calculate the number of unique deconvolutions it would add to the current PUMA library.4. Select the PUMA that maximizes this number and add it to the library.5. Repeat steps 2-4 until the desired number of PUMAs is reached or no further unique PUMAs can be added.Simulated Annealing

[0188] Simulated annealing is a probabilistic technique used to approximate the global optimum of a given function. In the context of PUMA generation, it can be used to iteratively improve a set of PUMAs, allowing for both uphill and downhill moves but gradually reducing the likelihood of downhill moves over time.Process:1. Start with an initial solution (set of PUMAs).2. At each iteration, slightly modify the set of PUMAs (e.g., changing one or more amino acids in a PUMA).3. If the new set has better or equal performance (in terms of unique deconvolution), accept it.4. If the new set has worse performance, accept it with a probability that decreases over time.5. Repeat this process until a satisfactory set of PUMAs is obtained or a predetermined number of iterations is reached.Machine Learning-Based Optimization

[0189] Machine learning models can be trained to predict the likelihood of unique deconvolution of a given PUMA-protein combination. These models can then guide the generation and selection of PUMAs.Process:1. Train a machine learning model on a dataset of PUMA-protein combinations and their deconvolution outcomes.2. Use the model to predict the deconvolution success of new PUMA-protein combinations.3. Iteratively generate and evaluate PUMAs using the model, selecting those with high predicted success.

[0190] Each of these heuristic approaches offers a different method for tackling the complex problem of generating a unique PUMA library for protein sequencing. They can be employed individually or in combination to optimize the PUMA generation process, taking into account the unique requirements of the sequencing application, such as error tolerance and the size of the protein library. The choice of heuristic method would depend on the specific characteristics of the problem at hand, including the complexity of the PUMA and protein sequences, the size of the dataset, and the computational resources available.Alternative Considerations in PUMA Sequence Design and Complexity

[0191] In addition to the traditional fixed-length PUMAs, an alternative embodiment involves the concept of variable-length PUMAs. This approach tailors the length of each PUMA sequence, adapting to the complexity and diversity of the protein environment it is designed to operate within.

[0192] In addition to the above error correction methods, the cycle or order of the amino acids in either the PUMA or protein may be substituted, misread, or not present, and may be accounted for in the PUMA pool design.

[0193] In the context of PUMA sequence analysis, while Hamming distance offers a robust measure for error detection and correction, alternative string distance measures may also be employed, each bringing advantages to address specific challenges in sequence comparison. Among these, the Levenshtein distance is particularly noteworthy for its ability to account for insertions and deletions, as well as substitutions, making it an invaluable tool in scenarios where sequence length variability or such types of errors are prevalent. The Damerau-Levenshtein distance may help account for the possibility of transpositions. Furthermore, the Jaccard distance, which evaluates the dissimilarity based on the presence or absence of elements in unordered sequences, offers an alternative perspective particularlysuited for set comparisons. Adaptations of dynamic programming approaches such as the Smith- Waterman and Needleman-Wunsch algorithms are pivotal for local and global sequence alignments, respectively, and may offer tailored solutions for biological sequence comparisons. For analog or highly quantized distance metrics such as described using the probability transition matrix, measurements such as one based off the Mahalanobis distance may be employed. It should be noted that this list is not exhaustive, and any suitable measures of the differences between strings or sets of amino acids may be used if needed.

[0194] In the case of studying protein-protein interactions, two protein sequences may be present and each convolved with the PUMA sequence in a single spatial location. This may necessitate applying the screening criteria above to a larger search space by either restricting the PUMA sequences to be a given hamming distance from any convolved protein pairs, a subset of expected convolved protein pairs, or a triply convolved sequence of two proteins and the PUMA sequence. In addition, in some aspects, two co-localized proteins that may be interacting could each have a PUMA tag, resulting in necessitating screening of PUMA codes against any combination of the four components. This may be limited to screening against suspected pairs or interacting proteins, such as p53 and Mdm2: p53 is a tumor suppressor, and Mdm2 is a regulator of p53, and have a strong interaction crucial in cell cycle regulation and cancer biology.

[0195] In addition to the informatic considerations above, PUMA design criteria may include chemical or physical constraints at to the properties of the peptides themselves. This could further reduce the space of available sequences, which may result in longer PUMA designs than theoretically shown above to account for this.

[0196] Other constraints may emerge in non-theoretical sequence spaces that necessitate longer PUMAs than described in the previous examples. Length may be added to these PUMAs without invalidating the methods and techniques above to create valid PUMA pools.Composition of PUMA Sequence Space and Network Deconvolution

[0197] Some embodiments relate to or include a library of PUMA sequences. Some embodiments include a PUMA sequence space. In some embodiments, the PUMA sequence space is base-20. In some embodiments, the PUMA sequence space is base-n. In some embodiments, n is an integer greater than or less than 20. In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50, or a range defined by any 2 of the aforementioned integers. In some embodiments, n is less than 2, less than 3, less than 4, less than 5, less than 6, less than 7, less than 8, less than 9, less than 10, less than 11, less than 12, less than 13, less than 14, less than 15, less than 16, less than 17, less than 18, less than 19, less than 20, less than 21, less than 22, less than 23, less than 24, less than 25, less than 26, less than 27, less than 28, less than 29, less than 30, less than 31, less than 32, less than 33, less than 34, less than 35, less than 36, less than 37, less than 38, less than 39, less than 40, less than 41, lessthan 42, less than 43, less than 44, less than 45, less than 46, less than 47, less than 48, less than 49, or less than 50. In some embodiments, n is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, or at least 50. In some embodiments, no PUMA sequence is the same as that of an analyte sequence in a given set of analytes. In some embodiments, no two PUMA sequences have the same sequence. In some embodiments, an insignificant fraction of PUMA sequences are the same as that of an analyte sequence in a given set of analytes. In some embodiments, an insignificant fraction of PUMA tags have the same sequence. In some embodiments, no PUMA sequence has hamming distance within 1, 2, 3, 4, 5, or more of an analyte sequence in a given set of analytes. In some embodiments, No two PUMA sequences have hamming distance within 1, 2, 3, 4, 5, or more.

[0198] In some aspects, the PUMA sequences are designed to maximize Hamming distance between sequences, rather than focusing solely on maintaining a minimum Hamming distance. This approach can further enhance the robustness and error-tolerance of the PUMA library.

[0199] Some embodiments relate to or include a library (e.g. a PUMA library, or a library or PUMA sequences). In some embodiments, the library includes a number of PUMA sequences. Some embodiments relate to or include a library wherein the number of PUMA sequences is >10, >100, >1000, or more than 1000. In some embodiments, the number of PUMA sequences is 5, 10, 15, 20, 25, 50, 75, 100, 150, 200, 250, 500, 750, 1000, 1250, 1500, 1750, 2000, 2500, or a range therebetween. In some embodiments the number of PUMA sequences is less than 5, less than 10, less than 15, less than 20, less than 25, less than 50, less than 75, less than 100, less than 150, less than 200, less than 250, less than 500, less than 750, less than 1000, less than 1250, less than 1500, less than 1750, less than 2000, or less than 2500. In some embodiments the number of PUMA sequences is at least 5, at least 10, at least 15, at least 20, at least 25, at least 50, at least 75, at least 100, at least 150, at least 200, at least 250, at least 500, at least 750, at least 1000, at least 1250, at least 1500, at least 1750, at least 2000, or at least 2500. In some embodiments that include a library, the given set of analytes is 100, 200, 1000 or more proteins. In some embodiments that include a library, the reverse translated and convolved PUMA information may be deconvoluted from the reverse translated information of any analyte in the set of analytes. In some embodiments that include a library, the reverse translated and convolved PUMA information may be deconvoluted from the reverse translated information of a majority of analytes in the set of analytes. In some embodiments that include a library, the reverse translated and convolved PUMA information may be deconvoluted from the reverse translated information for >10% of analytes in the set of analytes. In some embodiments that include a library, the reverse translated and convolvedPUMA information may be decon voluted from the reverse translated information for > 1 % of analytes in the set of analytes.

[0200] Some embodiments include a library of PUMA sequences for use in multiplexed assays. Some embodiments include a library of PUMA sequences for use in protein sequencing. Some embodiments include a library of PUMA sequences for use in single-cell protein sequencing. Some embodiments include a library of PUMA sequences for use in the field of spatial proteomics. Some embodiments include a library of PUMA sequences for use in the field of spatial multi-omics. Some embodiments include a library of PUMA sequences for use in the field of multi-omics. Some embodiments include a library of PUMA sequences for use in reverse translation protein sequencing.Spatial Proteomics

[0201] Spatial analysis adds another dimension to single-cell analysis. It seeks to analyze sample attributes of tissue, cell, and subcellular systems at high spatial resolution. Established and evolving spatial biology techniques include phenotypic analyses using microscopy, spatial transcriptomic tools, immuno-histochemistry (IHC), immuno-fluorescence (IF), and several other tools. These have created insights into biology through interrogation of whole tissue sections with single-cell resolution. Through these studies, researchers have observed that spatial relationships, cellular interactions, gradients, and microenvironments are drivers of cell differentiation, disease, and response to therapy. These among other observations suggest that a detailed understanding of how cells communicate, organize, and function within their environment is essential. Thus, interrogating cells within their spatial context may be crucial for understanding key aspects of cell biology, developmental biology, neurobiology, tumor biology, and important for drug and diagnostic development. For example, by screening large cohorts of patients with pathological complete responses and progressive disease, researchers determined spatial signatures associated with therapy response and resistance. [Fin, et al., Nat Cancer 4, 1036-1052 (2023)]

[0202] Spatially resolved analyses hold the promise to provide insight into new strategies to prevent or treat diseases, including infection, cancer, neurological conditions, and metabolic disorders. However, no techniques currently exists for spatially-resolved, high-depth, high-throughput, hypothesis-free interrogation of proteins. Current imaging and sorting methods rely on low-throughput, antibody-based detection, which can be a challenge for molecular targets without established antibodies. The current methods, compositions and assays address this gap.

[0203] A generalized method of the present disclosure is presented in FIG. 5. Novel concepts within the method include the disclosure of a peptide unique molecular association tag, (PUMA tag), PUMA code space, PUMA-conjugates, and PUMA surfaces. Each of these are explained in the Descriptions, Figures, and Examples herein.

[0204] In FIG. 6A-6C a tissue sample comprising a plurality of cells and a transfer substrate comprising PUMA tags is illustrated. In FIG. 6A PUMA constructs may be randomly-oriented, orderedor assigned to locations of the transfer substrate. In the figure FIG. 6B RG indicates a reactive group. It may be reactive to protein moieties directly or to a complementary reactive group that labels the protein analyte. Permeabilization of the cells releases proteins in proximity to PUMA-conjugates. The tissue sample may then be subjected to conditions sufficient to transfer the proteins to the TFS beads. The figure FIG. 6C illustrates the physical relationship and exemplary process steps to enable reverse translation. The number of proteins that may be captured per TFS bead is dependent on the bead size and surface density of PUMA-conjugates. For deep analysis is may be advantageous to capture proteins on an ultra-high density surface, and subsequently release them and recapture on a surface density suitable reverse translation operations.

[0205] Disclosed herein, in some embodiments, are spatial assay methods, comprising: (a) providing a plurality of amino acid identifier conjugates (PUMA-conjugates) each comprising: (x) a peptide, and (y) a reactive moiety capable of coupling to a protein, and optionally (z) an orthogonal reactive moiety capable of joining the PUMA-tagged protein conjugate to a solid support; (b) depositing the plurality of PUMA-conjugates on the solid support; (c) depositing a sample of cells on a surface of the solid support; (d) capturing proteins of the cells with the reactive moiety of the PUMA-conjugates, thereby joining the proteins of the cells with the PUMA-conjugates and generating PUMA-tagged protein conjugates; and (e) reverse translating the PUMA-tagged protein conjugates to obtain a memory oligonucleotide comprising sequences corresponding to identity and positional information of amino acids of the proteins, and corresponding with the PUMA tag; (f) obtaining sequence information for the memory oligonucleotide; and (g) based on the obtained sequence information, determining the identity and positional information of the amino acids of the proteins and assigning metadata that includes its original spatial location within the sample of cells.

[0206] Disclosed herein, in some embodiments, are spatial assay methods. A method may include providing amino acid identifier conjugates (PUMA-conjugates). PUMA-conjugates may each include a peptide. PUMA-conjugates may each include a reactive moiety. A reactive moiety may be capable of coupling to a protein. PUMA-conjugates may each include an orthogonal reactive moiety. An orthogonal reactive moiety may be capable of joining a PUMA-tagged protein conjugate to a support such as a solid support. A method may include depositing the plurality of PUMA-conjugates on the support. A method may include depositing a sample of cells on a surface of the support. A method may include capturing proteins of the cells. The capturing may be performed with the reactive moiety of the PUMA-conjugates. A method may include joining the proteins of the cells with the PUMA-conjugates. A method may include generating PUMA-tagged protein conjugates. A method may include reverse translating PUMA-tagged protein conjugates. The reverse translating may result in or produce an oligonucleotide such as a memory oligonucleotide. The oligonucleotide or memory oligonucleotide may a sequence corresponding to identity information of amino acids of proteins. The oligonucleotide or memory oligonucleotide may a sequence corresponding to positional information of amino acids of proteins. The oligonucleotide or memory oligonucleotide may a sequence corresponding with thePUMA tag. The oligonucleotide or memory oligonucleotide may include sequences corresponding to identity and positional information of amino acids of the proteins, and corresponding with the PUMA tag. A method may include obtaining sequence information for the memory oligonucleotide. A method may include determining identity information of the amino acids of the proteins. A method may include determining positional information of the amino acids of the proteins. A method may include assigning metadata that includes a protein’s spatial location or original spatial location within a sample (e.g. of cells). Determining identity or positional information may be based on sequence information (e.g. the obtained sequence information).Process 2000 - Spatial, PUMA-conjugate Tagging

[0207] Disclosed herein, in some embodiments, are methods for associating the metadata that includes spatial location together with identity and positional information of an amino acid residue of a protein, the method comprising: (a) providing a plurality of amino acid identifier conjugates (PUMA- conjugates) each comprising (x) a peptide, and (y) a reactive moiety capable of coupling to a protein, and optionally (z) an orthogonal reactive moiety capable of joining the PUMA-tagged protein conjugate to solid support; (b) depositing a plurality of PUMA-conjugates on a solid support; (c) depositing the sample of the plurality of cells on the surface of the solid support; (d) capturing cellular components from one or more cells of the plurality of cells with the capture molecule of the PUMA-conjugate, thereby joining the captured material from the one or more cells with a PUMA-conjugate; and (e) reverse translating the immobilized PUMA-tagged protein conjugate information; (f) obtaining sequence information for the memory oligo; and (g) based on the obtained sequence information, determining identity and positional information of an amino acid residue of the protein and assigning metadata that includes its original spatial location within the sample of the plurality of cells.

[0208] Disclosed herein, in some embodiments, are methods for associating the metadata that includes spatial location together with identity and positional information of an amino acid residue of a protein. The method may include providing a plurality of amino acid identifier conjugates (PUMA-conjugates). PUMA-conjugates may include a peptide and a reactive moiety capable of coupling to a protein, and optionally an orthogonal reactive moiety capable of joining the PUMA-tagged protein conjugate to solid support. The method may include depositing a plurality of PUMA-conjugates on a solid support. The method may include depositing the sample of the plurality of cells on the surface of the solid support. The method may include capturing cellular components from one or more cells of the plurality of cells with the capture molecule of the PUMA-conjugate. The method may include joining the captured material from the one or more cells with a PUMA-conjugate. The method may include reverse translating the immobilized PUMA-tagged protein conjugate information. The method may include obtaining sequence information for the memory oligo. The method may include based on the obtained sequence information, determining identity and positional information of an amino acid residue of theprotein and assigning metadata that includes its original spatial location within the sample of the plurality of cells.Process 2100 - Spatial, TFS Tagging

[0209] Disclosed herein, in some embodiments, are methods for associating the metadata that includes spatial location together with identity and positional information of an amino acid residue of a protein, the method comprising: (a) providing a sample having a plurality of cells on a support; (b) providing a transfer substrate comprising (x) a solid support having a complementary shape to interact with the sample, and (y) an array of “tri-functional support” (TFS) beads having a plurality of bead types defined by their PUMA sequence, and wherein each bead of a bead type comprises (a) a solid support surface, (b) an immobilized peptide unique molecular association (PUMA) tag having the same constant segment sequence, and (c) a reactive moiety(s) associated with each PUMA tag for binding an amino acid residue of a protein; (c) contacting the sample with the transfer substrate allowing the TFS beads to interact with the sample; (d) contacting the transfer substrate with reagents that permeabilize the cell membranes; (e) contacting the transfer substrate with processing reagents; (f) optionally dissociating the TFS beads from the solid support, collecting, and pooling the TFS beads; (g) optionally releasing the PUMA-tagged protein conjugates from the TFS beads; (h) optionally contacting the released PUMA-tagged protein conjugate pool with a solid support functionalized to support reverse translation; (i) reverse translating the immobilized PUMA-tagged protein conjugate information; (j) obtaining sequence information for the memory oligo; and (k) based on the obtained sequence information, determining identity and positional information of an amino acid residue of the protein and assigning metadata that includes its spatial location within the sample.

[0210] Disclosed herein, in some embodiments, are methods for associating the metadata that includes spatial location together with identity and positional information of an amino acid residue of a protein. The method may include providing a sample having a plurality of cells on a support. The method may include providing a transfer substrate. The substrate may include a solid support having a complementary shape to interact with the sample. The substrate may include an array of multifunctional support beads such as “tri-functional support” (TFS) beads having a plurality of bead types defined by their PUMA sequence. A bead of a bead type may include a solid support surface. A bead of a bead type may include an immobilized peptide unique molecular association (PUMA) tag having the same constant segment sequence. A bead of a bead type may include a reactive moiety(s) associated with each PUMA tag for binding an amino acid residue of a protein. The method may include contacting the sample with the transfer substrate allowing the beads (e.g. TFS beads) to interact with the sample. The method may include contacting the transfer substrate with reagents that permeabilize the cell membranes. The method may include contacting the transfer substrate with processing reagents. The method may include dissociating the TFS beads from the solid support. The method may include collecting or pooling the beads (e.g. TFS beads. The method may include releasing the PUMA-taggedprotein conjugates from the beads (e.g. TFS beads). The method may include contacting the released PUMA-tagged protein conjugate pool with a solid support functionalized to support reverse translation. The method may include reverse translating the immobilized PUMA-tagged protein conjugate information. The method may include obtaining sequence information for the memory oligo. The method may include: based on the obtained sequence information, determining identity and positional information of an amino acid residue of the protein and assigning metadata that includes its spatial location within the sample.Process 2200 - Spatial, S-TFS Tagging, Plurality

[0211] Disclosed herein, in some embodiments, are methods for associating the metadata that includes spatial location together with identity and positional information of a plurality of amino acid residue of a plurality of proteins, the method comprising: (a) providing a sample having a plurality of cells on a support; (b) providing a transfer substrate comprising (x) a solid support having a complementary shape to interact with the sample, and (y) an array of “tri-functional support” (TFS) beads having a plurality of bead types defined by their PUMA sequence, and wherein each bead of a bead type comprises (a) a solid support surface, (b) an immobilized peptide unique molecular association (PUMA) tag having the same constant segment sequence, and (c) a reactive moiety(s) associated with each PUMA tag for binding an amino acid residue of a protein; (c) contacting the sample with the transfer substrate allowing the TFS beads to interact with the sample; (d) contacting the transfer substrate with reagents that permeabilize the cell membranes; (e) contacting the transfer substrate with processing reagents; (f) optionally dissociating the TFS beads from the solid support, collecting, and pooling the TFS beads; (g) optionally releasing the PUMA-tagged protein conjugates from the TFS beads; (h) optionally contacting the released PUMA-tagged protein conjugate pool with a solid support functionalized to support reverse translation; (i) reverse translating the immobilized PUMA-tagged protein conjugate information; (j) obtaining sequence information for the memory oligos; and (k) based on the obtained sequence information, determining identity and positional information of a plurality of amino acid residue of a plurality of proteins and assigning metadata that includes their spatial location within the sample.

[0212] Disclosed herein, in some embodiments, are methods for associating the metadata that includes spatial location together with identity and positional information of a plurality of amino acid residue of a plurality of proteins. The method may include providing a sample having a plurality of cells on a support. The method may include providing a transfer substrate. The transfer substrate may include a solid support having a complementary shape to interact with the sample. The transfer substrate may include an array of multifunctional support beads such as “tri-functional support” (TFS) beads having a plurality of bead types defined by their PUMA sequence. A bead type may include a solid support surface. A bead type may include an immobilized peptide unique molecular association (PUMA) tag having the same constant segment sequence. A bead type may include a reactive moiety(s) associatedwith each PUMA tag for binding an amino acid residue of a protein. The method may include contacting the sample with the transfer substrate allowing the multifunctional support beads (e.g. TFS beads) to interact with the sample. The method may include contacting the transfer substrate with reagents that permeabilize the cell membranes. The method may include contacting the transfer substrate with processing reagents. The method may include optionally dissociating multifunctional support beads (e.g. TFS beads) from the solid support, collecting, and pooling the multifunctional support beads (e.g. TFS beads). The method may include optionally releasing the PUMA-tagged protein conjugates from the multifunctional support beads (e.g. TFS beads). The method may include optionally contacting the released PUMA-tagged protein conjugate pool with a solid support functionalized to support reverse translation. The method may include reverse translating the immobilized PUMA-tagged protein conjugate information. The method may include obtaining sequence information for the memory oligos. The method may include based on the obtained sequence information, determining identity and positional information of a plurality of amino acid residue of a plurality of proteins and assigning metadata that includes their spatial location within the sample.Process 2300 - Spatial, Deep Sequencing via Employing Expandable Elastomers

[0213] Disclosed herein, in some embodiments, are methods for associating the metadata that includes spatial location together with identity and positional information of an amino acid residue of a protein, the method comprising: (a- 1) providing a sample having a plurality of cells on a support; (a-2) soaking the sample in an expandable elastomer; (a-3) expanding the elastomer; completing steps (b) through (k) of a method described herein.

[0214] Disclosed herein, in some embodiments, are methods for associating the metadata that includes spatial location together with identity and positional information of an amino acid residue of a protein. The method may include providing a sample having a plurality of cells on a support. The method may include soaking the sample in an expandable elastomer. The method may include expanding the elastomer. The method may include completing additional method steps such as steps (b) through (k) of a method described herein.

[0215] In some aspects, the expanded elastomer sample (e.g. after step (a-3)) is sectioned and the sections are isolated into indexed compartments, such as the wells of a microtiter plate. Thereafter processing reagents and solid supports capable to facilitate reverse translation may be delivered to the compartments, and the released proteins may be reverse translated.Process 2400 - Spatial, Deep Sequencing by Employing Mechanical Devices

[0216] Disclosed herein, in some embodiments, are methods for associating the metadata that includes spatial location together with identity and positional information of an amino acid residue of a protein, the method comprising: (a- 1 ) providing a sample having a plurality of cells on a support; (a-2) directingselected cells of the sample via mechanical methods to indexed compartments, and thereafter reverse translating analytes of the sample. Mechanical means include picking cells using commercial instruments such as the Q-pix from Molecular Devices, or using a transfer stamp, for example an embossed stamp having nano-pins spaced appropriately, e.g., contact printing.

[0217] Disclosed herein, in some embodiments, are methods for associating the metadata that includes spatial location together with identity and positional information of an amino acid residue of a protein. The method may include providing a sample having a plurality of cells on a support. The method may include directing selected cells of the sample via mechanical methods to indexed compartments. The method may include reverse translating analytes of the sample.

[0218] In some aspects, permeability reagents are dispensed to selectively release proteins from specific regions of the sample, and the released proteins are directed to indexed compartments, such as microfluidic devices or wells of a micro well plate. In some aspects the reagents are dispensed using a piezo nozzle(s), or inkjet nozzle(s). In some embodiments, reagents are arrayed using commercially available instruments such as Beckton Coulter, Molecular devices, and others.Spatial Aspects

[0219] A biological sample may comprise a tissue sample comprising a plurality of cells. In some aspects, the tissue section is a tumor biopsy, a histological cross-section, a fresh sample, a FFPE sample, or a fresh-frozen sample, or other natural or synthetic biological structure.

[0220] In some aspects, the tissue sample is stained with H&E, immunofluorescent antibodies, etc.

[0221] In some aspects, the cell support is a slide, a tray, a culture dish, or any suitable support to which a sample may be associated.

[0222] In some aspects, the transfer substrate comprises a single layer, a double layer and / or a multilayer of one TFS bead type positioned at each a priori defined locus of an array on the transfer substrate.

[0223] In some aspects, beads are delivered from a solution of TFS beads that share the same PUMA constant segment sequence.

[0224] In some aspects, TFS beads having the same PUMA sequence are spotted in the same location of the transfer substrate. In some aspects the spotting solution comprises molecules to improve spot uniformity characteristics such as binder(s), buffer(s), and surfactant(s), e.g., sucrose monolaurate.

[0225] In some aspects, TFS beads having the same PUMA sequence are delivered using a pipet, and / or spray nozzle, and any other appropriate means.

[0226] In some aspects, the transfer substrate comprises a single layer, a double layer and / or a multilayer of many TFS bead types positioned at each a priori defined locus of an array on a solid support.

[0227] In some aspects, the transfer substrate or the cell support comprise a microfluidic device upon which the sample is deposited, and / or the TFS beads are deposited to facilitate various transferoperations and reagent removal and exchange. In some embodiments the microfluidic device further comprises active pixels, such as CMOS pixels or electrode pixels that may facilitate imaging, or drive fluidic movements, such as in cameras and / or electro wetting devices.

[0228] In some aspects, fiducial marks or fiducial beads are provided.

[0229] In some aspects, a pool of TFS beads is randomly loaded onto the transfer substrate to form a single layer, a double layer and / or a multilayer of randomly positioned bead types on a solid support.

[0230] In some aspects, beads are delivered from a solution of pooled TFS bead types.

[0231] In some aspects, a pool of TFS beads is delivered to the transfer substrate to adopt random positions, for example via spin coating, pipetting, and / or spray coating, or any appropriate method for positioning and / or delivering beads.

[0232] In some aspects, TFS beads are delivered by mechanically pressing them into a micro-sectioned sample. In some aspects they are embedded into the sample.

[0233] In some aspects the spotting solution comprises a hydrogel, components to control pH, surfactant(s), and / or salts to achieve uniform spotting conditions.

[0234] In some aspects the spot size is between lum and lOOum

[0235] In some aspects the transfer substrate is patterned with vias or wells to direct TFS beads into specific locations on the transfer substrate surface, for example, by nanoimprinting, embossing, photoetching, or any suitable method. And in some aspects, TFS beads are manipulated using mechanical forces to occupy pre-positioned vias or wells of a nanopatterned or micropatterned substrate. In some embodiments, TFS beads are manipulated using surface energy forces.

[0236] In some embodiments, the PUMA bead pool includes decode oligonucleotides that can be used to pre-map bead locations using techniques similar to those employed in microarray technologies. This may allow spatial tracking of different PUMA sequences in array formats.

[0237] Is some aspects, the transfer substrate surface energy is modulated using polymers that may comprise: poly(acrylamide)s, poly(methacrylamide)s, poly(N-alkyl acrylamide)s, poly(N,N-dialkyl acrylamide)s, poly(N-alkyl methacrylamide) s, poly(N,N-dialkylmethacrylamide)s, poly(methacrylate)s, poly(acrylate)s, poly(styrene)s, poly(vinyl ether)s, poly(vinyl ester)s, poly(hydroxyalkylacrylamide)s, poly(hydroxyalkylmethacrylates)s, poly(hydroxyalkylmethacrylamide)s, poly(vinyl phosphonate)s, poly(sulfobetaine)s, poly(hydroxyalkylacrylate)s, polyenes, polydienes, poly(carbonate)s, poly(imide)s, poly(amide)s, poly(ester)s, poly(ether)s, poly(sulfone)s, poly(glycol)s, poly(saccharide)s, poly(vinyl pyrrolidone)s, poly(vinyl imidazole)s, poly(vinyl imidazoleum halide)s, poly(ethylene imine)s, poly(ethylene glycol)s, poly(ethylene oxide)s, poly( vinyl chloride)s, poly(vinyl alcohol)s, poly( vinyl pyridine)s, poly(vinyl napthalene)s, poly(phenylene)s, poly(ketone)s, poly(ether ketone)s, fluoropolymers, perfluoropoly ethers, poly(oxazoline)s, poly(ester)s, poly(benzimidazole)s, poly(ether sulfone)s, poly(phenylene vinylene)s, poly(thiophene)s, poly(vinyl phenol)s, poly(caprolactone)s, poly(caprolactam)s, poly(lactide)s, poly(glycolide)s, poly(propylene glycol)s,poly(hydroxyalkanoate)s, poly(urethane)s, poly(urea)s, poly(tetramethylene oxide)s, dextran, poly(hyaluronic acid)s, poly(imines), poly(amine)s, poly(thioether)s, poly(phosphonate)s, poly(phosphazene)s, poly(siloxane)s, poly( vinyl fluoride)s, poly(vinylidene fluoride)s, poly(tetrafluoroethylene)s, poly(oligoethylene glycol methacrylate)s, poly(oligoethylene glycol acrylate)s, poly(oligoethylene glycol acrylamide)s, poly(oligoethylene glycol methacrylamide)s, chitosans, agaroses, alginates, poly(acrylic acid)s, poly(methacrylic acid)s, poly(N- isopropylacrylamide)s, poly(itaconate)s, poly(pyrrole)s, PEDOTs, poly(aniline)s, poly(acetylene)s, poly(carbazole)s, poly(acrylonitrile)s, poly(phosphorylcholine methacrylate)s, poly(phosphorylcholine acrylate)s, poly(phosphorylcholine acrylamide)s, poly(phosphorylcholine methacrylamide) s, poly(sulfobetaine methacrylate)s, poly(sulfobetaine acrylate)s, poly(sulfobetaine methacrylamide) s, poly(sulfobetaine acrylamide)s, poly(ethylene)s, poly(propylene)s, poly(ethylene terephthalate)s, poly(methyl methacrylate)s, poly(acryloxymorpholine)s, poly(dialkylaminoalkyl methacrylate)s, poly(vinyl boronic acid)s, cellulose acetates, cellulose nitrates, nylons, and substituted derivatives, in linear, branched, star, cyclic and crosslinked configurations, as copolymers, random copolymers, block and multiblock co-polymer configurations.

[0238] In some aspect, a 2-dimensional surface or a 3-dimensional hydrogel matrix may be substituted for the TFS beads. PUMA-conjugates may be incorporated into the hydrogel in ordered, semi-ordered, or random configurations using spin coating, nozzles, spotting, ink jet, or any suitable technology. Sample proteins may be transferred to the hydrogel via methods, such as Western blot or electrophoresis, known to those of ordinary skill in the art. PUMA-tagged proteins may subsequently be released or detached from the surface for processing.

[0239] In some aspects, TFS beads are delivered directly to the tissue sample comprising a plurality of cells, without the use of a separate transfer substrate. In some aspects the TFS beads are delivered in ordered or random configurations.

[0240] In some aspects, processing reagents comprise fixatives, for example aldehydes, such as glutaraldehyde and formaldehyde; acids, such as acetic acid and Davidson’s AFA, acetone, oxidizing agents, such as potassium permanganate, osmium tetroxide, potassium, sodium dichromate, and chromate), alcohols (e.g., methanol, ethanol), and / or linkers and conjugation agents.

[0241] In some aspects the permeabilizing reagents are chosen from a list of permeabilizing reagents and / or kits that include: organic solvents, such as methanol and acetone, detergents, such as saponin, Triton X-100, Tween-20, SDS, deoxycholate, and lytic enzymes, such as lysozyme, zymolase, salts, antibiotics, reducing agents, and reagents formulated and available commercially through BD Biosciences, ThermoFisher, Fisher Scientific, Sigma and others. [Jamur, et.al., Methods Mol Biol. (2010) 588:63-6 ; Naglak et.al., Bioprocess Technol. (1990) 9:177-205]

[0242] In some aspects, physical disruption, for example using acoustic or pressure waves, and / or heat or freeze thaw cycles and / or osmotic pressure are employed to permeabilize the cell.

[0243] In some aspects, processing reagents include: detergent(s), surfactant(s), chaotropic agent(s), reducing agent(s), alkylating agent(s), derivatizing agents, protease inhibitors, and / or enzymes used modify the cellular components, such as collagenase and hyaluronidase, DNase, RNase, CRISPR, and chemicals for manipulation of protein samples.

[0244] In some aspects, enzymes used for permeabilization and / or processing are engineered as described in [carrier protein provisional patent number ## / ###,###] which is herein incorporated by reference in its entirety, to render them spectators in reverse translation operations.

[0245] In some aspects, the sample may be subjected to deparaffinization and / or decrosslinking prior to analysis.

[0246] In some aspects, the captured and PUMA tagged proteins are released from the TFS bead surface by cleaving a chemical moiety, or by action of an endonuclease, exonuclease, reducing agents, CRISPR enzyme, or any suitable means.

[0247] In some aspects, the biological molecule tagged by the PUMA tag is any molecule that can be reverse translated.

[0248] In some embodiments the solid support, or TFS, comprises chemical moieties, which may support reverse translation as described herein, such as: amine, carboxyl, NHS, aldehyde, azide, alkyne, maleimide, thiol, hydrazine, oxyamine, tetrazine, an alkene, acryl, trans-cyclooctene (TCO), a DBCO, a bicyclononyne, a norbornene, a strained alkyne, strained alkene, and / or a derivative thereof.

[0249] In some aspects, a sample is processed with agents that target highly-abundant proteins for depletion from the analysis. For example, IgM antibodies are provided for agglutination of highly- abundant proteins.

[0250] In some aspects, a sample is processed with agents, such as IgG conjugates target selected proteins for enrichment in the analysis.

[0251] In some aspects, step (d) of process 2100 is optional. Non-permeabilized cells may be particularly useful when targeting cell surface receptors, membrane proteins or proteins otherwise associated with the exterior cell membrane.

[0252] In some aspects, the reactive moiety of the TFS may bind either directly or indirectly to an amino acid residue of the peptide.

[0253] In some aspects, reagents are removed and / or exchanged between any appropriate step(s) of the method disclosed herein. Methods to exchange reagents may include: iterative serial dilutions reagents in contact with sample cells, TFS beads, PUMA-conjugate membranes, or centrifugation, filtration, or another suitable means to exchange or remove reagents.

[0254] In some aspects, the metadata that includes spatial location or relative spatial location provides a probabilistic estimate of the possible subset of spatial locations or relative spatial locations from which the protein may have originated.

[0255] In some aspects, one or more operations of the method are repeated one or more times to increase a step yield of the method. For example, in specific aspects, operations (b), (c), (d), (e), (f), (g), and / or (h) are repeated one or more times to increase the step yield.

[0256] In some aspects, a mono layer, double layer, etc. of the plurality of cells on the substrate may be ablated and the steps of process 2100 may be repeated to provide additional information of the sample.

[0257] In some aspects the PUMA tag is joined to protein analytes via action of an enzyme. For example, sortase may be employed to add various amino acid sequences to the N-terminus of a protein. Exopeptidase may be used to prune incorporated amino acids sequences to remove amino acids superfluous to the function of the PUMA tag.

[0258] In some aspects, the PUMA tag joined to the TFS is an amino acid sequence between 4 and 16 monomers long with a free amino terminus.

[0259] In some aspects the TFS comprises a bead having diameter between 0.1 and lOOuM.

[0260] In some aspects, the PUMA tag is joined to the TFS using a trifunctional molecule.

[0261] In some aspects the complexity of the pool of TFS beads is any integer between 0 and 1E9.

[0262] In some aspects, an ITC-conjugate is employed to functionalize proteins for immobilization to a TFS. In some aspects, it may be advantageous to couple the ITC-conjugate to the protein, and subsequently join the ITC-conjugated protein to the TFS. This allows high concentration to afford fast kinetics and high yield of ITC-conjugate-functionalized protein with reduced expense and steric hinderance. Thus, a scheme is envisioned comprising: 1) an ITC-conjugate comprising a moiety (RG) for conjugating to the protein and a functional group (X) and 2) a complementary chemically-reactive moiety of the TFS, that can react to immobilize the protein analyte to the TFS. Examples of pairs of functional groups that may be suitable for joining the ITC-conjugate and TFS have been shown in PCT / US23 / 70077. The joining operation may be spontaneous, triggered (e.g., by light, temperature, solution composition change or other environmental change), or catalyzed.

[0263] In some embodiments, (d) and (e) are performed prior to step (c).

[0264] In some aspects, one or more of the operations of the method are performed in any suitable sequential order, or are simultaneously performed.

[0265] In some aspects, the method comprises (a) fragmenting peptides, protein, and / or protein complexes at any suitable step; (b) activating zero, 1, 2, or more moieties of each fragmented peptide, protein, and / or protein complex; and (c) joining the peptides to a TFS or solid support.

[0266] In some aspects, the density of PUMA-conjugates immobilized on the TFS surface is appropriate for in situ reverse translation.

[0267] In some aspects, subunits of a given protein are co-immobilized directly or through their interaction with native subunits to a TFS. Subsequently, the one or more subunits may be simultaneously reverse translated and associated with metadata that includes the spatial location in the original sample..

[0268] In certain aspects, methods described herein comprise cross-linking peptides, protein, and / or protein complexes to facilitate discovery, identification, and investigation of protein interactomes.

[0269] In some aspects, proteins that are tagged with one or more PUMA tags that may be assigned to a location or coordinate within a tissue sample to facilitate identification of the cell from which a protein originated.

[0270] In some aspects the PUMA tags comprise one or more non-natural amino acid R-groups

[0271] In some aspects, digital imaging mass spec may be employed to provide a map of peptides immobilized on a gridded surface and to which a tissue sample is applied

[0272] In some aspects, the PUMA pool is designed to include a varied number of sequences, ranging from smaller sets to considerably larger collections. Specifically, the pool might consist of 5, 10, 15, 20, 50, 100, or even larger numbers of sequences.

[0273] In some aspects, the protein library the PUMA pool is designed against may contain 5, 10, 50 100, 1000, 1 million, or more proteoforms.

[0274] In some aspects, the protein library the PUMA pool is designed against is unknown or partially unknown.

[0275] In some aspects, the PUMA pool consists of only 5, 6, 7, 10, 15 or another subset of canonical amino acids

[0276] In some aspects, the PUMA pool comprises 1, 2, 3, 4, 5 or more non canonical amino acids

[0277] In some aspects, the creation of PUMA pools incorporates a range of amino acids, including both natural and non-natural types.

[0278] In some aspects, PUMA pools are designed considering the convolution of PUMA and protein sequences, simplifying sequence deconvolution processes in complex proteomic analyses. This approach is critical in ensuring that each convolved sequence can be uniquely attributed to its constituent PUMA and protein pair.

[0279] In some aspects, the probability transition matrix is employed to calculate modified distance metrics between peptide sequences, accounting for substitutions and indels.

[0280] In some aspects, cycle or order of amino acids in the PUMA pool may have errors, and PUMA sequences may account for this possibility.

[0281] In some aspects, the PUMA generation process is scalable to large proteomes, with the potential to create extensive PUMA pools.

[0282] In some aspects, alternative algorithms are considered for PUMA generation. These may include computational methods capable of handling the increased complexity introduced by large sequence spaces and the inclusion of non-natural amino acids.

[0283] In some aspects, analysis of the resultant sequence data includes deconvolving sequences using a reference proteome.

[0284] In some aspects, analysis of the resultant sequence data does not include using a reference proteome.

[0285] In some aspects, analysis of the resultant sequence employs knowledge of the error correcting code based design to detect or correct errors in PUMA sequence.

[0286] In some aspects, analysis of the resultant sequence data employs statistical methods to differentiate between probable convolutions.

[0287] In some aspects, neural networks and deep learning models are used to interpret complex patterns in data.

[0288] In some aspects, quality scores are assigned to each deconvolution for use in interpreting results.

[0289] In some aspects, analysis of the resultant sequence data employs statistical methods to differentiate between probable convolutions. Neural networks and deep learning models may be used to interpret complex patterns in data. Quality scores may be assigned to each deconvolution for use in interpreting results. Quality scores may be based on the relative hamming distances between the nearest neighbor sequences to the measured sequence.

[0290] In some aspects, quality scores are based on the relative hamming distances between the nearest neighbor sequences to the measured sequence.

[0291] In some aspects, the PUMA pool design is tailored for specific applications, such as targeted proteomics or biomarker discovery, where the focus might be on a subset of proteins with particular characteristics.

[0292] In some aspects, the PUMA pool design accommodates potential sequencing errors, ensuring that the PUMAs remain identifiable even when inaccuracies occur during the sequencing process.

[0293] In some aspects, the data derived from PUMA pools can be used for further research, such as understanding protein interactions, studying disease mechanisms, or developing therapeutic strategies.

[0294] In some aspects, integrating PUMA pool designs with existing protein databases and bioinformatics tools enhances protein identification and characterization processes.

[0295] In some aspects, the PUMA pool data is integrated with other omics datasets, such as genomic, transcriptomic, or metabolomic data, to provide a more comprehensive understanding of biological processes.

[0296] In some aspects, the integration involves correlating PUMA-derived proteomic data with genetic variations or expression profiles observed in genomic and transcriptomic analyses.

[0297] In some aspects, the method of process 2100 comprises capturing an image of the sample on the solid substrate and providing coordinate and / or morphological analyses.

[0298] In some aspects, the method of process 2100 comprises capturing immuno-fluorescence information.

[0299] In some aspects, the method of process 2100 comprises capturing molecular transcription information of the sample on the solid substrate.

[0300] In some aspects, the method of process 2100 comprises correlating image information, transcriptional information, and / or immuno-fluorescence information with reverse translation information of step (k).

[0301] In some aspects, analysis includes assigning cell type, for example using a T-SNE analysis.

[0302] In some aspects, the assay may be used to quantify differences in the amount and / or activity of proteins across spatially-defined locations of a sample

[0303] In some aspects, the assay may be used to quantify differences in the amount and / or activity of proteins between different samples

[0304] In some aspects, one or more of the steps in the herein described processes are automated, including mechanical, fluidic, pneumatic, electrical, optical, and / or control systems.

[0305] In some aspects, one or more of the steps in the herein described processes are automated, in conjunction with automation of operations to ascertain imaging, morphological, transcriptional, and / or protein related information of the sample.PUMA Conjugates

[0306] Disclosed herein, in some embodiments, are peptide unique molecular association (PUMA)- conjugates. A PUMA conjugate may include a PUMA tag. A PUMA conjugate may include a reactive moiety, such as a reactive moiety for coupling to an amino acid. A PUMA conjugate may include an orthogonal reactive moiety, such as a moiety that binds a solid support.

[0307] Disclosed herein, in some embodiments, are peptide unique molecular association (PUMA)- conjugates, comprising: (a) a PUMA tag; (b) a reactive moiety for coupling to an amino acid; and (c) optionally, an orthogonal reactive moiety that binds a solid support.

[0308] In some embodiments, use of the PUMA tag provides metadata relating: a molecule, spatial location, relative spatial location, condition, cell, or sample of origin, for molecule(s) that associate with the PUMA tag. In some embodiments, the PUMA tag comprises repeats of a sequence of amino acids. In some embodiments, the PUMA tag comprises a constant portion and a variable portion. In some embodiments, the PUMA tag may comprise 1, 2, or more portions each with the same or a different function. In some embodiments, the PUMA tag comprises a natural amino acid. In some embodiments, the PUMA tag comprises a non-natural amino acid.

[0309] In some embodiments, a PUMA tag is capped with a terminated sequence. A terminated sequence may include or contain a specific cleavage site. In some embodiments, the PUMA tag is initially capped with a terminated sequence containing a specific cleavage site. This can allow for sequential rather than simultaneous sequencing of the protein analyte and PUMA tag. In some embodiments, the protein analyte is first sequenced while the PUMA tag remains protected. Subsequently, an enzyme specific to the cleavage site may remove the cap, which may allow sequencing of the PUMA tag. This sequential approach may be used to eliminates a need for deconvolution of simultaneous signals from the protein and PUMA tag.

[0310] Protection of the PUMA tag N-terminus may be achieved through various strategies. In some embodiments, the PUMA tag is synthesized with an N-terminal enzymatically cleavable sequence followed by acetylation, preventing PITC coupling during Edman degradation. Suitable enzyme / recognition sequence pairs may include: Factor Xa protease (recognition sequence IEGR), enterokinase (recognition sequence DDDDK), PreScission protease (recognition sequence LEVLFQ / GP), thrombin (recognition sequence LVPR / GS), TEV protease (recognition sequence ENLYFQ / G), and SUMOstar protease (recognition sequence QFGX-SUMO). The recognition sequences may be modified to optimize cleavage efficiency while maintaining protection of the PUMA tag.

[0311] Alternative protection strategies may employ chemical linkers with specific cleavage conditions. Some such strategies include: photocleavable linkers such as 2-nitrobenzyl derivatives that can be cleaved by UV light exposure; pH-sensitive linkers like hydrazones or acetals that cleave under acidic conditions; redox-sensitive linkers containing disulfide bonds that cleave under reducing conditions; and click chemistry-based linkers that can be cleaved through bio-orthogonal reactions. The choice of chemical linker may be optimized based on compatibility with the sequencing chemistry and desired experimental conditions.

[0312] Following protein sequencing or prior to PUMA tag deprotection, a sequenced protein N- terminus may be capped. Following protein sequencing and prior to PUMA tag deprotection, the sequenced protein N-terminus may be capped to prevent interference with subsequent PUMA sequencing. Capping strategies can include acylation with acetic anhydride, succinic anhydride, or other activated esters; PEGylation using activated PEG derivatives; or addition of bulky protecting groups such as Fmoc or Boc. In some embodiments, capping reaction conditions are optimized to help ensure complete protection while maintaining compatibility with the subsequent deprotection and sequencing steps.Protein Reverse Translation

[0313] Disclosed herein, in some embodiments, is protein reverse translation. The protein reverse translation may include identifying amino acid sequences of proteins or peptides that have been immobilized, using a nucleic acid tag.

[0314] Some aspects that may be included when performing protein reverse translation, and that may be useful in conjunction with other embodiments are shown in FIG. 10. N-terminal amino acids may be sequentially removed from a peptide using a conjugate such as a tri-functional molecule in a series of cycles, each of which results in immobilization of one amino acid complex adjacent to the anchor point of its cognate peptide. Multiple cycles may create a lawn of spatially localized complexes holding cycle DNA, such as is depicted in the figure where the large sphere represents a protein localization, and the smaller spheres represent its isolated amino acid localizations. In some embodiments, at this stage, a cycle or amino acid position number is known, but an amino acid identity is not yet determined.Following removal of protecting groups from isolated complexes and transition from an anhydrous to an aqueous environment, amino acid identity may be appended to isolated complexes via recognition by an affinity construct that brings identity information in the form of DNA into proximity of the cycle DNA. ‘Identity’ and ‘cycle’ DNA may be ligated in a high-fidelity reaction. Extension-ligation of regional DNA into a long construct that reflects the original peptide information, as shown in the lowest geometry panel, may be analyzed using next generation sequencing.

[0315] Some embodiments include determining identity information of an amino acid residue of a peptide. Some embodiments include positional information of an amino acid residue of a peptide. The peptide may be coupled to a solid support. Some embodiments include providing the peptide. Some embodiments include coupling the peptide to the solid support. The peptide may be coupled to the solid support such that a N-terminal amino acid residue of the peptide is not directly coupled to the solid support. The N-terminal amino acid (NTAA) residue may be exposed to reaction conditions. Some embodiments include providing a chemically-reactive conjugate (CRC). The CRC may include a cycle tag. The cycle tag may include a cycle nucleic acid, which may be associated with a cycle number. The CRC may include a reactive moiety. The reactive moiety bind the NTAA. The reactive moiety cleave the NTAA. The CRC may include an immobilizing moiety, which may be for immobilization to the solid support. Some embodiments include coupling the N-terminal amino acid with a first reactive moiety (e.g. ITC) or a first reactive moiety conjugate (e.g. an ITC conjugate). Some embodiments include providing a CRC comprising a second reactive moiety that binds the first reactive moiety or first reactive moiety conjugate. The second reactive moiety may react with the first reactive moiety or first reactive moiety conjugate. Some embodiments include contacting the peptide with the CRC. Some embodiments include coupling the CRC to the NTAA of the peptide to form a conjugate complex. Some embodiments include coupling the CRC to the first reactive moiety or first reactive moiety conjugate coupled to the NTAA of the peptide to form the conjugate complex. Some embodiments include immobilizing the conjugate complex to the solid support, e.g. via the immobilizing moiety. Some embodiments include cleaving and thereby separating the N-terminal amino acid residue from the peptide. Some embodiments include exposing the next amino acid residue as an NTAA residue on the cleaved peptide. Some embodiments include providing an immobilized amino acid complex. The immobilized amino acid complex may include the cleaved and separated N-terminal amino acid residue. Some embodiments include contacting the immobilized amino acid complex with a binding agent. The binding agent may include a binding moiety, which may preferentially bind to the immobilized amino acid complex. The binding agent may include a recode tag. The recode tag may include a recode nucleic acid. The recode nucleic acid may correspond with the binding agent. Some embodiments include forming an affinity complex. The affinity complex may include an immobilized amino acid complex. The affinity complex may include a binding agent. Some embodiments include bringing a cycle tag into proximity with a recode tag. Some embodiments include bringing a cycle tag into proximity with a recode tag within each formed affinity complex. Some embodiments include transferring informationof a recode tag. Some embodiments include generating a recode block. Some embodiments include transferring information of the nucleic acid recode tag associated with the first binding agent and the cycle tag of the first immobilized conjugate complex to generate a first recode block. Some embodiments include obtaining sequence information for the recode block. Some embodiments include generating a memory oligonucleotide, which may comprise multiple recode blocks. Some embodiments include determining identity or positional information of an amino acid residue of the peptide. Some embodiments include performing a method on multiple polypeptides. The multiple polypeptides may be immobilized, such as together on a solid support. Some embodiments include performing a method on multiple polypeptides together. Some embodiments include performing a method on multiple polypeptides simultaneously.

[0316] Disclosed herein, in some embodiments, are methods for determining identity and positional information of an amino acid residue of a peptide coupled to a solid support, the method comprising: (a) coupling the peptide to the solid support such that a N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions; (b) performing (bl) or (b2): (bl) providing a chemically-reactive conjugate comprising (x) a cycle tag comprising a cycle nucleic acid associated with a cycle number; (y) a reactive moiety for binding and cleaving the N- terminal amino acid residue of the peptide; and (z) a immobilizing moiety for immobilization to the solid support; or (b2) coupling the N-terminal amino acid with a first reactive moiety (e.g. an ITC conjugate), and providing a chemically-reactive conjugate comprising (x) a cycle tag comprising a cycle nucleic acid associated with a cycle number; (y) a second reactive moiety that binds or reacts with the first reactive moiety (e.g. that binds or reacts with the ITC conjugate); and (z) a immobilizing moiety for immobilization to the solid support; (c) contacting the peptide with the chemically-reactive conjugate of (bl) thereby coupling the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex, or contacting the peptide with the chemically-reactive conjugate of (b2) thereby coupling the chemically-reactive conjugate to the first reactive moiety coupled to the N-terminal amino acid of the peptide to form the conjugate complex; (d) immobilizing the conjugate complex to the solid support via the immobilizing moiety; (e) cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby exposing the next amino acid residue as the N-terminal amino acid residue on the cleaved peptide, and thereby providing an immobilized amino acid complex, the immobilized amino acid complex comprising the cleaved and separated N-terminal amino acid residue; (f) contacting the immobilized amino acid complex with a binding agent, the binding agent comprising: a binding moiety for preferentially binding to the immobilized amino acid complex; and a recode tag comprising a recode nucleic acid corresponding with the binding agent, thereby forming an affinity complex, the affinity complex comprising an immobilized amino acid complex and a binding agent and thereby bringing a cycle tag into proximity with a recode tag within each formed affinity complex; (g) transferring the information of the nucleic acid recode tag associated with the first binding agent and the cycle tag of the first immobilizedconjugate complex to generate a first recode block; (h) obtaining sequence information for the recode block; and (i) based on the obtained sequence information, determining identity and positional information of an amino acid residue of the peptide. The process may be performed on multiple polypeptides, such as multiple immobilized polypeptides together or simultaneously.

[0317] Some embodiments relating to reverse translation of immobilized protein information are further described in US provisional no. 63 / 620,076, and PCT application nos. PCT / US23 / 70077 and PCT / US23 / 72498, which are all herein incorporated by reference in their entirety. Performing reverse translation protein sequencing may comprise performing some aspects of a method for determining identity and positional information of an amino acid residue of a peptide such as a peptide coupled to a solid support.

[0318] Determining identity or positional information of an amino acid residue may include reverse translation protein sequencing or vice versa. Reverse translation protein sequencing may include recoding an amino acid sequence to a nucleic acid sequence, and then decoding or sequencing the nucleic acid sequence. Recoding may include the use of reagents or methods herein, or described in WO2024015875 or W02024040236. Reverse translation protein sequencing may include providing a peptide to a solid support (e.g. such that a N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions). Reverse translation protein sequencing may include providing a chemically reactive conjugate (CRC). A CRC may include a cycle tag comprising a cycle nucleic acid associated with a cycle number, a reactive moiety for binding an N- terminal amino acid residue of the peptide, and an immobilizing moiety for immobilization to the solid support. Reverse translation protein sequencing may include contacting a peptide with a CRC. Reverse translation protein sequencing may include forming a conjugate complex (e.g. upon contact of a peptide with a CRC). Reverse translation protein sequencing may include immobilizing a conjugate complex to a solid support. Reverse translation protein sequencing may include cleaving an N-terminal peptide. Reverse translation protein sequencing may include separating an N-terminal amino acid residue from a peptide. Reverse translation protein sequencing may include providing an immobilized amino acid complex. An immobilized amino acid complex may include cleaved or separated N-terminal amino acid residue. Reverse translation protein sequencing may include contacting an immobilized amino acid complex with a binding agent. A binding agent may include: a binding moiety for preferentially binding to the immobilized amino acid complex, and a recode tag comprising a recode nucleic acid corresponding with the binding agent. Reverse translation protein sequencing may include forming an affinity complex. An affinity complex may include an immobilized amino acid complex and a binding agent. Reverse translation protein sequencing may include bringing a cycle tag into proximity with a recode tag within an affinity complex. Reverse translation protein sequencing may include transferring information of a recode nucleic acid to a cycle nucleic acid (e.g. of an immobilized conjugate complex). Reverse translation protein sequencing may include generating a recode block. Reverse translation protein sequencing may include obtaining sequence information of a recode block. Reverse translationprotein sequencing may include formation of a memory oligo. Forming a memory oligo may include joining two or more recode blocks together. Forming a memory oligo may include combining sequences or sequence information of two or more recode blocks. Reverse translation protein sequencing may include obtaining sequence information of a memory oligo. Reverse translation protein sequencing may include determining identity and positional information of an amino acid residue of a peptide based on obtained sequence information (e.g. obtained information of one or more recode blocks, or obtained information of a memory oligo).Automation and Devices

[0319] It is recognized that automation may increase efficiency, improve reproducibility, and reduce cost. In some aspects automated systems including control and software are employed to accomplish operations of the assay. Thus, integrated automated submodules may be envisioned to include imaging instrumentation, fluidic instrumentation, and reverse translation instrumentation, FIG. 9 schematically illustrates multi-component integration that supports multi-omic analyses.

[0320] Note that additional or augmented solid support surfaces, including microfluid devices, may be integrated to support multi-omic analyses, for example, a surface that supports transcriptomics may be fitted with TFS beads, such that both types of analysis may proceed from the same substrate. Several steps are synergistic and may be accomplished in parallel for the multi-analyte methods.DEFINITIONS

[0321] Note that as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to “an oligo” refers to one or more oligos, and so forth. Additionally, it is to be understood that terms such as "left," "right," "top," "bottom," "front," "rear," "side," "height," "length," "width," "upper," "lower," "interior," "exterior," "inner," "outer" that may be used herein merely describe points of reference and do not necessarily limit embodiments of the present disclosure to any particular orientation or configuration. Furthermore, terms such as "first," "second," "third," etc., merely identify one of a number of portions, components, steps, operations, functions, and / or points of reference as disclosed herein, and likewise do not necessarily limit embodiments of the present disclosure to any particular configuration or orientation.

[0322] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. All publications mentioned herein are incorporated by reference for the purpose of describing and disclosing devices, compounds, compositions, and methods that may be used in connection with the presently described disclosure.

[0323] Where a range of values is provided, it is understood that each intervening value, between the upper and lower limit of that range and any other stated or intervening value in that stated range isencompassed within the disclosure. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges, and are also encompassed within the disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure.

[0324] As used in the present disclosure, the term “amino acid” and notation “AA” refer to natural d-, 1-, non-natural, and post-translationally modified amino acids. An “N-terminal amino acid” refers to an amino acid that has a free amine group, and is linked to only one other amino acid of the peptide through an amide bond. Similarly, a “C-terminal amino acid” refers to an amino acid that has a free carboxyl group, and is linked to only one other amino acid of the peptide through an amide bond.

[0325] The term “AA tag” refers to a nucleic acid molecule of any length, but typically in the range 5- 20 bases, which contains a sequence that is defined to represent a particular amino acid or class of amino acids that share structural or functional similarity. If recoding a polymer that does not comprise amino acids, then the AA tag sequence may be defined to represent a particular monomer or class of monomers that share structural or functional similarity. It may also refer to any construct that enables a method of subsequent identification of the cycle information, such as a mass tag.

[0326] The terms “analyze” and “analyzing” refer to assigning a sequence, and / or quantification, and / or identity to the macromolecule, or a part of the macromolecule analyte.

[0327] The term “assembly oligo” (e.g., an assembly Oligo) refers to a nucleic acid capable of hybridizing to a memory oligo tethered to a solid support and / or hydrogel. Assembly oligos may be utilized to facilitate ligation assembly of a complementary DNA strand to a memory oligo that is tethered to the hydrogel surface and or solid support as a template. Ligation assembly of a complementary strand avoids the need for polymerase extension through tethered nucleic acids to create a solution phase nucleic acid representative of the analyte sequence. An assembly oligo comprises a sequence complementary to a cycle tag sequence and a sequence complementary to an amino acid sequence.

[0328] The term “binding agent” refers to an entity comprised of a binding moiety joined with a recode tag. The binding moiety and recode tag may be joined by a linker.

[0329] The term “binding moiety” refers to a molecule or macromolecule that recognizes and binds with a target analyte or a feature of the target analyte. Exemplary binding moieties include: antibodies, F(ab’)2, Fab, and scFv regions, nanobodies, DNA aptamers, RNA aptamers, modified aptamers, photoactive or non-photoactive cage compounds, oligo peptide permease (Opp), amino-acyl t-RNA synthetase (aaRS), periplasmic binding proteins (PBP), dipeptide permease (Dpp), proton dependent oligopeptide transporters (POT), modified aminopeptidases, modified amino acyl tRNA synthetases, modified anticalins, modified ClpS, Eectin, or clathrates. A binding moiety may form a covalent association or non-covalent association with target analytes, which include immobilized conjugate complexes, such as an immobilized PTC-AA-cycle tag-conjugate complex. The binding moiety may exhibit preferential binding to one conjugate complex over another one depending on the amino acid ofthe complex. The binding moiety may bind preferentially to classes of amino acids that are structurally or functionally similar within the conjugate complex.

[0330] In addition to caged drugs and bioactive small molecules, amino acids and derivatized amino acids offer a number of possibilities for caging. For example, amines, carboxylates, and amino acid side chains offer a number of easily caged functional groups. More particularly, caged serine, threonine, tyrosine, cysteine, methionine, aspartate, glutamate, and lysine have all been reported; see Pirrung et al., Synthesis of photodeprotectable serine derivatives - caged serine, Bioorg. Med. Chem. Lett. 2, 1489-1492 (1992); Tatsu et al., Solid-phase synthesis of caged peptides using tyrosine modified with a photocleavable protecting group, Biochem. Biophys. Res. Comm. 227, 688-693 (1996); Gee, K.R., Carpenter, B. K., and Hess, G.P., Synthesis, photochemistry, and biological characterization of photolabile protecting groups for carboxylic acids and neurotransmitters, Met. Enz. 291, 30-50 (1998); Tatsu et al., Synthesis of caged peptides using caged lysine: Application to the synthesis of caged AIP, a highly specific inhibitor of calmodulin-dependent protein kinase II, Bioorg. Med. Chem. Lett. 9, 1093- 1096 (1999); Okuno, T., Hirota, S., and Yamauchi, 01., Folding character of cytochrome c studied by onitrobenzyl modification of methionine 65 and subsequent ultraviolet light irradiation, Biochem. 39, 7538-7545 (2000).

[0331] The terms “biochip” and “microarray” refer to consumable devices that support fluidic operations and further support a recode workflow. In some embodiments, these could include a flowcell used directly by an NGS sequencing instrument in a DNA sequencing process.

[0332] The term “biologically or synthetically-derived sample” refers to a sample of macromolecules that has its origins from a biological process, such as a cell lysate solution, or has origins from a sample created using synthetic biology techniques, or a sample of macromolecules created using purely chemical synthesis, for example a solution of synthetic peptides, synthetic nucleic acids, or chemically- synthesized polymers.

[0333] The term “C / AA tag” refers to a nucleic acid molecule of any length, but typically in the range 5-40 bases, having a sequence that represents a particular amino acid and cycle of a single-molecule sequencing workflow. The length of a C / AA tag may differ for different cycles of the workflow. The C / AA tag may optionally comprise additional nucleic acid sequences that direct assembly of memory oligos in subsequent steps, such as unifying assembly sequences which facilitate recode block assembly irrespective of the order of assembly. In certain examples, a C / AA tag may optionally comprise a restriction endonuclease sequence, and / or a sequence facilitating amplification of recode blocks, or any sequence functionality disclosed herein. The term, “C / AA tag” may also refer to any construct that enables a method of subsequent identification of the cycle information, such as a mass tag, or a chemically reactive moiety. It provides identifying amino acid (or monomer subunit) information for its associated binding agent. It may uniquely identify one amino acid or may identify a class of amino acids with structural and / or functional similarity. A C / AA tag may provide a probabilistic estimate as to the identity of the amino acid, and thereby provide sufficient information for analysis.

[0334] The term “chemically-reactive conjugate” refers to a conjugate comprising (a) a reactive moiety(ies) that can bind and cleave a terminal amino acid, (b) a reactive moiety that allows immobilization to a solid support, and (c) a cycle tag with identifying information regarding the workflow cycle. In certain embodiments, various moieties may be replaced with reactive moieties for attaching said groups.

[0335] The term “codespace” refers to the universe of codes that are associated with cycle tags and AA tags and are used to represent workflow cycle and monomer identity information, respectively. Codespace is defined by a set of rules that provide practical separation distance between codes and improve fidelity and accuracy while reading information. For example, Hamming distance theory, or other modern digital code space theories (e.g., Lee, Levenshtein-Tenengolts, Reed-Solomon, or others) may be applied to assign codes and enable error detection and error correction capability and account for: 1) NGS sequencing errors during analysis, 2) errors in oligonucleotide synthesis, 3) errors in reagents used in the recoding process, 3) errors that occur during assembly of recode blocks, 4) errors that occur during assembly of memory oligos, or combinations of errors that may occur during any step in the determination of protein sequence and protein abundance by recoding amino acid polymers into DNA polymers and analyzing.

[0336] The term “cognate binding agent” refers to a binding agent that was designed to, and that binds with high relative affinity to, a cognate target analyte or a feature or portion of the cognate target analyte. This is contrasted with a “non-cognate binding agent”, that was not designed to bind to, and thus interacts with low relative affinity to, a non-cognate target analyte or a feature or portion of the noncognate target analyte, such that the non-cognate binding agent does not effectively transfer recode tag information to the recode block under conditions appropriate for recode block assembly by cognate binding agents.

[0337] The terms “conjugate complex” and “immobilized conjugate complex” refer to a chemically- reactive conjugate having been joined optionally as appropriate within the context to: an amino acid (e.g., a monomer of the macromolecular analyte), a peptide, a linker, a solid support, and / or a cycle tag.

[0338] The term "complementary" refers to Watson-Crick base pairing between nucleotides and specifically refers to nucleotides hydrogen bonded to one another with thymine or uracil residues linked to adenine residues by two hydrogen bonds and cytosine and guanine residues linked by three hydrogen bonds. In general, a nucleic acid includes a nucleotide sequence described as having a "percent complementarity" or “percent homology” to a specified second nucleotide sequence. For example, a nucleotide sequence may have 80%, 90%, or 100% complementarity to a specified second nucleotide sequence, indicating that 8 of 10, 9 of 10 or 10 of 10 nucleotides of a sequence are complementary to the specified second nucleotide sequence.

[0339] The term "Convolved Sequences" refers to the composite sequences generated from the combination of protein and PUMA sequences. In this context, convolution implies the merging oroverlapping of amino acid sequences from both proteins and PUMAs, resulting in a sequence where the original source (protein or PUMA) of each amino acid is not distinguishable.

[0340] The term “cycle tag” (e.g., “cycleTag”) refers to a nucleic acid molecule of any length, but typically in the range 5-20 bases, having a sequence that is defined to represent a particular cycle of the recode workflow. The length of a cycle tag may differ for different cycles of the workflow. The cycle tag may optionally comprise additional nucleic acid sequences that direct assembly of memory oligos in subsequent steps, such as universal assembly sequences which facilitate recode block assembly irrespective of the order of assembly. In certain examples, a cycle tag may optionally comprise a restriction endonuclease sequence. The term, “cycle tag” may also refer to any construct that enables a method of subsequent identification of the cycle information, such as a mass tag.

[0341] The term “deprotecting” refers to removing protecting moieties that preserve the integrity of a functional group during exposure to conditions and potential reactants that may otherwise react to alter the functional group. Exemplary protecting agents for nucleic acids include: FMOC, acetyl (Ac), benzoyl (Bz), dimethylformamidine (DMFA), and phenoxyacetyl (PAC). See, Radhakrishnan P. Iyer, Current Protocols in Nucleic Acid Chemistry.

[0342] The term “Hamming Distance” describes the distance between two strings or vectors of equal length. It is the number of positions at which the corresponding symbols are different. It measures the minimum number of substitutions required to transform one string into another. In this context, this represents the number of amino acid substitutions in a given sequence needed to transform into another sequence in a given sequence space or pool.

[0343] The terms "homology" or "identity" or "similarity" refer to sequence similarity between two peptides or between two nucleic acid molecules.

[0344] The term “hydrogel” refers to synthetic polymers, natural polymers, and / or hybrid polymers. Exemplary monomers that may form the hydrogel include one or more: acrylamide, acrylate, vinyl pyridine, dihydroxy methacrylates, other methacrylates, HEMA, PHEMA, PVA, HPMC, PLGA, PEG, etc., in linear, branched, and crosslinked configurations, block co-polymers configurations, or other configurations conducive to sequencing macromolecules. See, Faisal Raza, Hajra Zafar, Ying Zhu, Yuan Ren, Aftab -Ullah, Asif Ullah Khan, Xinyi He, Han, Md Aquib, Kofi Oti Boakye-Yiadom and Liang Ge, A Review on Recent Advances in Stabilizing Peptides / Proteins upon Fabrication in Hydrogels from Biodegradable Polymers, Pharmaceutics 2018, 10, 16. A hydrogel may be associated with a solid support through covalent or non-covalent interactions. The hydrogel may further comprise orthogonal conjugation chemistry modalities to support the recode workflow.

[0345] The terms “ith,” “(i-l)th”, etc., refer to an arbitrary position in the macromolecular analyte and it’ s nearest neighbor.

[0346] The term “initiation oligo” refers to any nucleic acid that when immobilized is capable of initiating assembly of co-localized nucleic acids. Initiation oligos may comprise a sequence facilitating assembly of recode blocks, determination of a shared relative spatial location, amplification of nucleicacids, and / or ULI, U-AS, U-HS, hyb tags, ligation complements, or combinations of these or other sequences described herein for the analysis of peptides. It is of any length, but typically in the range 15- 80 bases and the length of one hyb tag may differ from that of another. In certain examples, an initiation oligo may optionally comprise a restriction endonuclease sequence. Initiation oligos may be linked via a linker or may comprise a linker. It is also recognized that the initiation oligos may be another molecular format that is not a nucleic acid, and that information may be joined with an initiation oligo via a chemical reaction.

[0347] The term “ITC-Conjugate” refers to a molecule having an amine-reactive group (RG), including but not limited to isothiocyanate, alkyl isothiocyanate, aryl isothiocyanate, substituted aryl isothiocyanate, isoselenocyanate, alkyl selenocyanate, aryl selenocyanate, substituted aryl selenocyanate and a functional group capable of being joined to a CRC, or other complementary reactive element. The ITC-conjugate may include substituents that influence the reactivity or physicochemical characteristics of the molecule, such as fluoro, nitro, halo, carboxyl, cyano, pyridyl, ether, thioether, amide, carbonate, carbamate, tertiary amino, quaternary amino groups and combinations thereof. The ITC conjugate may possess heterocyclic structures including imidazole, pyrazole, pyrazines, thiophene, furan, pyrrole, pyran, pyrimidine, oxazole, thiazole. It is recognized that in the case of C-terminal sequencing the RG group would be reactive to the carboxyl terminus.

[0348] The term “ligation oligo” (e.g., “ligationOligo”) refers to a nucleic acid that becomes ligated to a cycle tag of an immobilized conjugate complex when appropriately directed by a cognate binding agent via hybridization to the recode tag of the cognate binding agent. Ligation oligos may, in certain embodiments, hold information related to amino acid and workflow cycle assembly, and are complementary to the recode tag of a cognate binding agent. It is also recognized that the ligation oligo may be another molecular format that is not a nucleic acid, and that recodes amino acid and workflow cycle information that can be joined with a cycle tag via a chemical reaction. In certain embodiments, ligation oligos may optionally comprise a sequence facilitating ligation, extension: ligation, or chemical ligation of a recode block to another other recode block irrespective of the order of assembly. For example, by including a 3’ and / or 5’ universal assembly sequence on a plurality of recode blocks such that at least two recode blocks share the same universal assembly sequence, assembly of such recode blocks into a memory oligo, in any given order, is enabled.

[0349] The term “linker” or “spacer” refers to a molecule used to join two or more molecules. The composition of the molecule may be a polymer, a monomer or combination of both. A linker may further comprise reactive elements that promote covalent and / or non-covalent conjugation between molecules. It may be of any suitable length. Exemplary linkers include those used to join a binding agent to a recode tag, or a protein to a solid support, or a PUMA tag to a solid support, or a cycle tag to other elements of a conjugate complex, e.g. a molecule having a NHS-ester at one end and an azide at the other end of a PEG molecule, or a molecule having a biotin at one end and an maleimide moiety at the other end of a nucleic acid.

[0350] The term “linking oligo” (e.g., linkingOligo”) refers to a nucleic acid capable of promoting ligation between a recode block associated with a given workflow cycle and a second recode block associated with any other workflow cycle of the recoding process. Linking oligos are useful to complete the assembly of a memory oligo, because they can substitute for errors, e.g., in upstream processes that resulted incomplete or unexpected recode block sequence for one or more workflow cycles, no recode block assembly for one or more workflow cycles, or steric effects that prevent interaction between and assembly of recode blocks. Linking oligos may optionally comprise a sequence complementary to the cycle tag sequence of one workflow cycle and the cycle tag sequence of any other workflow cycle. Ligation of recode blocks via linking oligos may create a lack of information related to the recode block that was skipped in the assembly of the memory oligo. In this case it is recognized that the memory oligo may still be valuable for analysis of macromolecular information, since information may be inferred during analysis that an unknown (or multiple unknown) monomers separate the positions of known monomers, and mapping to references sequence allows macromolecule sequence and identity information. In certain embodiments, linking oligos may optionally comprise a sequence for promoting ligation between a recode block associated with a workflow cycle and a second recode block associate with another workflow cycle of the recoding process. For example, such ligation may be promoted via complementarity between universal assembly sequences of the cycle tag and / or the recode tag.

[0351] The term “location linker” refers to any molecule configured to attach a peptide to a solid support, and further configured to bind to a nucleic acid. In some examples, a location linker refers to a molecule with 3 or more functional elements that facilitate the attachment of a peptide, a nucleic acid, and a solid support. In some examples, the nucleic acid can be a UMI that carries code information related to a location of isolation for partitioned immobilized PTC-conjugates.

[0352] The term “location oligo” (e.g., “locationOligo”) refers to a nucleic acid of any suitable length, but typically in the range 10-40 bases, that contains a sequence that represents the x, y, z coordinates of an immobilized macromolecular analyte and is held in proximity to a macromolecule via a location linker. Location oligos are useful to transfer location information to spatially-adjacent immobilized recode blocks.

[0353] The terms “macromolecule” and “macromolecular polymer” refer to a high molecular weight molecule composed of subunits. Examples of macromolecules include, but are not limited to, protein complexes such as a photosynthetic reaction center antenna complex, multi-subunit proteins such as a photosynthetic reaction center or a pore protein, single subunit proteins such as cytochrome-c, protein fragments, peptides, polypeptides, nucleic acids, carbohydrates, and polymers such as urethane or acrylamide. “Macromolecule” also describes natural and synthetic combinations of two or more macromolecular types, such as a peptide covalently bound to a nucleic acid, or a lectin bound to a carbohydrate though electrostatic, van der waals forces, or any non-covalent forces.

[0354] The term “memory oligo” (e.g., “memoryOligo”) refers to a construct that comprises location information, monomer relative positional information, and / or monomer identity information. It istypically assembled by aggregating the information of recode blocks. Typically, a memory oligo comprises information for one associated macromolecular analyte. However, it is recognized that there are embodiments where a memory oligo comprises identifying information for one or more macromolecular analytes. Optionally, a memory oligo may further comprise: sample indexes, UMIs, universal priming sites, linkers, and other identifiers of macromolecule provenance. The length of a memory oligo will typically be between 25 and 25,000 base pairs. When perfectly assembled, the length of the memory oligo equals the sum of the lengths of provenance identifiers plus the lengths of cycle tag and AA tag sequences multiplied by the number of workflow cycles. It is recognized that cycle tag lengths may be different for different workflow cycles. Note that imperfect assembly of a recode block may produce a memory oligo with shorter or longer lengths than the perfectly assembled memory oligos and that are valuable for analysis of the macromolecule, since cycle and amino acid (e.g., monomer) information is transferred to adjacent registers of the memory oligo. It is further recognized that sequential assembly of recode block information into a memory oligo is not required to provide a memory oligo for analysis that is useful for macromolecule analyte analysis.

[0355] The term “metadata conjugate” refers to a conjugate comprising (a) a binding moiety(ies) that can bind to an attribute of the peptide analyte (b) a reactive moiety that allows immobilization to a solid support, and (c) a metadata tag with identifying information regarding the cognate attribute of the immobilized peptide.

[0356] The term “metadata tag” may refer to a nucleic acid molecule of any length, but typically in the range 5-40 bases, having a sequence that is a priori defined to represent a particular attribute of a peptide. The length of a metadata tag may differ for different attributes. The metadata tag may optionally comprise additional nucleic acid sequences that direct assembly of memory oligos in subsequent steps, such as unifying assembly sequences which facilitate recode block assembly irrespective of the order of assembly. In certain examples, a metadata tag may optionally comprise a restriction endonuclease sequence. The composition of a metadata tag may be DNA, RNA, LNA, PNA, XNA, TNA, BNA, NA with both backbone and base modifications, chemically protected nucleic acids, or a combination thereof. The term metadata tag can also be used to describe any entity that provides metadata information or functionality. For example, a metadata tag may be a peptide, a biotinylation moiety, a small molecule, a reactive small molecule, a click reagent, or other.

[0357] The term “n” refers to the length of the target macromolecular analyte, or the workflow cycle number. It also refers to terminal subunit of the macromolecular analyte, e.g., nth subunit. Accordingly, the next subunit is denoted as n-1, then the n-2, and so on down the length of the peptide. Theses labels can be assigned starting from the N-terminal or the C-terminal end of a macromolecule.

[0358] The terms “n-1”, n-2”, etc., refer to a cycle prior to the last cycle and, so on. It can also refer to a nearest and a next-nearest subunit molecule to the terminal subunit of a macromolecular analyte.

[0359] The term “polynucleic acid” or “polynucleotide” refers to a polymer of deoxyribonucleotides linked by 3'-5' phosphodiester bonds. This also includes polymers with nucleotide analogs and non-natural nucleotides such as Iso-G and Iso-C. This also includes nucleotides linked by thiophosphate bonds or peptidyl bonds such as in PNA. This also covers RNA and polymers with a modified ribose moiety or moieties, such as LNA, XNA, or BNA.

[0360] The terms “nucleic acid sequencing,” “NGS,” or “next generation sequencing” refer to high- throughput methods to determine the sequence of a nucleic acid polymer. These methods are exemplified by commercially available products from Illumina, Pacific Biosciences, and Oxford Nanopore.

[0361] The term “peptide” or “polypeptide” refers to a chain of two (2) or more amino acids, and no discrimination in terms of length is implied by the terms: peptide, polypeptide, or protein. Similarly, no discrimination or restriction is implied in terms of 1-, d-, non-natural, or post-translationally modified amino acids monomers that comprise the peptide.

[0362] The term “peptide unique molecular association tag” (PUMA tag), refers to a peptide of any length, but typically 4 to 16 amino acids in length. The peptide unique molecular association tag sequence may be any sequence of natural, derivatized, synthetic, or non-natural amino acids, but typically abides by the described in the written description. It may be used to provide metadata relating to molecule, spatial location, relative spatial location, condition, cell, or sample of origin, or any useful identifier for the molecule(s) associated with the PUMA molecule(s). Reverse translating a PUMA tag may provide a probabilistic estimate as to the provenance of molecule, cell, sample, condition, spatial information, or any useful identifier, and thereby provide sufficient information for analysis. The length of a PUMA tag may differ for different analyte molecules. The PUMA tag may optionally comprise a moiety(s) that facilitate joining with a solid support or TFS, such as an azide, alkyne, amine, carboxyl, or any of the reactive groups described in PCT / US23 / 70077. It may further comprise fluorophores or haptens, which may aid in development of assays. The PUMA tag may optionally comprise a protease digestion target sequence, a linker, and / or any sequence functionality disclosed herein. The PUMA tag may comprise a constant portion and a variable portion. For example, a constant portion shared by all tagged molecules of a given cell, and a variable portion that is unique to each tagged molecule of a cell. The sequence of amino acids of an individual PUMA tag sequence may be repeated one or more times. For example, the PUMA tag may comprise T-E-C-H-T-E-C-H-linker-azide (SEQ ID NO. 263). A PUMA tag may comprise 1 , 2, or more portions each with the same or a different function.

[0363] The term “peptide unique molecular association tag conjugate” (PUMA-conjugate), refers to a conjugate comprising (a) a PUMA tag; and (b) a reactive moiety capable of coupling to a protein, and optionally (c) an orthogonal reactive moiety capable of joining a PUMA-tagged protein conjugate to a solid support.

[0364] The term “PITC-conjugate” refers to a chemically-reactive conjugate that has not been reacted with an amino acid or a solid support. It is recognized that the qualifier “PITC” is representative terminology to describe any number of molecules (or sets of molecules) that can function similarly to bind to N-terminal or C-terminal amino acids and cleave the terminal subunit.

[0365] The terms conjugate complex, “PTC-conjugate,” and “PTC-AA-cycle tag-conjugate complex”, refer to a chemically-reactive conjugate that has been reacted with an amino acid, but not necessarily been immobilized to a solid support. It is recognized that the qualifier “PTC” is representative terminology to describe any number of alternative molecules (or sets of molecules) that can function similarly to bind to N-terminal or C-terminal amino acids and cleave the terminal subunit. The terms “immobilized conjugate complex,” “immobilized PTC-conjugate,” and “immobilized PTC-AA-cycle tag-conjugate complex” refer to a chemically-reactive conjugate that has been reacted with an amino acid been immobilized to a solid support. It is recognized that the qualifier “PTC” is representative terminology to describe any number of alternative molecules (or sets of molecules) that can function similarly to bind to N-terminal or C-terminal amino acids and cleave the terminal subunit.

[0366] The term “post-translational modification” refers to any modification of an 1-, d-, or non-natural amino acid, either biologically or synthetically. The modifications can occur at the terminal amine, the terminal carboxyl, or any reactive moiety of a peptide. Examples include, but are not limited to, phosphorylation, glycosylation, glycanation, methylation, acetylation, ubiquitination, carboxylation, hydroxylation, biotinylation, pegylation, and succinylation. Further information regarding post- translational modifications may be found in, DOI: 10.1021 / acs.biochem.7b00861. Biochemistry 2018, 57, 177-185, which is herein incorporated by reference in its entirety.

[0367] A "Protein Library" may generally refer to a curated collection or database of proteoforms, which are the various forms of proteins present in a specific cell, tissue, or organism, or other collection of peptide or proteins in a given experiment or set of experiments. This library serves as a reference against which Peptide Unique Molecular Association (PUMA) sequences are designed and analyzed. The library encompasses some expected protein sequences, optionally including their variants and modifications, and may be used to ensure the specificity and uniqueness of PUMA sequences relative to existing proteomic data. A "Protein Library" in the broader sense encompasses not only the various forms of proteins present in specific cells, tissues, or organisms but also any set of amino acid sequences utilized in a given experiment or set of experiments or other sequencing application. This library can include, but is not limited to, peptides derived from known proteins, synthetic peptide sequences, fragmented proteins, and post-translationally modified proteins. Its scope extends beyond full-length proteins to incorporate any sequence of amino acids of interest, whether naturally occurring or synthetically designed. A “Protein Library” may refer to a predefined set of peptides or proteins, each represented as a sequence of amino acids.

[0368] A "PUMA Pool" is a designed or physically realized group of PUMA sequences. It represents a collection of peptide sequences. The pool may vary in size and composition, comprising a number of PUMAs with unique sequence characteristics. These PUMAs are used for a variety of applications including, but not limited to, protein tagging, tracking, and identification in complex biological samples.

[0369] The term “PUMA Library” refers to a set of PUMAs to be generated, with the requirement that each PUMA maintains a specific Hamming distance from the peptides or proteins in the protein library and from each other.

[0370] The term “recode block” (e.g., “recodeBlock”) refers a construct created by interaction between a cycle tag of an immobilized conjugate complex and the recode tag of a cognate binding agent. Typically, a recode block is a chimeric nucleic acid molecule that contains the information relating the workflow cycle and the amino acid, or class of amino acid, composition that comprises the conjugate complex. Further, the recode block holds information to direct assembly of a memory oligo, and / or amplify the recode block. A recode block may be formed by utilizing an extension-ligation method to transfer information from the recode tag to the recode block, or via a ligation reaction under appropriate conditions in the presence of ligase and ligation oligo, or otherwise transferring information from the cycle tag to a separate entity. A recode block may be formed by utilizing an extension-ligation method to transfer information from the cycle tag to the recode block, or via a ligation reaction under appropriate conditions in the presence of ligase and ligation oligo, or otherwise transferring information from the cycle tag to a separate tag. A recode block may be formed by otherwise transferring information from the cycle tag and recode tag to a separate entity. A recode block may be formed by transferring information from the cycle tag into a new sequence of nucleic acids that is not the cycle tag or its complement, but otherwise indicates the cycle of the workflow. The format of a recode block is not necessarily a nucleic acid. It may also take the form of mass tags that could be used to assign identity for cycle and amino acids of the cognate conjugate complex, or other modalities that represent the information of the immobilized conjugate complex, and are amenable to group that information for analysis.

[0371] The term “recode tag” (e.g., “recodeTag”) refers to a nucleic acid molecule of any length, but typically in the range 15-60 bases, having a sequence comprised of an ith cycle tag complement, an AA tag complement, and an (i-l)th cycle tag complement. It provides identifying amino acid (or monomer subunit) information for its associated binding agent. It may uniquely identify one amino acid or may identify a class of amino acids with structural and / or functional similarity. A recode tag may provide a probabilistic estimate as to the identity of the amino acid component of an immobilized PTC-AA-cycle tag-conjugate complex, and thereby provide sufficient information for analysis. In certain embodiments, a recode tag may optionally comprise the ith cycle tag complement, an AA tag complement, and / or a universal assembly sequence or a complement of the universal assembly sequence that aids in the assembly of a memory oligo. In certain embodiments, a recode tag may optionally comprise a universal assembly sequence at both the 3’ and 5’ ends to facilitate memory oligo assembly without regard to the order of assembly of constituent recode blocks. In further embodiments, a recode tag may comprise a sequence facilitating amplification of recode blocks. In some embodiments, a recode tag may not comprise the ith cycle tag complement, but a site for attaching said cycle tag or its complement, orinformation identifying to the cycle of the workflow that is not identical to said cycle tag or its complement.

[0372] The terms “reverse translate”, “reverse translated”, “reverse translation”, “reverse translating”, “reverse translation operations”, etc. refer to reverse translating macromolecular information as further described in US provisional no. 63 / 620,076, and PCT / US23 / 70077, which are herein incorporated by reference in their entirety.

[0373] A “sample” may include a biological sample. Examples of samples include fluids or biofluids that include biological material such as cell or peptides.

[0374] The term “sample index” refers to an identifier incorporated during a post-recode preparation of a DNA library for NGS analysis, or an identifier that can be ligated as a component of a memory oligo during its assembly, and used during NGS analysis to identify the provenance of oligonucleotides in the DNA library.

[0375] The term “solid support” or “surface” refers to any solid material substrate in planar form, spherical form, or a combination of forms including, but not limited to: a solid bead, a porous bead, a solid planar material, a porous planar material, a patterned or non-patterned solid material, a nanoparticle, or a inorganic or polymeric microsphere, or a capillary. For example, the solid support may comprise a glass slide or wafer, a silicon slide or wafer, a PC, PTC, polyethylene (PE), high density polyethylene (HDPE), or other plastic slide, a teflon, nylon, nitrocellulose membrane, or borosilicate capillary, a ceramic surface or a gold surface. Particles and beads may be formed from polystyrene, cross-linked polystyrene, agarose, acrylamide, silica, silica oxide, metal oxides. Beads or nanoparticles may be magnetic or paramagnetic to support separation or purification processes. Solid supports may be passivated with glass, silicon oxide, alkyl silanes, functionalized silanes, tantalum pentoxide, DLC diamond-like carbon, or other passivation agents. A “solid support,” including membranes, may be passivated or activated via corona or other plasma treatments methods. Solid supports may further be assembled with other components to facilitate fluid transport and / or detection (e.g., flowcell, biochip, a microtiter plate. Solid supports may comprise an associated hydrogel that supports joining components for macromolecule recoding and / or analysis workflows. In certain examples, the term, “solid support” may include any of the described solid supports above further associated with a hydrogel.

[0376] The term “splint” refers to a nucleic acid with complementarity to the 5’ end of one nucleic acid and the 3’ end of another nucleic acid, such that hybridization of the splint to both nucleic acids brings the 5 ’and 3’ ends into proximity to promote either chemical or biological ligation.

[0377] The term “strobe sequencing” refers to a method of sequencing (e.g., nucleic acids, peptides, and other polymers) wherein short gapped reads, or interspersed subreads, are generated from a contiguous fragment rather than a single uninterrupted read. Such subreads are referred to as “strobe” or “strobed” reads.

[0378] The term “Tri-Functional Support” (TFS) refers to a solid support comprising (x) a solid support surface, (y) immobilized peptide unique molecular association tags (PUMA tags), and (z)reactive moiety(s) associated with each PUMA tag for binding an amino acid residue of a protein. A TFS bead type is defined by its PUMA tag sequence. A pool of TFS beads refers to a mixture, typically of roughly equal proportions, of bead types having PUMA tag sequences that differ from one another. In certain embodiments, various moieties may be replaced with reactive moieties for attaching said groups. A TFS may take any form and / or material as described for a “solid support” and may further comprise a hydrogel, or other functionalization or passivation molecules, especially those that may facilitate a reverse translation operation(s). A tri-functional molecule of the TFS may further comprise: a moiety(s) that facilitates solubility, a moiety(s) reduces non-specific interactions, a sequence(s) or moiety(s) that facilitates removal of the PUMA-tagged protein from the TFS surface, such as cleavable linker moieties as described herein, or any combination of the aforementioned components. A TFS may be used during operations related to single-cell assays, and / or may be used during operations related to spatial assays, or both. When used to capture proteins from a tissue section for spatial analysis the TFS may lack various functional groups that support reverse translation operations, and may provide a higher density of capture site or larger size appropriate for protein capture.

[0379] As used in the present disclosure, the term “unique molecular identifier” or “UMI” refers to a nucleic acid molecule of length 10 to 40 bases that can be assembled into, e.g., the memory oligo and provides unique identification for in silico deconvolution of NGS sequencing data as to a specific memory oligo.

[0380] The term “universal priming site” or “universal primer” refers to a nucleic acid molecule, which may be used for library amplification and / or during NGS. Exemplary universal priming sequences can include P5, P7, P5’, P7’, SBS Read 1, and SBS Read 2 primers.

[0381] The term “universal sequence” or “universal assembly sequence” or “universal amplification sequence” refers to a common complementary polynucleotide sequence that can be appended to a 3’ and / or 5’ end of a tag, e.g., a recode tag, for facilitating amplification thereof with common primers or assembly into an oligo, e.g., a memory oligo. In certain embodiments, a universal sequence comprises a repetitive sequence, e.g., a dinucleotide repetitive sequence such as (GT)n, or other relatively short nucleotide motif. The universal sequence may be silent during sequencing of the oligo to facilitate efficient detection and analysis of the assembled constituents of the oligo.

[0382] The term, “workflow cycle” or “cycle” refers to the iteration number of any one of the operations of a process flow or method described herein.

[0383] Some embodiments refer to a sequence. The sequence may be included in the accompanying sequence listing. Any discrepancies between the sequence listing and specification may usually be resolved by referring to the sequence as described in the specification.

[0384] References to PUMA sequence are employed and may be included or named as in Table 6.Table 6: Example PUMA Sequence

[0385] Note that as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to “an oligo” refers to one or more oligos, and so forth. Additionally, it is to be understood that terms such as "left," "right," "top," "bottom," "front," "rear," "side," "height," "length," "width," "upper," "lower," "interior," "exterior," "inner," "outer" that may be used herein merely describe points of reference and do not necessarily limit embodiments of the present disclosure to any particular orientation or configuration. Furthermore, terms such as "first," "second," "third," etc., merely identify one of a number of portions, components, steps, operations, functions, and / or points of reference as disclosed herein, and likewise do not necessarily limit embodiments of the present disclosure to any particular configuration or orientation.

[0386] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. All publications mentioned herein are incorporated by reference for the purpose of describing and disclosing devices, methods and cell populations that may be used in connection with the presently described disclosure.

[0387] Where a range of values is provided, it is understood that each intervening value, between the upper and lower limit of that range and any other stated or intervening value in that stated range is encompassed within the disclosure. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges, and are also encompassed within the disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure.

[0388] In this description, numerous specific details are set forth to provide a more thorough understanding of the present disclosure. However, it will be apparent to one of ordinary skill in the art that the present disclosure may be practiced without one or more of these specific details. In other instances, well-known features and procedures well known to those skilled in the art have not been described in order to avoid obscuring the disclosure.

[0389] The functionalities described in connection with one embodiment are intended to be applicable to the additional embodiments described herein except where expressly stated or where the feature or function is incompatible with the additional embodiments. For example, where a given feature or function is expressly described in connection with one embodiment but not expressly mentioned in connection with an alternative embodiment, it should be understood that the feature or function may bedeployed, utilized, or implemented in connection with the alternative embodiment unless the feature or function is incompatible with the alternative embodiment.

[0390] The practice of the techniques described herein may employ, unless otherwise indicated, conventional techniques and descriptions of organic chemistry, polymer technology, molecular biology (including recombinant techniques), cell biology, proteomics, biochemistry, and sequencing technology, which are within the skill of those who practice in the art. Such conventional techniques include polymer array synthesis, hybridization and ligation of polynucleotides and other polymers, and detection of hybridization using a label. Specific illustrations of suitable techniques can be had by reference to the examples herein. However, other equivalent conventional procedures can, of course, also be used. Such conventional techniques and descriptions can be found in standard laboratory manuals such as Green et al., Eds. (1999), Genome Analysis: A Laboratory Manual Series (Vols. I-IV); Weiner, Gabriel, Stephens, Eds. (2007), Genetic Variation: A Laboratory Manual; Dieffenbach, Dveksler, Eds. (2003), PCR Primer: A Laboratory Manual; Mount (2004), Bioinformatics: Sequence and Genome Analysis; Sambrook and Russell (2006), Condensed Protocols from Molecular Cloning: A Laboratory Manual; and Sambrook and Russell (2002), Molecular Cloning: A Laboratory Manual (all from Cold Spring Harbor Laboratory Press); Stryer, L. (1995) Biochemistry (4th Ed.) W.H. Freeman, New York N.Y.; Gait, “Oligonucleotide Synthesis: A Practical Approach” 1984, IRL Press, London; Nelson and Cox (2000), Lehninger, Principles of Biochemistry 3rd Ed., W. H. Freeman Pub., New York, N.Y.; Berg et al. (2002) Biochemistry, 5th Ed., W.H. Freeman Pub., New York, N.Y.; all of which are herein incorporated in their entirety by reference for all purposes.Table 7: Example Reference Protein Library SequencesEXAMPLES

[0391] The examples are put forth so as to provide those of ordinary skill in the art with a complete disclosure and description of how to make and use the present disclosure, and are not intended to limit the scope of what the inventors regard as their disclosure, nor are they intended to represent or imply that the experiments below are all of or the only experiments performed. It will be appreciated by persons skilled in the art that numerous variations and / or modifications may be made to the disclosure as shown in the specific aspects without departing from the spirit or scope of the disclosure as broadly described. The present aspects are, therefore, to be considered in all respects as illustrative and not restrictive.Example 1: Single-cell Proteomic Analysis Incorporating PUMA Tags

[0392] This example describes the process utilized for the incorporation of PUMA tags using the Single-cell Proteomic analysis disclosed herein.

[0393] For this example, a commercial cell sorter is used to singulate and isolate cells into 100 wells of a 1536-well plate (Scionion, cellenONE). Note that additional single-cell isolation and reagent addition systems are also compatible (e.g., 10X Chromium). To the partitioned cells is added luL of 0.3% TritonX-100 and the plate is incubated for 5 minutes at RT. A 3uL aliquot of Albumin Depletion Kit (Abeam, Cat. ab241023) beads is added to each well. They provide Cibacron Blue 3G-A immobilized on beads and will make albumin less accessible and less able to participate in subsequent steps. A luL aliquot of a unique 50uM PUMA-conjugate is added to each well and incubated with the permeabilized, depleted cells for 60 minutes at room temperature to drive a TCO-tetrazine click reaction. Subsequently, the PUMA-protein conjugates are captured onto azide functional beads at a density of approximately one protein per lum2 by addition of a 3uL aliquot of a 10%w / v beads into each well. This equates to the surface area required to support reverse translation of approximately 3E9 proteins. Beads are incubated for 60 minutes with mild shaking under standard CuAAC click reaction conditions to ensure complete capture, and excess reagents are washed, the beads are dried, and ready for reverse translation as described in Example 4. Analysis proceeds as described in Example 4 below.Example 2: Spatial Proteomic Analysis Incorporating PUMA Tags

[0394] This example describes the process utilized for the incorporation of PUMA tags using the Spatial Proteomic analysis disclosed herein.

[0395] A solid support for deposition of a tissue micro-sectioned sample is prepared by spin coating 500uL of hydrogel polymer using a Sigma Chemat precision spin-coater at 500 rpm for 1 minute onto a corning glass slide. The hydrogel polymer is obtained by co-polymerization of acrylamide with modified acrylate-based monomers having sidechains that include azide functional groups. Briefly, a RAFT polymerization of acrylamide and azido acrylate monomers follows procedures as described byPalmiero et.al. Polymer (2016), 98, 156-164. Combinatorial methods and a commercial microarray printer (Scionion) allow deposition of PUMA-conjugates in a precise pattern at 80um resolution across a 10-mm square pattern as a proof of principle. The hydrogel is partially dried, in preparation for deposition of a tissue sample.

[0396] A fresh-frozen oncology tissue specimen and matched control are obtained (Precision Biospecimens). The tissue specimens are sectioned using a cryomicrotome. The transfer substrate prepared with PUMA-conjugates is inverted and pressed against the microtomed tissue sections such that the PUMA-conjugates are in contact with the tissue and the tissue adheres to the transfer substrate. The transfer substrate with tissue section is fitted into a slide holder (Grace Bio Labs) to facilitate addition of permeabilizing reagents. Tissue sections are incubated in a 0.2% Tween 20 in PBS at 22C for 30 minutes. During this time PUMA-conjugates interact with locally-released cellular components to create PUMA-protein conjugates. The permeation solutions for the sample and control transfer substrates are recovered by pipet and transferred to centrifuge tubes for further processing, including protein sidechain derivatization with a ImM iodoacetic acid solution. Azido-functionalized beads that are capable of supporting reverse translation are added to the tube under conditions (lOmM ascorbic acid, 2mM PMDETA, and 0.5mM Cu2+ catalyst) that support conjugation of the PUMA- protein conjugates. Reverse translation proceeds as described in Example 4 below. Analysis proceeds as described in Example 4 below. Alternatively, after recovery and processing of the permeation solutions, high density azido-functionalized beads are added that serve to capture the PUMA-protein conjugates under conditions that support conjugation (lOmM ascorbic acid, 2mM PMDETA, and 0.5mM Cu2+ catalyst). After a washing step, the cleavable disulfide linker in the PUMA-protein conjugate is cleaved using a ImM TCEP solution, and the PUMA-protein conjugates are released. The PUMA-protein conjugates may then be captured via reaction with maleimide groups on the reverse translation surface. Reverse translation proceeds as described in Example 4 below. Analysis proceeds as described in Example 4 below.Example 3: Preparing a Single TFS Bead Type

[0397] This example describes the preparation of a single Tri-Functional Support bead type that may be associated with analytes, and the release of PUMA-analyte conjugates from the prepared TFS beads. A Tri-Functional Support (TFS) bead comprises (a) a solid support surface, (b) an immobilized peptide unique molecular association tag (PUMA tag), and (c) a reactive moiety(s), associated with each PUMA tag, for binding an amino acid residue of an analyte, as illustrated in an exemplary configuration in FIG. 4. Accordingly, a TFS bead may be useful during operations related to single-cell assays, and / or may be used during operations related to spatial assays, or both. In the present example to prepare one TFS beadtype, 5um spherical silica beads (Bangs SS05003) solid surface were procured. The PUMA tag is a synthetic 9 amino acid peptide having a C-terminal azide group (Biomatik). A TFS bead type is defined by its PUMA tag sequence, therefore, this bead type is defined by the sequence of the 9 aminoacids of the synthetic peptide, which is: A-W-S-M-E-T-E-C-H-azide (SEQ ID NO: 174). A custom trifunctional molecule, N-(Propargyl-PEG2)-DBCO-PEG3-Amine, TFA salt (PDA, Broadpharm Cat# 29932) is used to assemble the various elements into a functional TFS bead type. Assembly proceeded as follows:

[0398] Azide-functionalized beads are prepared by exchanging 1 mL of 5% w / v beads into ethanol, then incubating in a 2.5% v / v ethanolic solution of 3-aminopropyltriethoxysilane (Gelest, SIA0611.0) for 60 minutes at RT. The silane solution is removed by washing twice with ethanol and twice with acetone, acetone is removed, and the beads are dried at 120C for 2 hrs. The surface amine groups are converted to azide groups using azidoethyl-SS-propionic acid (Broadpharm, Cat. 24063) under mild standard NHS / EDC coupling conditions. Subsequently azide-functionalized beads are resuspended in DI water, creating ImL of a 5% w / v slurry.

[0399] The PUMA tag (0.5 mg) is dissolved in 0.5 mL water to create an ~0.8 mM solution. PDA is incubated with PUMA tags at approximately equimolar concentration in phosphate buffer, pH 7.2 for 60 minutes, and the conjugated product was purified by SAX HPLC using a 0% to 70% linear gradient over 20 minutes, A = PBS, B= PBS, IM NaCl. Dibenzocyclooctyne (DBCO) reagents represents a class of click chemistry labeling reagents that reacts with azide-tagged molecules without the need for catalyst to form stable triazoles, while the more stable alkyne group requires catalysis, typically via Cul-i- to couple with azides. Thus, in the next step, the heretofore unreacted propargyl group of PDA is coupled to the azide bead by incubation of 1.0 mg of PUMA-PDA conjugate with 50mg of azide beads in a solution of 20mM ascorbate, ImM copper sulfate, and ImM tris(3- hydroxypropyltriazolylmethyl)amine, phosphate buffer, pH 7.2 for 60 minutes at room temperature. Beads are washed to remove unreacted components, and finally, the amine of the original PDA molecule is reacted with NHS-TCO. The beads are exchanged into PBS pH 7.5 and allowed to settle. The supernatant is removed, leaving the wet bead pack. A 1 mM solution of NHS-TCO (Broadpharm, Cat. 22417) is prepared in DMSO and added to the beads to cover the beads pack, pH 9.0 carbonate buffer is added, the mixture is vortexed and allowed to incubate for 4 hours at RT. Subsequently, the beads are washed with DMSO and water and finally exchanged into PBS and stored at 4C.

[0400] These TFS beads are complementarity reactive to proteins which are labeled with methyltetrazine-amine HC1 salt (Broadpharm, Cat. 22433) under mild standard NHS-EDC coupling conditions. Methyltetrazine-amine labels at carboxyl groups, such as the C-terminal carboxyl or the carboxyl sidechain of aspartic or glutamic acid to yield tetrazine-labeled analytes having free N- terminal amine for participation in reverse translation. Alternately an ITC-tetrazine conjugate synthesized by conversion of amine-PEG2-methyltetrazine with thiophosgene, using TEA base catalysis in THF for 4 hours at RT may be used to label amines of protein analytes, which may subsequently react to form a TFS bead associated with analytes.

[0401] In this example, the TFS bead does not comprise reactive groups that support reverse translation operations. Instead, the surface density of PUMA-conjugates is tuned for maximum conjugation of proteins and PUMA tags. Protein analyte loading between 1E5 and 1E6 molecules per bead is achieved. Release of PUMA-analyte conjugates from a TFS beads may be accomplished through multiple methods. In this example, a ImM dithiothreitol solution or ImM TCEP solution is added to the beads to release the PUMA-analyte conjugates. Subsequently, the PUMA-analyte construct may be captured to a surface that supports reverse translation using stable maleimide bonds.Example 4: Characterization of the Reverse Translation Product of a Peptide Having NonNative Amino Acids

[0402] The reverse translated information of PUMA tags may be readily deconvoluted from reverse translated information of a protein analyte when the PUMA tags comprise non-natural amino acids. In this example, a synthetic peptide having FAM and biotin derivatized to the epsilon-amine of lysine sidechains was analyzed. Several constructs were created.Peptide Functional Beads

[0403] The peptide (0.5 mg, -1100 g / mol, sequence from N-terminus to C-terminus: {Lys-FAM} {Ser} {Ser} {Eys-biotin} {Ser}-pra) (SEQ ID NO: 264) was dissolved in 0.5 mF water to create an -1 mM solution. Peptide immobilization reaction was initiated by combining the reactants in Table 8 (volumes in pF). The reaction was conducted at 50C for 1 hr on a rotator. The beads were subsequently washed by exchanging 1 mF volumes of solutions in the following order: 100 mM pH 9.6 carbonate buffer, DMSO, water, 100 mM pH 9.6 carbonate buffer, water, DMSO. After each addition, the beads were resuspended through shaking, centrifuged (21,000 ref 1 min), and the supernatant was removed. The second DMSO solution was incubated with the beads at 57C for 4 min, followed by serial exchange into acetonitrile, water, 100 mM pH 9.6 carbonate buffer, and water, as described above. Eastly the beads were incubated in carbonate buffer for 10 min at ambient temperature, and analyzed using a fluorescent plate reader (545 nm excitation, 586 nm emission) to confirm FAM signal indicative of immobilized peptide.Table 8: Example Reaction CompositionCRC

[0404] Briefly, ~10 flg oligonucleotide having a 5’ alkyne and internal amine functional group was dissolved in 0.5 mL water to create an ~1 mM solution. The amine of the oligonucleotide was reacted with NHS-TCO in DMSO carbonate buffer pH 9.0 (50:50), and purified via desalting (Glen Research, Gel-Pak) followed by spin filtration (EMD Millipore3 kDa). The process was repeated with additional oligonucleotide sequences to produce additional CRC molecules.Binding Agent

[0405] An anti-biotin binding agent was constructed by incubating a 5’ biotin-labeled DNA recode tag oligonucleotide (IDT) at lOOuM concentration with streptavidin (Sigma Cat. 11721674001) at Img / mL at RT for 30 minutes at a stoichiometry such that biotin binding sites on streptavidin are not fully occupied and purified by SAX HPLC using a 0% to 70% linear gradient over 20 minutes, A = PBS, B= PBS, IM NaCl.. Similarly, an anti-FAM binding agent was constructed. First the antibody (Sigma, F5636) was azidylated using NHS-PEG- Azide (Broadpharm, Cat. 21857) by incubating the reagent with the antibody at RT for 120 minutes. Excess reagent was removed by SAX HPLC using a 0% to 70% linear gradient over 20 minutes, A = PBS, B= PBS, IM NaCl. A 5’ alkyne-labeled DNA recode tag oligonucleotide was coupled under conditions and using protocols that are well known (lOmM ascorbic acid, 2mM PMDETA, and 0.5mM Cu2+ catalyst, Presolski et al. (2011) Copper-Catalyzed Azide-Alkyne Click Chemistry for Bioconjugation. Current Protocols in Chemical Biology. 3(4), 153— 162; Hong et al., (2009) Analysis and Optimization of Copper Catalyzed Azide- Alkyne Cycloaddition for Bioconjugation. Angew. Chem. Int. Ed., 48(52), 9879-9883). Excess oligo was removed by SAX HPLC using a 0% to 70% linear gradient over 20 minutes, A = PBS, B= PBS, IM NaCl.Reverse Translation

[0406] Reverse translation of the peptide proceeded according to methods described in US provisional no. 63 / 620,076, and PCT / US23 / 70077. Briefly, beads having a surface functionalized to support reverse translation and immobilized proteins were provided, as above. Cycles of chemistry whereby amino acids of the PUMA tag were isolated and immobilized on a solid support were accomplished. To derivatize the N-terminal amine of the immobilized peptide with an ITC-conjugate 50uL of a 20mM solution of methyltetrazineisothiocyanate in DMSO and 2.5uL of diisopropylethylamine were contacted to a 1 mg sample of silica beads functionalized with the model PUMA tag under anhydrous conditions and incubated for 30 minutes at 50C. A CRC described above, where X’ is TCO was provided. A lOuL aliquot of lOuM CRC was contacted with the beads and the tetrazine moiety of the N-terminally bound ITC-tetrazine conjugate was allowed to couple with the TCO moiety of the CRC for 10 minutes atRT. Unreacted CRC was washed from the surface extensively using 65% DMSO in water. The solution was exchanged for click reaction buffer (neutral pH PBS, 2mM PMDETA, ImM copper sulfate, lOmM sodium ascorbate) and the alkyne group of the N-terminally-bound CRC was allowed to react with the surface-bound azide groups of the peptide functionalized beads for 30 min at room temperature. Cleaving the N-terminal amino acid to form the immobilized amino acid complex was completed by exchanging the beads into dry acetonitrile, removing excess solvent, and incubating in 60uL of anhydrous trifluoroacetic acid for 20 minutes at 50C. These steps were repeated 4 times.

[0407] Recognition of the isolated amino-acid complexes and assembly of recode blocks also proceeded according to methods described in US provisional no. 63 / 620,076, and PCT / US23 / 70077. Briefly, 10 nM anti-FAM and lOnM of anti-biotin binding agents in PBS buffer were incubated with the beads for 60 minutes at RT. As a negative control (condition 4), some of the beads were separated and buffer without anti-FAM binding agent was incubated. The binding agents were washed with 10 mL of buffer, and the beads were resuspended in lx T4 DNA ligase buffer (NEB Cat. B0202), with 10 nM ligation oligo and 20 units / pL T4 DNA ligase (NEB cat. M0202). Ligation proceeded for 60 mins at RT. The negative control sample was treated similarly, however in the absence of ligation oligo and T4 DNA ligase. Following ligation, the beads were washed using buffer with standard surfactants to remove excess ligation oligo and resuspended. To further assemble recode blocks into a memory oligo the beads incubated in an extension mix solution of lx isothermal amplification buffer (NEB Cat. B0537), 8 mM magnesium sulfate, 1.4 mM each dATP, dCTP, dGTP, and dTTP, 8 units Bst 2.0 DNA polymerase (NEB Cat. M0537), and 100 nM of a 3’ blocked initiator oligo in a total reaction volume of 25 pL. The mixture was incubated at 55 °C for 10 minutes in a thermocycler, cooled to room temperature, and then beads were washed with buffer to remove reactants. As a second negative control (condition 3), initiator oligo was not included in the assembly mix. In the next step of memory oligo assembly, the beads were resuspended roughly 50uL of lx NEBuffer 4 (NEB Cat. B7004) having 10 units of T7 exonuclease (NEB Cat. M0263) for 30 minutes at RT. A third negative control (condition 2) was similarly treated, except without addition of exonuclease. The beads were washed and resuspended in lx isothermal amplification buffer (NEB Cat. B0537) and again treated with the extension mix, washed, and then resuspended using deionized water.

[0408] A qPCR method to analyze the recode block and memory oligo process included: qPCR master mix (NEB Cat. M3003) which contains SYBR Green intercalating dye and Taq DNA polymerase. Each sample (conditions 1-4 plus standards) was amplified using three different sets of PCR primers: primers specific to the anti-FAM recode, primers specific to the anti-biotin recode block, and primers specific to the memory oligo. PCR and melt curve analysis were performed on an Applied Biosystems Quant Studio 3 real-time PCR machine. Further, the PCR products were diluted 5-fold and run on a 10% TBE polyacrylamide gel (ThermoFisher, Cat. EC62752BOX) at 200 V constant voltage for 30 minutes, then stained with lx SYBR Green dye in TE buffer at pH 7.6 for 30 minutes and imaged on an Invitrogen iB right FL 1500 imaging system. Table 9 shows lane identification of the gel in FIG. 8.Table 9: Lane Identification of PCR Results from FIG. 8Analysis

[0409] PCR results of FIG. 7 indicated that the recode block for the lysine-FAM (condition 1) was amplified with a mean Ct of 22 and generates a product with a melt temperature of 79.1 °C, whereas the negative control (condition 4) has a mean Ct of 33 and generates a product with a melt temperature of 77.2°C, indicative of primer dimer product. Gel results in FIG. 8 corroborate that the recode block identified in condition 1 (lane 4) is of the correct length and matches the standard (lane 6) while the negative control does not result in the intended product (lane 5.) Similarly, the conditions related to memory oligo indicate that recode blocks generated from non-natural amino acids of a PUMA tag may be assembled into a memory oligo. Gel results in FIG. 8 corroborate the formation of the memory oligo in condition 1 (lane 7), while negative controls (conditions 2-4, lanes 8-10) had little or no apparent memory oligo formation.Example 5: Preparation of a PUMA Construct and TFS Bead

[0410] Following the synthetic scheme shown in Fig. 11, 1 mg Peptide #20 (NH2-I-K-H-Q-R-S-Pra (SEQ ID NO. 265), where Pra is propargyl glycine) was dissolved in 100 pL water to create a 11.6 mM solution. A 20pL aliquot of Peptide #20 solution was combined with 1 pL neat BP-24507 (N-bis(Azido- PEG3)-N-(BocNH-PEG2)), 15 pL 484 mM THPTA in water, 5 pL 291 mM CuSO4 in water and 20 pL dimethylformamide. To this mixture, 15 pL 10% sodium ascorbate in water was added. The mixture was heated at 40C for 2 hrs, and the product, designated Pep20-24507, was isolated via reverse phaseHPLC (sol A: 35 mM TEAA+5% ACN in water; sol B: ACN, linear gradient 5% to 95%) at a retention time of 18.4 min. The retention time of Peptide #20 was 1.8 min and the retention time of BP-24507 was 23.8 min by comparison. The mass of the product Pep20-24507 (1514 g / mol) was confirmed via mass spectrometry ([M+H]+ m / z = 1515 and [M+2H)]2+ m / z = 758) (FIG. 12).

[0411] In the next step, 0.5 mg peptide #39 (NH2-A-S-S-Pra, 358 g / mol) was dissolved in 100 pL water to form a 13.9 mM solution. Pep20-24507 isolated from the previous step (solution in 20 pL water) was combined with 20 pL of the peptide #39 solution, along with 15 pL 484 mM THPTA in water, 5 pL 291 mM CuSO4 in water and 20 pL dimethylformamide. To this mixture, 15 pL 10% sodium ascorbate in water was added. The mixture was heated at 50C for 3 hrs, and the product (PUMA- 1) was isolated by reverse phase HPLC with a retention time of 11 min. The product (PUMA-1) mass (1872 g / mol) was confirmed by mass spectrometry ([M+2H]2+ m / z= 937) (FIG. 13). The PUMA-1 peptide N-termini can be PITC coupled, and the Boc protected amine group deprotected in order to create a functional handle (amino group) for subsequent immobilization of PUMA-1 onto a surface. Either acid or thermal methods for Boc deprotection may be employed.Example 6: Ligation and Amplification of Oligonucleotides in situ

[0412] This example describes the synthesis of functionalized silica beads, streptavidin oligo conjugate, successful ligation and PCR amplification of oligonucleotides on a solid support to form an exemplar memory oligonucleotide using a custom peptide conjugated to silica beads. FIG. 14 illustrates an embodiment in which the C-terminus of the peptide was covalently linked to the bead surface and a chemically-reactive conjugate containing an oligonucleotide (PPO-[ / 5Phos / ATGAGTG / iFormInd / AGGGAAATAGCTTCTGGTCGAACTAGTTGTTCGTCAA (SEQ ID NO: 266)]-SOC) was reacted with the N-terminal amine of the peptide. Streptavidin was labelled with a second oligonucleotide (Syst#002-SOC-[ / 5Phos / GAACGTG / iFormInd / CTTCTGATGAAGTTTGGAGACAAATTGC GTGGGAGCA (SEQ ID NO: 267)]) and bound via biotin-streptavidin interaction to form a model affinity complex. The two oligonucleotides, now in close proximity, were ligated using a sequencespecific splint oligonucleotide and T4 DNA ligase. After ligation and amplification of the resultant surface-bound oligonucleotide, qPCR was performed using primers specifically designed to amplify the ligated product, thus, amplification only occurred if the complete ligation product was present. As a positive control, the ligation reaction was performed in solution in the absence of the peptide. Similar Ct values were observed for the solution ligation and on-bead ligation conditions (FIG 15). As a negative control, the beads were prepared and incubated without the addition of T4 DNA ligase. Melt curve analysis indicated the desired product was produced (Tm = 80 °C) in both the bead and the solution ligation samples, while any product produced in the negative control sample was non-specific off-target amplification (Tm ~ 68 °C).

[0413] FIG. 15 shows qPCR amplification curves indicating the successful ligation of the products on bead thereby showing steps of the method: (f) contacting the immobilized amino acid complex with abinding agent, the binding agent comprising: a binding moiety for preferentially binding to the immobilized amino acid complex, and a recode tag comprising a recode nucleic acid corresponding with the binding agent, thereby forming an affinity complex, the affinity complex comprising an immobilized amino acid complex and the binding agent and thereby bringing the cycle tag into proximity with the recode tag within the affinity complex, (g) transferring information of the recode nucleic acid to the cycle nucleic acid of the immobilized conjugate complex to generate a recode block; and finally (h) obtaining sequence information of the recode block, in this case via PCR, melt temperature analysis, and Sanger sequencing.Synthesis of Azide-Functionalized Silica Beads:

[0414] Five mg of amine-functional silica beads (CD Biosciences DNG-F046, 20 pm, 5 wt%, 4 pmol / g NH2) were added to a 0.5 mL protein lo-bind tube. Solution exchange was accomplished using an Eppendorf benchtop fixed angle centrifuge at 21000 x g for 1 minute. Supernatants were carefully removed by pipette, ensuring the bead pellet remained undisturbed. Surface amine groups were converted to azide using 34 mg of Azido acetic acid NHS ester (AA-NHS, Broadpharm BP-22467) under standard NHS-coupling conditions.

[0415] Following conversion, the beads were rinsed twice with deionized water, followed by a single wash with 200 mM carbonate buffer (pH 9.6). Subsequent washing steps were conducted thrice with deionized water. After the last wash, the supernatant was removed, and the bead pellet was resuspended in a solution containing 1 mg of fluorescamine (Aldrich cat F9015) dissolved in 1 mL of DMSO to test for residual amine. After allowing this mixture to react for 10 minutes at room temperature, the rinsed beads and supernatant solutions were transferred to a 96-well plate, and fluorescence was measured using a plate reader, yielding a bead RFU value of 1.42xlOA6. Residual amines were capped using a solution of 1.32 M succinic anhydride in 0.32 mL of dimethylformamide (DMF), 10% Diisopropylethylamine (DIPEA, Aldrich cat D125806). Following reaction at 60C for 2 hours, excess reactant was removed by serially washing with DMF DMSO, and water. A fluorescamine assay was again performed to check for residual amines after succinilation and reported acceptably low background signal. Finally, the beads were suspended in IX SSPE buffer (prepared from Aldrich cat 15591043 20X stock) and stored at 4°C, shielded from light.Synthesis of Peptide 5 -Functionalized Silica Beads:

[0416] The azide-functionalized beads prepared as described above served as the starting material. These were combined with a 200 mM phosphate buffer at pH 7, 50 mM THPTA, 10 mM CuSO4, and 0.8 mM of Peptide 5, which has the sequence {lys-biotin} {ser} {ser} {lys-FAM} {ser}-Pra (SEQ ID NO: 268); where “Pra” denotes a propargyl glycine group at the C-terminus. Freshly prepared 100 mM sodium ascorbate was also added.Synthesis of streptavidin- oligo conjugate

[0417] Streptavidin (SA, Sigma - SA101) was solubilized in PBS buffer to 100 pM, yielding approximately 2 mL. NHS-PEG4-DBCO (BP-22288) was prepared at a lOmM concentration in DMSO. NHS-PEG4-DBCO was added to the Streptavidin in a 2-fold molar excess, targeting 1-2 linkers per SA molecule. The reaction proceeded at room temperature for 60 minutes. Unreacted NHS-PEG4-DBCO was removed via serial rinses using a 10k MWCO spin column (Sigma UFC5010). The conjugate was stored at -20°C. Purity and yield were evaluated using absorption values at 280 nm and 309 nm.

[0418] / 5Phos / GAACGTG / iFormInd / CTTCTGATGAAGTTTGGAGACAAATTGCGTGGGAGCA (SEQ ID NO: 267) was reacted for 16h at 40C with aminoxy-PEG-azide in the presence of 5- aminoindole catalyst at pH 6.5, and purified via HPLC, dried, and resuspended in nuclease-free DI water. Verification of the correct product was conducted via high-performance liquid chromatography (HPLC) and electrospray ionization time-of-flight mass spectrometry (ESI-TOF-MS). Subsequently, the N3-oligo conjugate at 100 pM was mixed with the DBCO-SA conjugate at a 1.3:1 molar ratio of oligo to DBCO-SA. The reaction proceeded for 2 hours at room temperature. Unreacted N3-oligo was removed via replicate rinses through a 30k MWCO spin column with PBS and the conjugate concentration was determined to be -120 pM.Preparation of Beads having Affinity Complexes

[0419] The chemically-reactive conjugate, PPO-Sysl, as described above (see synthesis of PPO, a trifunctional CRC) with the SEQ ID 266 substituted for SEQ ID 269), purified by HPLC, were introduced to beads in 200 mM carbonate (pH 9.6) and allowed to couple to the N-terminal alpha amine of the peptide. Beads were then rinsed with SSPE buffer.

[0420] The SA-oligo conjugate was introduced to the bead at nM concentration, and the excess removed by copious rinsing with PBS.Ligation of Oligos on Bead Surface:

[0421] The oligonucleotide ligation steps utilized several components, including T4 ligase (NEB cat M0202S), T4 DNA Ligase Reaction Buffer 10X (NEB cat B0202SVIAL), and a 1 M NaCl solution. Nuclease-free water was used throughout the process. In 0.2mL PCR tubes, the following reactions were mixed and prepared as described in Table 10:Table 10: Example Reaction Mixture

[0422] 4pL of T4 ligase buffer (10X) and IpL of T4 ligase was added to each tube (except no T4 ligase in the no-ligase control), and gently pipette mixed. The reaction proceeded at RT for 1 hour, then the ligate was heat inactivated at 65 °C for 10 min. qPCR of Ligation Products and Controls:

[0423] The real-time PCR was performed using SYBR Green Master Mix (Bio-Rad cat 1708880). qPCR cycling was run on a standard mode with an initial denaturation step of 3 minutes at 95 °C, followed by 40 cycles of 10 seconds at 95°C and 30 seconds at 60°C. The melt curve stage started at 65°C, increasing by 0.5°C every 5 seconds until 95°C. Primers for amplification included SysOOl PR1 and Sys002 PR3, and appropriate controls were set to assess the efficiency of the qPCR reaction. Data analysis was performed using a qPCR software suite.Example 7: Use of Alternative Nucleic Acid Tags and Transfer of Information into Polymerizable Molecules

[0424] This example demonstrates methods for incorporating PUMA tag information from non- polymerizable nucleic acid analogs into sequenceable DNA constructs. In certain embodiments, the method comprises utilizing a peptide nucleic acid (PNA) molecule that contains unique molecular identifier information, wherein said PNA molecule cannot directly participate in enzymatic ligation or polymerization reactions. The method employs a bridging oligonucleotide comprising a first domain complementary to the PNA sequence and a second domain that functions as a splint to facilitate ligation. The bridging oligonucleotide enables the transfer of sequence information from the PNA to a ligation product comprising a recode tag oligonucleotide, which contains information about a detected molecule, and a ligation oligonucleotide, which contains identifier information. Through the action of the bridging oligonucleotide, the recode tag and ligation oligonucleotide are brought into proximity and joined via enzymatic ligation to form a recode block. The resulting recode block incorporates both the molecular identity information of the recode tag and the identifier information of the PNA molecule. This method enables the integration of information from non-enzymatically active nucleic acid analogs into amplifiable and sequenceable DNA constructs.

[0425] This example illustrates a method for transferring information from non-polymerizable nucleic acids to polymerizable nucleic acids by showing the transfer of information from PNA to DNA. A model peptide (PEP6) having the sequence {pTyr,Ser,Lys-FAM,Ser,Lys-Biotin,Ser-Pra} (SEQ ID NO: 270) was prepared and conjugated to solid support beads according to methods described in Example 6. The peptide was subjected to five cycles of sequential amino acid removal, wherein cycles 1 and 5 employed tetrazine-modified phenylisothiocyanate (Tz-PITC) in conjunction with chemically-reactive conjugates PPOT2 and PPOT3 (PPOT: N-(Pentynoyl)-5'-oligopeptidenucleic acid-3'-(8-trans(cyclooct-5-enyloxyacetyl)lysine)), respectively, while cycles 2-4 utilized unmodified PITC. The chemicallyreactive conjugates PPOT2 and PPOT3 comprised PNA cycle tags having the sequences CTT GCA CAG AAG ACT (SEQ ID NO: 271, PNA2, 3’ to 5’) and ACT TCA AAC CTC TGT (SEQ ID NO: 272, PNA3, 3' to 5'), respectively.

[0426] Bridge oligonucleotides complementary to the PNA cycle tags were synthesized as follows: BR-001 (5’ to 3’): GAA CGT GTC TTC TGA CCA ACT CAT GTT GGA CGA AAG TAC AAT GTC / 3ddC / (SEQ ID NO: 273, complementary to PPOT2 PNA)

[0427] BR-002 (5’ to 3’): TGA AGT TTG GAG ACA CCA ACT CAT GTT GGA CAT TCG CAA CTA TTC C (SEQ ID NO: 274, complementary to PPOT3)

[0428] The following recode tag-modified proteins employed as binding agents were:

[0429] Anti-phosphotyrosine antibody (Sigma 05-1050X) conjugated to C2-PNA-RT-002 (SEQ ID NO 275):( / 5Phos / TCC AAC ATG TGT TGG GCC TGA TTG TCA GCG +G+A+G +T+G+T +ATT / iUniAmM / TT / 3ddC / )

[0430] Streptavidin conjugated to C2-PNA-RT-003 (SEQ ID NO: 276): ( / 5Phos / TCC AAC ATG TGT TGG CGA GAG CTG TTT CCG +G+A+G +T+G+T +ATT / iUniAmM / TT / 3ddC / )

[0431] Ligation oligonucleotides specific to each cycle were utilized:

[0432] C2-PNA-LO-001 for PPOT2 (SEQ ID NO: 277): (TTT TTC CAT GGA GTG TAG ACA TTG TAC TTT CG)

[0433] C2-PNA-LO-002 for PPOT3 (SEQ ID NO: 278): (TTT TTC CAT GGA GTG TAG AAT AGT TGC GAA TG)

[0434] In this model system, ‘+’ designates a LNA base.

[0435] The beads containing immobilized amino acid complexes from cycles 1 and 5 were processed according to the following protocol. Unless otherwise specified, all wash steps and incubations were performed in buffer comprising PBS supplemented with BSA and Pluronic acid and Tween surfactants. The binding agents were introduced at 10 nM and incubated with the beads for 1 hour at room temperature. Following a wash step, bridge oligonucleotides were added at 10 pM and incubated for 30 minutes at room temperature. After washing, ligation oligonucleotides (10 pM) were combined with T4 DNA ligase and incubated for 1 hour at room temperature to facilitate formation of the recode blocks. The beads were then subjected to a final wash step.

[0436] For analysis, the processed beads were transferred to PCR reaction vessels. Amplification was performed using primers complementary to the terminal sequences of the fully formed recode blocks according to methods known in the art. PCR analysis demonstrated significant product enrichment in samples containing the peptide compared to background controls. Ct values for the recode block representing phosphotyrosine were 25.8 for the condition with peptide PEP6 (SEQ ID NO: 270) and 29.1 for the control without PEP6 (SEQ ID NO: 270), and for the recode block representing the biotin tag was 20.7 for the condition with PEP6 (SEQ ID NO: 270) and 28.4 for the control condition without PEP6 (SEQ ID NO: 270), confirming successful transfer of sequence information from the non-polymerizable PNA cycle tags to amplifiable DNA constructs through the bridge oligonucleotide - mediated ligation process.

[0437] This example validates the core technical capability required for implementing PNA-based combinatorial PUMA codes by demonstrating successful transfer of sequence information from non- polymerizable PNA tags to amplifiable DNA constructs. The bridge oligonucleotide system is useful for enabling efficient conversion of PNA-encoded positional and sequence information into DNA constructs suitable for downstream analysis, as evidenced by successful PCR amplification showing significant product enrichment in peptide-containing samples (Ct values improved by 3.3-7.7 cycles). The demonstrated ability to transfer information from PNA to DNA through bridge oligonucleotide- mediated ligation is useful for directly enabling the implementation of the combinatorial PNA-based PUMA coding system described in Example 8, where each position's PNA sequence must be converted to DNA for analysis. This validation establishes that complex PUMA codes can be constructed using linked PNA sequences while maintaining compatibility with standard DNA sequencing workflows, thereby enabling the generation of large code spaces from relatively small sets of PNA sequences.Example 8: Combinatorial PUMA Coding Using Linked PNA Sets

[0438] This example describes methods for expanding PUMA code space through combinatorial use of linked peptide nucleic acid (PNA) sequences. The method may employ four PNA sequences joined by flexible linkers into a single molecule, where each position can accommodate one of 1000 possible PNA sequences, thereby enabling generation of 1000A4 (1 trillion) unique codes while requiring synthesis of only 4000 distinct PNA sequences.

[0439] Each linked PNA construct comprises four PNA sequences (designated Position A, B, C, and D) joined by flexible linkers of sufficient length to ensure independent accessibility of each sequence. The constructs are synthesized such that each position incorporates one member of its corresponding PNA set (Set A, B, C, or D), where each set may comprise 1000 unique PNA sequences designed according to the following criteria:

[0440] (a) PNA sequences within each set maintain minimum Hamming distance of 3 from all other sequences in the same set (b) Sequences are 8-12 monomers in length (c) Each sequence incorporates universal regions for bridge oligonucleotide binding (d) Sequences are designed to minimize secondary structure formation (e) GC content is maintained between 40-60%

[0441] The linked PNA constructs may be synthesized using standard solid-phase PNA synthesis methods with appropriate linker incorporation between positions. The linkers (e.g., polyethylene glycol units, oligoglycine spacers, or other suitable polymeric linkers) may be designed to provide sufficient spatial separation between PNA sequences to ensure independent accessibility. The constructs may include a single attachment point for conjugation to the TFS.

[0442] Transfer of positional information from PNAs to amplifiable DNA may be accomplished using bridge oligonucleotides according to methods described in Example 7. Four distinct bridgeIl loligonucleotides may be employed, each designed to: (a) Hybridize specifically to PNAs from its corresponding position (b) Present a unique positional identifier sequence (c) Enable ligation-based information transfer to DNA constructs

[0443] The bridge oligonucleotides may facilitate transfer of both sequence and positional information from the PNA codes to DNA through a modified version of the bridge-mediated ligation process described in Example 7. The resulting DNA constructs may comprise:Position-specific identifiers indicating which position contributed each sequence The specific sequence information from each position Universal amplification handles

[0444] In some embodiments, the DNA constructs are amplified using universal primers and sequenced using next-generation sequencing platforms. The sequence data is decon voluted to determine: (a) Which position (A, B, C, or D) contributed each sequence component (b) The specific PNA sequence present at each position (c) The complete combinatorial code representing that TFS

[0445] This combinatorial approach would enable generation of 1000A4 unique PUMA codes while requiring synthesis of only 4000 distinct PNA sequences (1000 per set). The system is readily scalable by increasing either the number of sequences per set or the number of positions in the linked construct. For example, using 5 positions with 1000 sequences per set would yield 1000A5 unique codes while requiring only 5000 distinct PNA sequences.

[0446] In alternative embodiments, the position-specific PNA sequences may be conjugated to separate attachment points on the TFS surface using orthogonal chemistries rather than being pre-linked into a single molecule.

[0447] The high code space achieved through this combinatorial approach enables single-cell and spatial proteomic applications requiring extensive unique identifiers. Error correction is facilitated by the position-specific nature of the code components and the maintained minimum Hamming distance within each set. The system is compatible with existing nucleic acid sequencing workflows through the bridge oligonucleotide-mediated transfer of information to amplifiable DNA constructs.Example 9: Sequential Sequencing of Peptide and Protected PUMA Tag

[0448] A TFS bead can be prepared as described in Example 3, with the modification that the PUMA tag would be synthesized with an N-terminal I-E-G-R sequence (SEQ ID NO: 279), followed by acetylation of the N-terminus. The resulting PUMA tag would have the structure: Ac-I-E-G-R-A-W-S- M-E-T-E-C-H-azide (SEQ ID NO: 280). This design incorporates the factor Xa cleavage site (I-E-G- R, SEQ ID NO: 279) for later deprotection. Alternatively, other protease-cleavable sequences could be used, such as enterokinase (D-D-D-D-K, SEQ ID NO: 281) or TEV protease (E-N-L-Y-P, SEQ ID NO: 282) recognition sites, depending on the specific experimental requirements.

[0449] Proteins are captured and prepared for sequencing as described in Example 4, and the sample is derived from various sources, including cell lysates, tissue homogenates, or purified protein mixtures.Prior to sequencing, the sample may undergo additional preparation steps such as reduction and alkylation of cysteine residues to prevent disulfide bond formation during the sequencing process.

[0450] Reverse translation of the peptide proceeds according to methods described in US provisional no. 63 / 620,076, and PCT / US23 / 70077, with the protected PUMA tag remaining intact during this process. The sequencing chemistry, whether Edman degradation or another method, can be optimized to ensure compatibility with the protection mechanism of the PUMA tag. For Edman-like deconstruction of proteins as described in US provisional no. 63 / 620,076, and PCT / US23 / 70077, reaction with acetic anhydride in pyridine 2:1 for 30 minutes at room temperature effectively renders the N-terminus unreactive and therefore the PUMA tag inert during cycles of phenylisothiocyanate (PITC) exposure.

[0451] Following completing n cycles of deconstruction of the peptide analyte, the peptide analyte may be capped with acetic anhydride as described above, making it inert toward further deconstruction and to preventing interference with subsequent PUMA tag deconstruction. While acylation with acetic anhydride in the presence of N,N-diisopropylethylamine in anhydrous DMF for 30 minutes at room temperature is one option, alternative protection strategies may also be employed. These might include PEGylation of the N-terminus using mPEG-SCM (Creative Pegworks, cat# PJK-208-5g), or the addition of other chemical protecting groups that are stable under the conditions used for PUMA tag deprotection.

[0452] The PUMA tag is activated by incubating in the presence of factor Xa protease (10 units / mL in 20 mM Tris-HCl, pH 8.0, 100 mM NaCl, 2 mM CaC12) for 2 hours at 23°C. For alternative protease- cleavable sequences, the corresponding protease and optimal reaction conditions would be used. Following the protease treatment, thorough washing steps are performed to remove the cleaved protecting group, e.g., Ac-I-E-G-R (SEQ ID NO: 279), and the protease.

[0453] Reverse translation of the peptide proceeds according to methods described in US provisional no. 63 / 620,076, and PCT / US23 / 70077, with the protected peptide analyte remaining intact during this process.. The number of sequencing cycles would be determined by the length of the PUMA tag. In this example, 9 cycles would be sufficient to sequence the entire PUMA tag (A-W-S-M-E-T-E-C-H, SEQ ID NO: 174).

[0454] Recode blocks for both the peptide and PUMA tag are assembled according to methods described in US provisional no. 63 / 620,076, and PCT / US23 / 70077,with cycle information differentiating the information associated with the analyte from the information associated with the PUMA tag. Analyze proceeds as described in Example 4. The analysis software would be configured to recognize the transition point between peptide and PUMA tag sequencing, allowing for separate processing of the two datasets. For the peptide sequence, the software would perform standard de novo sequencing analysis, potentially incorporating reference database matching if appropriate. For the PUMA tag sequence, the software would compare the obtained sequence against a predefined library of PUMA tags to identify the exact tag and its associated metadata.

[0455] The sequential nature of this process allows clear differentiation between the peptide sequence and the PUMA tag sequence, facilitating accurate assignment of metadata to the corresponding peptide. This approach eliminates the need for computational deconvolution of mixed peptide and PUMA tag signals, potentially improving the accuracy and confidence of both the peptide sequencing and metadata assignment.

[0456] Additional quality control steps could be implemented, such as the inclusion of control beads with known peptide and PUMA tag sequences to assess the efficiency of the protection, deprotection, and sequencing steps. Furthermore, multiple technical replicates could be performed to increase confidence in the obtained sequences and to account for any potential sequencing errors or inefficiencies in the protection / deprotection process.

[0457] While this disclosure is satisfied by embodiments in many different forms, as described in detail in connection with preferred embodiments of the disclosure, it is understood that the present disclosure is to be considered as exemplary of the principles of the disclosure and is not intended to limit the disclosure to the specific embodiments illustrated and described herein. Numerous variations may be made by persons skilled in the art without departure from the spirit of the disclosure. The scope of the disclosure will be measured by the appended claims and their equivalents. The abstract and the title are not to be construed as limiting the scope of the present disclosure, as their purpose is to enable the appropriate authorities, as well as the general public, to quickly determine the general nature of the disclosure. In the claims that follow, unless the term “means” is used, none of the features or elements recited therein should be construed as means-plus-function limitations pursuant to 35 U.S.C. §112, '||6.

Claims

CLAIMS1. A single-cell reverse translation method, comprising:(a) partitioning cells of a sample;(b) contacting the partitioned cells with reagents that permeabilize cell membranes of the cells, and release proteins from the cells;(c) providing a trifunctional support (TFS) to the cells, the TFS comprising:(x) a solid support surface,(y) an immobilized peptide unique molecular association (PUMA) tag, and(z) a reactive moiety for binding an amino acid residue of the peptide;(d) coupling a protein of the cells to the TFS;(e) reverse translating the PUMA tag and protein coupled to the TFS to generate a memory oligonucleotide comprising a sequence corresponding to identity and positional information of amino acids of the protein and PUMA tag;(f) obtaining sequence information for the memory oligonucleotide; and(g) based on the obtained sequence information, determining the identity and positional information of the amino acid residue of the protein and assigning metadata to the sequence information that includes information of a cell of origin of the protein.

2. A spatial assay method, comprising:(a) providing a plurality of amino acid identifier conjugates (PUMA-conjugates) each comprising:(x) a peptide, and(y) a reactive moiety capable of coupling to a protein, and optionally(z) an orthogonal reactive moiety capable of joining the PUMA-tagged protein conjugate to a solid support;(b) depositing the plurality of PUMA-conjugates on the solid support;(c) depositing a sample of cells on a surface of the solid support;(d) capturing proteins of the cells with the reactive moiety of the PUMA-conjugates, thereby joining the proteins of the cells with the PUMA-conjugates and generating PUMA-tagged protein conjugates; and(e) reverse translating the PUMA-tagged protein conjugates to obtain a memory oligonucleotide comprising sequences corresponding to identity and positional information of amino acids of the proteins, and corresponding with the PUMA tag;(f) obtaining sequence information for the memory oligonucleotide; and(g) based on the obtained sequence information, determining the identity and positional information of the amino acids of the proteins and assigning metadata that includes its original spatial location within the sample of cells.

3. The method of claim 2, wherein the PUMA-conjugates each comprise the orthogonal reactive moiety.

4. A peptide unique molecular association (PUMA)-conjugate, comprising:(a) a PUMA tag; and(b) a reactive moiety for coupling to an amino acid.

5. The PUMA-conjugate of claim 4, further comprising an orthogonal reactive moiety that binds a solid support.

6. The PUMA-conjugate of claim 4 or 5, wherein use of the PUMA tag provides metadata relating: a molecule, spatial location, relative spatial location, condition, cell, or sample of origin, for molecule(s) that associate with the PUMA tag.

7. The PUMA-conjugate of claim 4 or 5, wherein the PUMA tag comprises repeats of a sequence of amino acids.

8. The PUMA-conjugate of claim 4 or 5, wherein the PUMA tag comprises a constant portion and a variable portion.

9. The PUMA-conjugate of claim 4 or 5, wherein the PUMA tag may comprise 1, 2, or more portions each with the same or a different function.

10. The PUMA-conjugate of claim 4 or 5, wherein the PUMA tag comprises a natural amino acid.

11. The PUMA-conjugate of claim 4 or 5, wherein the PUMA tag comprises a non-natural amino acid.

12. A trifunctional support (TFS), comprising:(x) a solid support surface;(y) immobilized peptide unique molecular association tags (PUMA tags); and(z) reactive moieties associated with the PUMA tags, wherein the reactive moieties each bind an amino acid.

13. The TFS of claim 12, wherein the solid support comprises a hydrogel, a functionalization molecule, or a passivation molecule.

14. The TFS of claim 12, wherein the solid support comprises a bead.

15. The TFS of claim 14, wherein the bead comprises a bead type corresponding with its PUMA tag sequence.

16. The TFS of claim 12, further comprising a moiety that facilitates solubility, a moiety that reduces non-specific interactions, a sequence or moiety that facilitates removal of a PUMA-tagged protein from the solid support surface, such as a cleavable linker, or a combination thereof.

17. A pool of TFSs beads of any one of claims 12-16, comprising a first TFS of a first bead type, and a second TFS of a second bead type.

18. The pool of TFS beads of claim 17, wherein amounts of the first and second TFS are mixed in about equal proportions19. The pool of TFS beads of claim 17, wherein bead types of the first and second TFS have PUMA tag sequences that differ from one another.

20. Use of the TFS of any one of claims 12-16 in a single-cell assay, a spatial assay, or both.

21. Use of the TFS of any one of claims 12-16 to capture proteins from a tissue section for spatial analysis, wherein the TFS lacks a functional group that supports reverse translation, and provides a higher density of capture site or larger size appropriate for protein capture.

Citation Information

Patent Citations

  • Methods and compositions for polypeptide analysis

    US20200348307A1

  • Systems and methods for biomolecule preparation

    US20230212322A1