In-solution methods for determination of protein information by recoding amino acid polymers into DNA polymers

By recoding amino acid polymers into DNA polymers and employing boundary-targeted sequencing, the method addresses inefficiencies in current proteomic analysis, enhancing detection sensitivity and reducing costs for high-throughput protein sequencing.

WO2026080906A1PCT designated stage Publication Date: 2026-04-16ABRUS BIO INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/050604
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-04-30
Filing Date
2025-10-10
Publication Date
2026-04-16

AI Technical Summary

Technical Problem

Current proteomic analysis tools lack sensitivity, accuracy, and efficiency for high-throughput detection of protein sequences and concentrations, particularly for low-abundance molecules, leading to incomplete molecular profiling and high sequencing costs in spatial transcriptomics and single-molecule technologies.

Method used

A method involving chemically-reactive conjugates to recode amino acid polymers into DNA polymers, creating spatially defined zones with multiple oligonucleotide identifiers around analyte binding sites, and a boundary-targeted sequencing approach to reconstruct spatial arrangements, enhancing detection efficiency and reducing sequencing demands.

Benefits of technology

This method improves detection sensitivity and reduces sequencing costs by ensuring high-throughput, accurate, and efficient analysis of protein sequences and concentrations, particularly for low-abundance molecules, while preserving spatial information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025050604_16042026_PF_FP_ABST
    Figure US2025050604_16042026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to compositions of matter, methods, and systems for analyzing polymeric macromolecules, including polymeric macromolecules such as peptides, polypeptides, and proteins. Some such embodiments include recoding the polymeric macromolecule into nucleic acid sequences. Some steps may be performed in solution without a surface anchor.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No: 062954-509001 WOIN-SOLUTION METHODS FOR DETERMINATION OF PROTEIN INFORMATION BY RECODING AMINO ACID POLYMERS INTO DNA POLYMERSRELATED APPLICATIONS

[0001] This application claims the benefit of U.S. provisional application nos. 63 / 706,591, filed October 11, 2024; 63 / 777,516, filed March 25, 2025; and 63 / 797,290, filed April 30, 2025; which are all incorporated herein by reference.SEQUENCE LISTING

[0002] The present application is being filed along with a Sequence Listing in electronic format. The Sequence Listing is provided as a file entitled 063954-509001WO_seqs.xml, created October 10, 2025, which is 50,035 bytes. The information in the Sequence Listing is incorporated by reference in its entirety.FIELD

[0003] The present disclosure relates to compositions of matter, methods, and systems for analyzing polymeric macromolecules, including polymeric macromolecules such as peptides, polypeptides, and proteins, and for enhancing molecular detection and spatial analysis in molecular biology, diagnostics, and single-molecule studies. In some cases, the disclosure pertains to:1) Methods for determining identity and positional information of amino acid residues by recoding amino acid polymers into DNA polymers for subsequent DNA sequencing and analysis;2) Systems that create spatially defined zones containing multiple copies of oligonucleotide identifiers surrounding an analyte binding site to enhance detection efficiency; and3) Methods that reconstruct spatial arrangements of molecular zones or other spatially distributed oligonucleotide features through DNA sequence-based adjacency mapping.BACKGROUND

[0004] Proteins are essential for cellular function, and their sequences and concentrations are key indicators of cell health. Abnormalities in either can signal disease, but current tools for sensitive, accurate, and cost-effective proteome analysis are lacking. Early detection of such irregularities is crucial for diagnosing and treating diseases like cancer. Improved tools for assessing protein sequences and concentrations are needed.

[0005] Tools that provide accurate high-throughput and high-resolution proteomic information may enable the discovery of new biomarkers, provide accurate measurement of low-abundance proteins, identification of post-translational modifications, and better monitoring of proteome dynamics. This progress mayAttorney Docket No: 062954-509001 WO enhance early disease detection, support therapeutic discovery, and improve patient care. Advanced methods and systems for high-throughput, sensitive, and accurate proteomic analysis are needed.

[0006] Single-molecule detection and spatial biology technologies increasingly rely on molecular identifiers to label analytes such as peptides, proteins, nucleic acids, or small molecules. Recent advances in single-molecule analysis have introduced various methods for molecular detection, but they are often limited by the "one molecule finding one molecule" paradigm, which relies on probabilistic interactions that can be inefficient and unreliable. A common approach involves assigning a unique molecular identifier (UMI) to each target molecule; however, this method can suffer from low efficiency due to the probabilistic nature of one-to-one binding events. Many molecules of interest may be missed.

[0007] The probabilistic nature of UMI-based approaches poses challenges for applications requiring high detection efficiency and sensitivity. In single-molecule protein sequencing methods, as demonstrated by Zheng et al. (2024), the transfer of DNA barcodes to individual amino acids cleaved from peptides is inherently limited by the efficiency of a single barcode molecule finding a single amino acid in a local environment. Similar challenges exist in diagnostic assays for rare biomarkers, spatial transcriptomics, and other single-molecule technologies.

[0008] Current technologies, including established spatial transcriptomics platforms such as 10X Genomics Visium, utilize spatially barcoded oligonucleotides but operate on a one-to-one binding principle between analytes and molecular barcodes. This binding paradigm results in inefficient capture, particularly for low-abundance molecules, leading to information loss and incomplete molecular profiling. Methods employing multiple distinct probes per target fail to achieve the efficiency of having multiple identical molecular identifiers surrounding each analyte binding site.

[0009] Current technologies such as clonal cluster generation create dense regions of identical DNA oligonucleotides on surfaces. However, these clusters are not designed to surround a centralized analyte binding site, nor do they define discrete, analyte-centric territories. The oligonucleotide clusters in such systems are largely passive carriers of spatial barcodes without localized functionalization for active molecular interactions or analyte capture.

[0010] Current spatial mapping approaches present limitations when applied to discrete molecular territories with sharp boundaries. DNA microscopy approaches, as developed by Weinstein et al. (2019), treat each oligonucleotide as an individual pixel in the final image. This approach can necessitate sequencing the entirety of the sample to reconstruct spatial relationships, resulting in high sequencing costs, particularly when mapping large homogeneous regions containing identical identifiers. Iterative proximity ligation approaches, as developed by Boulgakov et al. (2018), record interactions between individual oligonucleotides across the entire sample, often requiring exhaustive sequencing of pairwise interactions to reconstruct spatial arrangements.Attorney Docket No: 062954-509001 WO[Oil] These pixel-based approaches face scaling challenges when applied to systems with discrete territories where large regions contain identical molecular identifiers. In such systems, sequencing every oligonucleotide or every possible pairwise interaction typically becomes inefficient, as it generates redundant data from the homogeneous interiors of each territory. An as-yet undeveloped selective approach that focuses sequencing efforts specifically on the boundaries between different territories, where meaningful spatial information resides, would be advantageous.SUMMARY

[0012] The present disclosure relates to or includes compositions of matter, methods, and systems for analyzing polymeric macromolecules, including peptides, polypeptides, and proteins, in a highly-parallel and high-throughput manner via recoding their sequences into DNA polymers.

[0013] Disclosed herein, in some embodiments, are methods for determining amino acid identity and position, comprising: (a) providing a peptide, wherein the peptide is coupled to a solid support such that a N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions; (b) providing a chemically -reactive conjugate, the chemically-reactive conjugate comprising: (x) a location nucleic acid, (y) a reactive moiety for binding the N-terminal amino acid residue of the peptide, and (z) an immobilizing moiety for immobilization to the solid support; (c) contacting the peptide with the chemically-reactive conjugate, thereby coupling the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex; (d) immobilizing the conjugate complex to the solid support via the immobilizing moiety; (e) cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby exposing a next amino acid residue as a second N-terminal amino acid residue on the cleaved peptide and providing an immobilized location complex, the immobilized location complex comprising the cleaved and separated N-terminal amino acid residue and the location nucleic acid; (f) providing a second chemically-reactive conjugate, the second chemically-reactive conjugate comprising: (xx) a cycle tag comprising a cycle nucleic acid associated with a cycle number, and (yy) a second reactive moiety for binding and cleaving the second N-terminal amino acid residue of the peptide; (g) contacting the second N-terminal amino acid residue of the peptide with the second chemically-reactive conjugate, thereby coupling the second chemically-reactive conjugate to the second N-terminal amino acid of the peptide to form a coupled-conjugate-complex; (h) forming a transfer complex comprising the immobilized location complex and the coupled-conjugate- complex, thereby bringing the location nucleic acid into proximity with the cycle nucleic acid; (i) within the transfer complex, joining the location nucleic acid or a reverse complement thereof to the cycle nucleic acid to form a recode unit, or otherwise joining information of the location nucleic acid and the cycle nucleic acid, thereby creating an immobilized recode unit comprising location and cycleAttorney Docket No: 062954-509001 WO information of the location nucleic acid and cycle nucleic acid; (j) cleaving and thereby separating the second N-terminal amino acid residue from the peptide, thereby exposing another next amino acid residue as a third N-terminal amino acid residue on the cleaved peptide and liberating a recode unit complex from the solid support, the liberated recode unit complex comprising the cleaved and separated second N- terminal amino acid residue the recode unit, wherein the recode unit is no longer immobilized; (k) contacting the recode unit complex with a binding agent, the binding agent comprising: a binding moiety for preferentially binding to the second N-terminal amino acid residue of the recode unit complex, and a recode tag comprising a recode nucleic acid corresponding with the binding agent, thereby forming an affinity complex, the affinity complex comprising the recode unit complex and the binding agent, and thereby bringing the recode nucleic acid into proximity with the recode unit; (1) transferring information of the recode nucleic acid to the recode unit to generate a recode block; (m) obtaining sequence information of the recode block; and (n) based on the obtained sequence information, determining identity and positional information of the second amino acid residue of the peptide. In some embodiments, forming the transfer complex comprises hybridizing a portion of the location nucleic acid with a portion of the cycle nucleic acid. In some embodiments, joining the location nucleic acid or a reverse complement thereof to the cycle nucleic acid comprises extending a 3’ end of the cycle tag using the location tag as a template. Some embodiments include repeating (f) through (j) for subsequent amino acids of the peptide. Some embodiments include washing the immobilized location complex before (f) or before (g). Some embodiments include washing the coupled-conjugate-complex before (h). Some embodiments include determining a likely three-dimensional structure of the peptide based on the obtained sequence information. In some embodiments, the location nucleic acid comprises deaza-DNA. In some embodiments, the cycle nucleic acid comprises deaza-DNA. In some embodiments, the recode nucleic acid comprises DNA or RNA. In some embodiments, the recode nucleic acid comprises a hybridization- capable code. In some embodiments, the recode nucleic acid comprises a non-colliding code in an additive vector space. In some embodiments, any of (b)-(j) are performed in the presence of a Lewis acid, and in the absence of trifluoroacetic acid. In some embodiments, the binding moiety comprises a peptide, antibody, antibody fragment, antibody derivative, or aptamer. In some embodiments, the binding moiety binds to a natural amino acid, a post-translationally modified amino acid, a derivatized version of an amino acid, a derivatized or stabilized version of a post-translationally modified amino acid, a synthetic amino acid, an amino acid with a specific side chain, an amino acid with a phosphorylated side chain, an amino acid with a glycosylated side chain, an amino acid with a methylation modification, or a D-amino acid, phenylthiohydantoin (PTH) derivative or anilinothiazolinone (ATZ) derivative of an amino acid, or binds to a combination thereof. In some embodiments, the binding moiety binds covalently or non- covalently to the recode unit complex. In some embodiments, the solid support comprises a bead, a plate,Attorney Docket No: 062954-509001 WO a chip, a glass slide, silica, a resin, a gel, a hydrogel, a membrane, polystyrene, a metal, nitrocellulose, a mineral, plastic, polyacrylamide, latex, or ceramic. In some embodiments, the peptide comprises a hormone, neurotransmitter, enzyme, antibody, viral protein, bacterial protein, synthetic peptide, bioactive peptide, peptide hormone, oligopeptide, polypeptide, fusion protein, cyclic peptide, branched peptide, recombinant protein, tumor marker, therapeutic peptide, antigenic peptide, or signaling peptide. In some embodiments, the peptide is derived from a cell lysate, blood sample, plasma sample, serum sample, tissue biopsy, saliva sample, urine sample, cerebrospinal fluid sample, sweat sample, synovial fluid sample, fecal sample, gut microbiome sample, environmental water sample, soil sample, bacterial culture, viral culture, organoid, tumor biopsy, sputum sample, or hair sample. In some embodiments, the peptide is associated with a disease. In some embodiments, transferring information comprises performing nucleic acid amplification, enzymatic ligation, splint ligation, chemical ligation, template-assisted ligation, use of a ligase enzyme, use of a splint oligonucleotide, use of a catalyst, use of a bridging molecule, use of a condensation agent, use of a coupling reagent, use of a polymerase enzyme, use of a complementary nucleic acid sequence, use of a nicking enzyme, use of a nucleic acid modifying enzyme, use of a recombinase, use of a strand-displacing polymerase, use of a single-strand binding protein, a click chemistry reaction, a phosphodiester bond formation, or a peptide nucleic acid-mediated ligation. In some embodiments, the information of the recode nucleic acid comprises a sequence of the recode nucleic acid or a reverse complement of the sequence of the recode nucleic acid. In some embodiments, transferring information comprises joining the recode nucleic acid or a reverse complement of the recode nucleic acid with the cycle nucleic acid. In some embodiments, the second N-terminal amino acid is identified as amino acid position 2, 3, 4, 5, 6, 7, 8, 9, 10, or grearter, within the peptide, based on sequenced information of the cycle nucleic acid. In some embodiments, the second chemically reactive conjugate includes a moiety for joining to a second solid support.

[0014] Disclosed herein, in some embodiments, are methods for determining identity and positional information of a plurality of amino acid residues of a peptide, the peptide comprising n amino acid residues, the method comprising: (a) coupling the peptide to a solid support such that a N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions; (b) providing a chemically -reactive conjugate, the chemically-reactive conjugate comprising: (x) a location nucleic acid, (y) a reactive moiety for binding and cleaving the N-terminal amino acid residue of the peptide and exposing a next amino acid residue as a N-terminal amino acid residue on the cleaved peptide, and (z) an immobilizing moiety for immobilization to the solid support; (c) contacting the peptide with the chemically- reactive conjugate, thereby coupling the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex; (d) immobilizing the conjugate complex to the solid support via the immobilizing moiety; (e) cleaving and thereby separating the N-terminal amino acid residue from theAttorney Docket No: 062954-509001 WO peptide, thereby exposing the next amino acid residue as a N-terminal amino acid residue on the cleaved peptide and providing an immobilized location complex, the immobilized location complex comprising the cleaved and separated N-terminal amino acid residue and the location nucleic acid; (f) providing a second chemically-reactive conjugate comprising: (xx) a cycle tag comprising a cycle nucleic acid associated with a cycle number, and (yy) a second reactive moiety for binding and cleaving the N-terminal amino acid residue of the peptide and exposing a next amino acid residue as an additional N-terminal amino acid residue on the cleaved peptide, and optionally (zz) a moiety for joining to an optional second solid support; (g) contacting the peptide with the second chemically-reactive conjugate, thereby coupling the second chemically-reactive conjugate to the additional N-terminal amino acid of the peptide to form a coupled- conjugate-complex; (h) forming one or more transfer complexes, each transfer complex comprising an immobilized location complex and a coupled-conjugate-complex, thereby bringing a location nucleic acid into proximity with a cycle nucleic acid within each formed transfer complex; (i) within each formed transfer complex, joining a location nucleic acid or a reverse complement thereof to a cycle nucleic acid to form an immobilized recode unit, or otherwise joining information of the location nucleic acid and the cycle nucleic acid, thereby creating one or a plurality of immobilized recode units, each recode unit corresponding with a formed transfer complex; (j) cleaving and thereby separating the additional N-terminal amino acid residue from the peptide, thereby exposing the next amino acid residue as a further N-terminal amino acid residue on the cleaved peptide and liberating the immobilized recode unit complex from the solid support, the liberated recode unit complex comprising the cleaved and separated additional N-terminal amino acid residue, cycle and location information; (k) repeating (f) through (j) n-1 times to liberate and collect pools having a plurality of recode unit complexes, each additional plurality of recode unit complexes comprising information associated with cycle 2 to n, accordingly; (1) contacting the collection of recode units complexes with binding agents, the binding agents comprising: a binding moiety for preferentially binding to one, or to a subset, of the recode unit complexes, and a recode tag comprising a recode nucleic acid corresponding with the binding agent, thereby forming one or more affinity complexes, each affinity complex comprising an recode unit complex and the binding agent, thereby bringing a recode nucleic acid into proximity with a recode tag within each formed affinity complex; (m) within each formed affinity complex, joining a recode unit or a reverse complement thereof to a recode tag to form a recode block, or otherwise transferring information of the recode tag to the recode unit complex, thereby creating a plurality of recode blocks, each recode block corresponding with a formed affinity complex; (n) optionally, joining two or more members of the plurality of recode blocks to form a memory oligonucleotide; (o) obtaining sequence information for the recode blocks or memory oligonucleotides; and (p) based on the obtained sequence information, determining identity and positional information of a plurality of amino acid residues of the peptide. In some embodiments, (f)-(j) are repeated 2, 3, 4, or more times. In some embodiments, n isAttorney Docket No: 062954-509001 WO an integer greater than or equal to 2. In some embodiments, each binding agent comprises recode tags with a unique nucleic acid sequence. In some embodiments, a plurality of binding agents comprises recode tags with the same nucleic acid sequence. In some embodiments, the binding agents comprises recode tags which have a unique sequence portion and a common sequence portion. In some embodiments, the binding agents are contacted with combined pool(s) of recode unit complexes. The method of claim 25, further comprising washing the immobilized amino acid complex before said contacting the immobilized amino acid complex with a coupled-conjugate-complex. Some embodiments include washing the coupled-conjugate-complex before said contacting the immobilized amino acid complex with a coupled-conjugate-complex. Some embodiments include determining a likely three-dimensional structure of the peptide based on the sequence information. In some embodiments, the recode nucleic acid comprises a hybridization-capable code. In some embodiments, the recode nucleic acid comprises DNA. In some embodiments, the recode nucleic acid comprises non-colliding codes in an additive vector space. In some embodiments, the location nucleic acid comprises DNA. In some embodiments, the cycle nucleic acid comprises DNA. In some embodiments, the peptide comprises a hormone, neurotransmitter, enzyme, antibody, viral protein, bacterial protein, synthetic peptide, bioactive peptide, peptide hormone, oligopeptide, polypeptide, fusion protein, cyclic peptide, branched peptide, recombinant protein, tumor marker, therapeutic peptide, antigenic peptide, or signaling peptide. In some embodiments, a plurality of peptides are immobilized to the solid support and analyzed in concert. In some embodiments, obtaining the sequence information for the memory oligonucleotide comprises performing sequencing. In some embodiments, obtaining the sequence information for the memory oligonucleotide comprises melt curve analysis for multi-stage encoding. In some embodiments, the binding moiety comprises an antibody or a fragment thereof, or an aptamer. In some embodiments, the binding moiety is covalently bound or non-covalently bound to the amino acid. In some embodiments, the binding moiety binds to a natural amino acid, a derivatized amino acid, a synthetic amino acid, or a D- amino acid. In some embodiments, the binding moiety binds to a natural amino acid, a post-translationally modified amino acid, a derivatized version of an amino acid, a derivatized or stabilized version of a post- translationally modified amino acid, a synthetic amino acid, an amino acid with a specific side chain, an amino acid with a phosphorylated side chain, an amino acid with a glycosylated side chain, an amino acid with a methylation modification, or a D-amino acid, phenylthiohydantoin (PTH) derivative or anilinothiazolinone (ATZ) derivative of an amino acid, or binds to a combination thereof. In some embodiments, the solid support comprises a bead, a plate, or a chip. In some embodiments, the solid support comprises glass slide, silica, a resin, a gel, a hydrogel, a membrane, polystyrene, a metal, nitrocellulose, a mineral, plastic, polyacrylamide, latex, or ceramic. Some embodiments include protecting or deprotecting the location tag at any appropriate step of the workflow between (b) and (k). In some embodiments, transferring information comprises performing nucleic acid amplification, enzymatic ligation, splintAttorney Docket No: 062954-509001 WO ligation, chemical ligation, template-assisted ligation, use of a ligase enzyme, use of a splint oligonucleotide, use of a catalyst, use of a bridging molecule, use of a condensation agent, use of a coupling reagent, use of a polymerase enzyme, use of a complementary nucleic acid sequence, use of a nicking enzyme, use of a nucleic acid modifying enzyme, use of a recombinase, use of a strand-displacing polymerase, use of a single-strand binding protein, a click chemistry reaction, a phosphodiester bond formation, or a peptide nucleic acid-mediated ligation. In some embodiments, the information of the recode nucleic acid comprises a sequence of the recode nucleic acid or a reverse complement of the sequence of the recode nucleic acid. In some embodiments, transferring information comprises joining the recode nucleic acid or a reverse complement of the recode nucleic acid with the cycle nucleic acid. In some embodiments, determining the identity and positional information of the plurality of amino acid residues of the peptide comprises determining the identity and positional information of all of the amino acid residues of the peptide. In some embodiments, determining the identity and positional information of the plurality of amino acid residues of the peptide comprises determining the identity and positional information of only a subset of the amino acid residues of the peptide. Some embodiments include identifying the peptide by comparing the identity and positional information of the plurality of amino acid residues to a database.

[0015] Disclosed herein, in some embodiments, are kits for determining identity and positional information of an amino acid residue of a peptide, comprising: a chemically -reactive conjugate comprising (a) a nucleic acid sequence tag and (b) a reactive moiety that couples to a N-terminal amino acid residue of a peptide, and thereby forms a conjugate complex comprising the chemically -reactive conjugate coupled to the N-terminal amino acid of the peptide; a binding agent comprising: a binding moiety for preferentially binding to the conjugate complex and a recode tag comprising a recode nucleic acid corresponding with the binding agent; and a reagent for transferring information of the recode nucleic acid to the cycle nucleic acid of the conjugate complex to generate a recode block.

[0016] Disclosed herein, in some embodiments, are chemically-reactive conjugates (CRCs) comprising or consisting of: (A) a nucleic acid sequence tag; and (B) a reactive moiety for binding and cleaving a N- terminal amino acid residue from a peptide or a cleavable derivative thereof. Some embodiments relate to a CRC represented by Formula III: l-ABA- B(Formula III),Attorney Docket No: 062954-509001 WO wherein A comprises a cycle tag, and B comprises a reactive moiety, and LAB comprises a linker with option central moiety (C). Some embodiments include a cleavable group between (A) and (B), between(B) and (C), between (A) and (C), between (A) and (B+C), between (B) and (A+C), or between (C) and (A+B), or any combination thereof. Some embodiments include a cleavable group between (A) and (B), between (B) and (C), between (A) and (C), between (A) and (B+C), between (B) and (A+C), or between(C) and (A+B), or any combination thereof. Some embodiments include a cleavable group between (A) and (B), between (B) and (C), or a combination thereof. In some embodiments, the reactive moiety comprises a phenyl isothiocyanate (PITC), an isothiocyanate (ITC), a dansyl chloride, a dinitrofluorobenzene (DNFB), an enzyme or peptide, or a combination or derivative thereof. In some embodiments, the reactive moiety specifically cleaves at a specific amino acid. In some embodiments, the reactive moiety cleaves more than a single amino acid or motif. In some embodiments, the immobilizing moiety comprises a protected thiol group, a protected amine group, or a carboxyl group, an azide, an alkyne, an alkene, an aryl boronic acid, an aryl halide, a haloalkyne, an acryl, a silylalkyne, a Si-H group, a protected or photoprotected reactive group, or a photoactivated reactive group. In some embodiments, the nucleic acid sequence tag is generated upon conjugating the nucleic acid sequence to a group for attaching a nucleic acid sequence comprising a protected oxyamine group, a protected thiol, a protected amine, a protected hydrazine, a tetrazine, an azide, an alkyne, an alkene, a trans-cyclooctene, a DBCO, a bicyclononyne, a norbornene, a strained alkyne, or a strained alkene, or a derivative thereof. In some embodiments, the reactive moiety comprises a group on the CRC for attaching to a cleavable derivatized N-terminal amino acid, comprising a tetrazine, an azide, an alkene, an alkyne, a trans-cyclooctene, a DBCO, a bicyclononyne, a norbornene, a strained alkyne, or a strained alkene, or a derivative thereof.

[0017] In a further aspect, disclosed herein is a method, comprising: removing an N-terminal amino acid from a protein in the presence of tris(pentafluorophenyl)borane (BCF). In some embodiments, removing the N-terminal amino acid from the protein is performed in the absence of trifluoroacetic acid (TFA). In some embodiments, removing the N-terminal amino acid from the protein comprises contacting the protein with a chemically reactive conjugate (CRC) comprising an Edman degradation reagent. In some embodiments, the CRC comprises a nucleic acid tag. In some embodiments, the nucleic acid tag comprises deaza-nucleotides. In some embodiments, the nucleic acid tag maintains its integrity. In some embodiments, the method further comprises sequencing the nucleic acid tag.Molecular Neighborhood Technologies

[0018] This disclosure also provides methods and compositions for creating spatially organized molecular territories that address fundamental detection efficiency limitations in single-molecule analysis.Attorney Docket No: 062954-509001 WO

[0019] In one aspect, the present disclosure provides methods and compositions for creating discrete molecular territories, herein referred to as ULI neighborhoods or ULI territories, comprising a central analyte binding site surrounded by multiple copies of a unique location identifier (ULI) oligonucleotide. This configuration directly addresses the low capture efficiency problem in conventional "one molecule finding one molecule" approaches by increasing the local concentration of identifying barcodes around each target molecule by 100- to 1000-fold (e.g., from one barcode to hundreds or thousands of local copies). Where prior art methods might miss a target due to the probabilistic nature of one-to-one binding, some embodiments in the present disclosure surround each target with hundreds to thousands of identical ULI copies, dramatically increasing the probability of successful detection even for low-abundance analytes. Where clonal amplification or bead-based systems simply increase the number of barcodes per zone, some embodiments in the present disclosure combine dense local barcoding with a precisely positioned analyte capture site. This central-peripheral design is useful for providing more efficient capture.

[0020] In alternative embodiments, the disclosure provides methods for incorporating multiple analyte binding sites within a single ULI neighborhood. One approach for achieving this may involve secondary modification of the initial analyte binding site with multi-functional linkers, such as multi-azide compounds that facilitate attachment of multiple analytes to a single primary binding site while maintaining the association with a single ULI territory. Other approaches may include branched DNA structures, dendrimeric extensions, or alternative chemistries that increase the valency of binding sites while preserving the spatial organization of ULI territories. These various strategies may facilitate multiplexed detection or higher analyte density within each defined molecular zone without altering the fundamental ULI neighborhood architecture.

[0021] The ULI neighborhood approach employs a novel bridge amplification strategy utilizing nonstandard nucleotides (iso-cytosine and iso-guanine) to precisely control amplicon endpoints. This controlled amplification process creates discrete territories. Some embodiments provide for use of a sparse distribution of analyte binding sites, each surrounded by its own defined molecular neighborhood containing multiple copies of zone-specific ULIs, addressing the challenge of precisely associating molecular identifiers with specific analyte binding sites.

[0022] In some aspects, ULI territories may be used to purify and tag analytes that have a common chemical moiety, such as an azide, a thiol, an amine, a carboxyl group, a hydroxyl group, an aldehyde, a ketone, an alkyne, a strained alkyne, a cyclooctyne, a norbomene, a tetrazine, a trans-cyclooctene, a maleimide, a hydrazine, an oxyamine, a biotin, a peptide tag, a His-tag, a FLAG-tag, a Myc-tag, combinations thereof, or other chemical moieities used for purification

[0023] In some aspects, a priori specified analyte molecules may be directed specifically to one ULI territory, or a class of ULI territories, where the ULI territories include a commonality, or common feature.Attorney Docket No: 062954-509001 WO

[0024] In another aspect, the disclosure provides a method for spatial reconstruction of molecular arrangements using junction-spanning DNA constructs that link adjacent oligonucleotide-defined territories. By focusing sequencing specifically on the informative boundaries between ULI territories, rather than the redundant interiors, the method can reduce sequencing requirements by 80% or more while preserving essential spatial information. This boundary-targeted approach directly solves the problem of sequencing inefficiency in existing DNA microscopy based spatial mapping technologies when applied to systems with large homogeneous regions.

[0025] The method utilizes dynamic hybridization regions and directional bridge oligonucleotides to selectively capture and encode adjacency relationships between territories into sequence data, allowing reconstruction of spatial layouts via computational graph analysis. Unlike standard DNA microscopy, which may rely on uncontrolled diffusion or stochastic ligation events to capture all molecular adjacency, some embodiments employ a deterministic coding scheme. It uses pre-assigned, orthogonal Zone Interaction Pairing (ZIP) code pairs embedded into ULI oligonucleotides, and directional bridge oligonucleotides that enforce orientation-specific ligation. This design ensures predictable, structure-aware encoding of adjacency between territories. This selective junction-based sequencing approach overcomes the fundamental limitation of pixel-based methods that treat each oligonucleotide as an individual point in space. This spatial mapping method is compatible with, but independent of, ULI neighborhoods and may be applied to alternative spatial systems such as random oligo deposition, lithographically patterned arrays, or surface-encoded molecular zones.

[0026] Unlike DNA microscopy (Weinstein et al.) or iterative proximity ligation (Boulgakov et al.), which treat each oligonucleotide as a unique spatial pixel and thus can require exhaustive sequencing, some embodiments’ spatial mapping approach selectively targets junctions between discrete territories. This boundary -targeted sequencing directly addresses the fundamental scaling problem in pixel-based methods when applied to systems with large homogeneous regions.

[0027] Some embodiments employ orthogonal ZIP codes embedded into ULI oligonucleotides, and directional bridge oligonucleotides that selectively hybridize across adjacent territories. These bridges enforce unidirectional encoding of adjacency, ensuring each territorial junction is captured without capturing interior region pairings of the same ULIs, and eliminate redundant ligation products that consume sequencing bandwidth in existing proximity ligation schemes.

[0028] Because spatial information is encoded only at zone junctions, sequencing demand scales with the number of territorial boundaries, not total oligonucleotide count, offering exponential scalability and providing for spatial reconstruction even in large, complex samples.

[0029] Together, these two technologies can form a unified framework that addresses two core limitations of existing systems: (1) the inefficiency of “one molecule finding one molecule” is overcome by ULIAttorney Docket No: 062954-509001 WO clustering, and (2) the sequencing inefficiency of pixel-based spatial mapping is resolved through junctionspecific detection focused on information-rich boundaries.

[0030] In further aspects, the disclosure provides methods for modifying ULI-containing oligonucleotides with strategically placed ZIP codes to provide boundary detection, for generating junction-spanning amplicons through directional extension and ligation, and for purifying these products via size-selection or sequence-specific enrichment.

[0031] Computational methods are also disclosed for reconstructing spatial arrangements from ZIP-linked junction sequences using spring layout algorithms adapted for discrete molecular territories. These analytical tools extract maximal spatial information from sparse junction data, overcoming the limitations of algorithms built for exhaustive, pixel-level datasets.

[0032] The disclosure offers a set of substantial advantages over prior spatial and molecular detection technologies:

[0033] Enhanced detection efficiency: Unlike existing spatial transcriptomics or barcoded bead arrays that depend on inefficient one-to-one binding events, the ULI neighborhood framework ensures that each analyte is surrounded by a dense cluster of identical molecular identifiers, dramatically improving capture sensitivity and reducing dropout.

[0034] Scalable spatial mapping: The ZIP-based spatial reconstruction approach sequences only the informative boundaries between zones. This reduces sequencing burden by orders of magnitude, scaling with junction count rather than total oligonucleotide number.

[0035] Orientation-specific mapping: The use of orthogonal, directionally encoded ZIP codes may facilitate precise, unambiguous reconstruction of spatial relationships, unlike prior undirected ligation systems.

[0036] Multi-analyte capacity: Some embodiments uniquely support multiplexed detection within a single spatial territory while maintaining a shared ULI identifier, extending functionality beyond what is contemplated in existing spatial barcoding methods.

[0037] Together, these innovations solve critical challenges in single-molecule spatial biology: boosting detection probability, minimizing sequencing demand, and facilitating robust, sequence-based reconstruction without optical imaging

[0038] Disclosed herein, is a method of producing a surface that presents spatially discrete molecular zones, the method comprising: (a) attaching first and second anchor oligonucleotides to a solid support; (b) extending a subset of the first anchor oligonucleotides with a template oligonucleotide that comprises a universal hybridization site (UHS) sequence, a unique location identifier (ULI) sequence, and a first primer sequence, thereby defining analyte binding locations; (c) performing bridge amplification using the first primer sequence and a second primer sequence on the second anchor oligonucleotides, whereby multipleAttorney Docket No: 062954-509001 WO copies of the ULI are generated around each analyte binding location to form a discrete molecular zone; and (d) modifying the analyte binding location (e.g. oligonucleotide) to permit covalent attachment of an analyte. In some embodiments, the template of step (b) comprises iG bases, iC bases, neither iC nor iG, or both iC and iG. In some embodiments, the anchor oligonucleotides do not comprise adenine or guanine bases. In some embodiments, the bridge amplification reaction solution does not comprise iso dCTP. In some embodiments, thr oligonucleotide of step (b) is supplied at <0.5 mol % of the total extension reaction oligonucleotides. In some embodiments, the chemical modification introduced in step (d) comprises incorporation of an alkyne modified deoxycytidine.

[0039] In a further aspect, dislclosed herein is a method of producing a surface that presents spatially discrete molecular zones, the method comprising: (a) attaching template oligonucleotides that comprise a universal hybridization site (UHS) sequence, a unique location identifier (ULI) sequence, and a first primer sequence, and thereby define analyte binding locations, and template oligonucleotides that comprises a second primer sequence capable to support bridge amplification in combination with the first primer sequence; (b) performing bridge amplification of the template oligonucleotides, whereby multiple copies of the ULI are generated around each analyte binding location to form a discrete molecular zone; and (c) modifying the analyte binding location (e.g. oligonucleotide) to permit covalent attachment of an analyte.

[0040] In a further aspect, disclosed herein is a surface comprising a plurality of discrete molecular zones, each zone containing a single analyte binding site and two or more copies of a zone specific ULI oligonucleotide, the ULI oligonucleotides of different zones being sequence distinguishable from one another. In some embodiments, the ULI oligonucleotides in each zone are bounded by an iCiC motif that is unpaired in the absence of iso dGTP. In some embodiments, each zone further comprises upstream and downstream ZIP codes that flank the ULI sequence, at least 80 % of adjacent zones differing in their ZIP code pair. In some embodiments, the analyte binding site comprises an alkyne group suitable for copper catalyzed azide-alkyne cyclo addition.

[0041] Also disclosed herein ia a method for determining adjacency relationships between discrete molecular zones of clonal ULIs, the method comprising: (a) copying, in each zone, one of said ZIP flanked oligonucleotides to generate a strand that terminates in a restriction site (RS) and is blocked at a modified nucleotide motif; (b) annealing and ligating a bridge oligonucleotide whose termini are complementary to unlike ZIP codes on neighbouring zones, thereby forming a junction spanning strand that contains the ULIs of two adjacent zones; (c) amplifying the junction spanning strand with primers that bind to adapter sequences distal to the ZIP codes to generate a junction amplicon; (d) removing the amplified junction amplicon from the surface; and (e) sequencing the junction amplicon and recording the two ULIs therein as adjacent nodes in a spatial graph. In some embodiments, the ZIP codes are 4-7 nucleotides in length and comprise at least one LNA base. In some embodiments, the bridge oligonucleotide carries a 5' phosphateAttorney Docket No: 062954-509001 WO and internal abasic spacers that suppress non specific hybridization to conserved primer regions. In some embodiments, the amplification of step (c) is carried out with primers complementary to CT1 and CT2 adapter sequences. In some embodiments, after step (d), the residual surface bound strand is digested at the RS site with a Type IIS restriction endonuclease to regenerate the primer site for a subsequent mapping cycle. In some embodiments, the spatial graph is laid out with a spring embedding algorithm in which edge weights are proportional to junction amplicon read counts. In some embodiments, machine learning inference is applied to the spatial graph to refine node positions using a training set of reference layouts.

[0042] In a further aspect, disclosed herein ia an oligonucleotide composition comprising N distinct oligonucleotide species. In some embodiments, each species comprises: a first region containing a subset of k-mers selected from a pool of N+l distinct k-mer sequences, wherein exactly one k-mer from said pool is absent from said first region, and a second region comprising the reverse complement of said absent k- mer, wherein any two different oligonucleotide species in said composition are capable of forming a heteroduplex via complementary k-mer pairing, and wherein no oligonucleotide species is capable of forming a homoduplex with itself.BRIEF DESCRIPTION OF THE DRAWINGS

[0043] FIG. 1 illustrates a diagram of an exemplary workflow for analyzing polymeric macromolecules, including polymeric macromolecules such as peptides, and proteins, according to embodiments of the present disclosure.

[0044] FIG. 2 schematically illustrates various operations of the recoding Process 200, according to embodiments of the present disclosure.

[0045] FIG. 3 schematically illustrates a mode of information transfer between an immobilized location complex and a coupled-conjugate-complex, during the operations of Process 200, according to embodiments of the present disclosure.

[0046] FIG. 4 schematically illustrates a mode of information transfer between a recode unit and a recode tag, during the operations Process 200, according to embodiments of the present disclosure.

[0047] FIG. 5 schematically illustrates assembly of information into an amplicon for NGS sequencing, during the operations Process 200, according to embodiments of the present disclosure.

[0048] FIG. 6 schematically illustrates the components and steps of isothermal assembly and amplification of protein sequence information, according to embodiments of the present disclosure.

[0049] FIG. 7 provides a 2-dimensional perspective of a solid support on which a plurality of immobilized macromolecules (filled circles) and a plurality of immobilized single amino acid complexes having location tags (open circles) reside, according to embodiments of the present disclosure. The overlapping areasAttorney Docket No: 062954-509001 WO present the opportunity for interaction between cycle tags and location nucleic acids during the operations of transferring information between cycle tags and location tags.

[0050] FIG. 8 schematically illustrates various operations of the recoding Process 300, according to embodiments of the present disclosure. Concatenation of recode units provides various advantages and efficiencies when obtaining sequence information via nanopore readout to determine identity and positional information of a plurality of amino acid residues of a plurality of peptides.

[0051] FIG. 9 schematically illustrates various operations of the recoding Process 400, according to embodiments of the present disclosure. In solution indexing provides various advantages and efficiencies when obtaining sequence information to determine identity and positional information of a plurality of amino acid residues of a plurality of peptides.

[0052] FIG. 10 depicts a variation of above processes, according to embodiments of the present disclosure.

[0053] FIG. 11 depicts a variation of above processes, according to embodiments of the present disclosure.

[0054] FIG. 12 depicts a variation of above processes, according to embodiments of the present disclosure.In step (o) TCE location may be at any position appropriate to allow time signal averaging of the AA component of the polymer - note that the polymer is not polymerizable (e.g., it could be added as a component of the assembly sequence. It could be added via a golden gate process where the incorporated base comprises a TCE. It could be a part of the cycleTag. It could be added during the copy step. It could be a chemical entity of the CRC.

[0055] FIG. 13 depicts a variation of above processes, according to embodiments of the present disclosure.

[0056] FIG. 14 depicts a variation of above processes, according to embodiments of the present disclosure.

[0057] FIG. 15 depicts a variation of above processes, according to embodiments of the present disclosure.

[0058] FIG. 16 depicts a schematic representation of spatially organized ULI neighborhoods on a surface. Different shades represent distinct ULI zones, each containing a central analyte binding site (large circle) surrounded by multiple copies of zone-specific ULI oligonucleotides (small circles). The natural boundaries formed between adjacent territories create a Voronoi-like pattern where each analyte binding site is surrounded by its unique molecular neighborhood.

[0059] FIG. 17 depicts a carboxyl-functionalized surface loaded with two distinct cytosine-thymine (CT) oligonucleotides in approximately equal proportions. The CT1 oligonucleotide ( / 5AmMC6 / -{CTl } -U) terminates with uracil, providing sites for subsequent enzymatic cleavage, while the CT2 oligonucleotide ( / 5AmMC6 / -{CT2}) provides complementary anchoring sites.

[0060] FIG. 18 depicts the three components used in the single extension reaction: (1) H5- / iC / iC / -A- {CT1'}2O at standard concentration, (2) / iG / iG / -H7'-{CT2'}2o at standard concentration, and (3) H7- / iC / iC / -UHS-ULLH5-RS'-{CT2'}2o at substantially lower concentration (e.g., 0.1%). The differential concentration creates a sparse distribution of potential analyte binding sites across the surface. AlthoughAttorney Docket No: 062954-509001 WO the low-concentration seeding approach produces broadly sparse distributions, Poisson-based variability may result in local regions of slightly higher or lower analyte binding site density. This variability can be acceptable in many applications, but optimization of loading concentrations may be beneficial depending on a possibly required spatial uniformity.

[0061] FIG. 19 depicts the surface after extension reaction completion, showing: (1) Multiple copies of CTl-U-iGiG-H5' (resulting from extension of CT1 with the first component), (2) Multiple copies of CT2- H7-iCiC (resulting from extension of CT2 with the second component), and (3) Sparse distribution of CT2-RS-H5'-ULI'-UHS'-iGiG-H7' (resulting from extension of CT2 with the third component), which define the location of the analyte binding sites.

[0062] FIG. 20 depicts the first bridge amplification cycle that copies the ULI to a proximal H7 anchor. The bridge amplification occurs between the analyte binding site strand and nearby H7 anchors on the surface, creating the first copy of the ULI in the zone surrounding the potential analyte binding site.

[0063] FIG. 21 depicts endonuclease application (e.g., BbsLHF) to cleave at the restriction site, removing the analyte binding site from participation in subsequent bridge amplification steps. The enzyme leaves a specific adenine-guanine-rich overhang that facilitates subsequent specific extension.

[0064] FIG. 22 depicts continued bridge amplification between the H5' and H7 anchors, using a nucleotide mixture that may include standard C / T nucleotides, deazaA / G nucleotides, and iso-guanine (isoG) nucleotides, but deliberately excludes iso-cytosine (isoC) nucleotides. This specific mixture ensures extension termination at iGiG stretches, creating defined ULI territory boundaries.

[0065] FIG. 23 depicts an Uracil-Specific Excision Reagent (USER) enzyme application to cleave at uracil sites, removing the non-target strand after amplification. This cleavage simplifies the molecular architecture in preparation for analyte binding site modification. A complementary strand may be annealed to the uracil-containing strand to provide for USER cleavage.

[0066] FIG. 24 depicts each ULI zone after USER cleavage, containing: (1) one copy of CT2-RSTail (the analyte binding site), (2) multiple copies of CT2-H7-iCiC-UHS-ULI-H5, and (3) multiple copies of CT1 (resulting from USER cleavage).

[0067] FIG. 25 depicts the secondary extension from the RSTail to introduce an alkyne modification to the analyte binding site. In one embodiment, an extension oligo such as GGAAAAAA-RsTail'-CT2' may be used to incorporate alkyne-modified cytosines; however, alternative sequences compatible with the extension primer may also be used. Incorporating alkyne-modified cytosines is suitable for subsequent conjugation to peptides or other analytes through click chemistry.

[0068] FIG. 26 depicts the mathematical foundation for spatial mapping via ULI territories with an exemplary layout. Panel A shows discrete ULI neighborhoods forming Voronoi-like regions with analyte binding sites at their centers. Panel B illustrates the Delaunay triangulation that connects adjacentAttorney Docket No: 062954-509001 WO territories through junction detection. Panel C demonstrates how this junction information is converted into a graph and then reconstructed into the original spatial arrangement using spring layout algorithms.

[0069] FIG. 27 depicts an exemplary ULI-ULI Neighborhood Junction. The diagram depicts an exemplary boundary between two distinct ULI territories (ULLA and ULLB). Information regarding adjacent neighborhoods (e.g., capturing ULLA to ULI-B junctions in one amplicon) contributes critically to mapping the spatial location of all ULI neighborhoods. In contrast, information from the homogeneous interior of neighborhoods (e.g., ULLA adjacent to another ULLA oligonucleotide) provides redundant data that does not contribute to spatial mapping. This selective boundary-focused approach can dramatically reduce sequencing requirements while preserving essential spatial information.

[0070] FIG. 28 depicts an exemplary workflow for selective ULLULI junction formation and adjacency encoding using ZIP codes and directionally constrained ligation. Each ULI neighborhood includes two distinct directional ZIP code sequences (ZIPa and ZIPb) positioned to facilitate asymmetric bridging with adjacent territories. In the first step, a polymerase extension reaction is initiated from an anchored primer (e.g., H5a') to copy the ULI region into a complementary strand; this reaction halts at an internal isocytosine (iCiC) motif due to the absence of iso-guanine in the reaction mixture. Next, the ZIPb region from an adjacent ULI hybridizes across the boundary and is ligated to the extended strand, completing the first half of the junction. Simultaneously, the complementary ZIPa region hybridizes to the target strand but cannot be ligated across the iCiC block. A second polymerase extension is initiated from a downstream UHS' primer, with iso-guanine now present in the nucleotide pool, allowing synthesis across the iCiC motif to produce an iso-guanine-rich blocking sequence (iGiG). Final ligation across the now- extended ZIPa region completes the junction, generating a single DNA amplicon that contains both ULI codes in a directionally resolved format suitable for spatial adjacency reconstruction.

[0071] FIG. 29A-30B depict an alternative short-ZIP boundary-mapping workflow. FIG. 29A: Surface composition immediately after the final bridge-amplification cycle and before USER treatment. Neighbouring ULI territories display surface-anchored CT1 / CT2 primer pairs; each ULI is flanked by ultra-short ZIP codes (ZA for ULIA, ZB for ULIB). FIG. 29B: A single primer pool (Zn-UHS-RS) primes from ZIPA' or ZIPB', copying the ULI, installing iso-cytosine residues, and embedding a downstream restriction-site (RS). USER enzyme nicks the original CT1 strand, and a brief T7 exonuclease digestion removes the clipped fragment, leaving the newly synthesised strand hybridised to the surface via its AG adapter. FIG. 30A: Resulting surface architecture after USER cleavage and exonuclease clean-up: every territory now presents a 3'-blocked, RS-tagged copy of its own ULI adjacent to the untouched complementary strand from its neighbour. FIG. 30B: Directional bridge oligonucleotides carrying complementary ZIP motifs anneal and are ligated only at boundaries whereAttorney Docket No: 062954-509001 WO different ZIP codes (e.g., ZA and ZB) meet, generating single-strand chimeras that fuse ULIA to ULIB and thereby encode the spatial adjacency.

[0072] FIG. 31 depicts an alternative ZIP boundary mapping method using complementary stretches of oligonucleotides

[0073] FIG. 32 depicts gel electrophoresis analysis of post-extension amplicons, demonstrating transfer of sequence information between co-immobilized cycle tags and location nucleic acids.

[0074] FIG. 33 depicts multi-cycle memory oligonucleotide formation whereby 1 to 5 recode blocks are incorporated using a heat denaturation method and amplified via PCR. Each recode block contributes an additional 33 bp to the amplicon length. Expected amplicon sizes are: 67 bp, 90 bp, 113 bp, 136 bp, 159 bp, and 182 bp for memory oligonucleotides comprising 0, 1, 2, 3, 4, and 5 recode blocks, respectively.

[0075] FIG. 34 depicts Ct values comparing samples with beads versus their corresponding supernatant solutions, from beads that underwent the polymerization step with or without Taq polymerase.

[0076] FIG. 35 depicts molecules of copied ULI per square micron derived from qPCR results on beads, comparing beads which had received Taq polymerase in the polymerization step with beads which had not received Taq polymerase.

[0077] FIG. 36 depicts qPCR analysis (on beads) of the number of copies per pm2of the ULI-containing sequence originating from the B361 oligonucleotide which was ligated to the bead surface at different concentrations along with B359 and B360, with either 15 or 30 bridge amplification cycles of amplification afterward.

[0078] FIG. 37 depicts superimposed confocal laser microscopy images of sample E beads, USER treated and hybridized to two fluorescent probes B421 (corresponding to Cy5 channel shown in red) and SyslDHI (corresponding to Cy3 channel shown in green) showing clusters containing two different ULI sequences.

[0079] FIG. 38 depicts matched confocal laser microscope images of sample C beads, USER treated and hybridized to two fluorescent probes B421 and SyslDHI where the left image is in the Cy5 channel (corresponding to B421) and the right image is the Cy3 channel (corresponding to SyslDHI).

[0080] FIG. 39 depicts confocal images of bridge-amplified beads stained with Picogreen DNA intercalating dye. Left panel shows non-USER treated sample C beads with bright spots corresponding to clusters. Middle panel shows B357 control bead (non-ligated, no bridge amplification, no USER) at matching microscope settings. Right panel shows B357 control bead at higher image brightness level.

[0081] FIG. 40A-41B depict HPLC chromatograms and mass spectra corresponding to deaza-modified oligonucleotide B209 subjected to incubation in 12% (w / v) tris(pentafluorophenyl)borane in anhydrous acetonitrile at 55°C for 15 hours. FIG. 40A: the LC trace of the stressed oligonucleotide. FIG. 40B: theAttorney Docket No: 062954-509001 WO mass spectrum of the stressed oligonucleotide. FIG. 41A: the LC trace of the unstressed control oligonucleotide. FIG. 41B: the mass spectrum of the unstressed control oligonucleotide.

[0082] FIG. 42A-42D depict HPLC chromatograms and mass spectra corresponding to natural DNA oligonucleotide Sys2PR5 subjected to incubation in 12% (w / v) tris(pentafluorophenyl)borane in anhydrous acetonitrile at 55°C for 15 hours. FIG. 42A: the LC trace of the stressed oligonucleotide. FIG. 42B: the mass spectrum of the stressed oligonucleotide. FIG. 42C: lower left panel shows the LC trace of the unstressed control oligonucleotide. FIG. 42D: the mass spectrum of the unstressed control oligonucleotide.

[0083] FIG. 43 depicts HPLC chromatogram and mass spectra corresponding to deaza-modified oligonucleotide B208 subjected to incubation in trifluoroacetic acid at 55°C for 22 hours. Left panel shows the LC trace with labeled peaks A, B, and C, and right panels show the mass spectra corresponding to peaks A, B, and C.

[0084] FIG. 44 depicts HPLC chromatogram and mass spectra corresponding to deaza-modified oligonucleotide B208 subjected to incubation in trifluoroacetic acid at 55°C for 22 hours. Lower middle panel shows the LC trace with arrows connecting to mass spectra for corresponding peaks labeled D, E, F, G, and H.

[0085] FIG. 45 depicts HPLC chromatogram and mass spectrum corresponding to unstressed deaza- modified oligonucleotide B208 control. Upper panel shows the LC trace and lower panel shows the mass spectrum.

[0086] FIG. 46 shows results of a non-denaturing gel showing PCR amplicons from cognate ZIP code combinations with size increments of 10 bp, with predicted amplicon sizes listed below respective lanes demonstrating unique banding patterns for each combination

[0087] FIG. 47 depicts HPLC analysis of samples: Initial, PITC-reacted, and TFA-Cleaved.

[0088] FIG. 48A depictsHPEC analysis of uncleaved pep9-PTC control showing the peptide-PTC peak at 20.2 minutes. FIG. 48B depicts HPEC overlay comparing cleavage of pep9-PTC with BCF in water, BCF in methanol, and TFA. The cleaved product (ala-ATZ) appears at 24.5 minutes. The broad peak at approximately 26.5 minutes corresponds to BCF.

[0089] FIG. 49A-49D depict chemical structures of modified phenylisothiocyanate (PITC) reagents used for on-surface Edman degradation with oligonucleotide-connected peptides. FIG. 49A shows tetrazinefunctional PITC variant 1 (TzPITC VI). FIG. 49B shows tetrazine-functional PITC variant 2 (TzPITC V2). FIG. 49C shows tetrazine-functional PITC variant 4 (TzPITC V4). These three TzPITC variants enable oligonucleotide attachment via inverse electron demand Diels- Alder chemistry with transcyclooctene reagents. FIG. 49D shows azido-PITC, which enables oligonucleotide attachment via strain- promoted azide-alkyne cycloaddition with DBCO-functionalized oligonucleotides.Attorney Docket No: 062954-509001 WODETAILED DESCRIPTION

[0090] Described herein are methods and compositions that may be useful for determining identity and positional information of amino acids of macromolecules, such as proteins, polypeptides, peptides.

[0091] In some embodiments, methods are provided for determining identity and positional information of amino acid residues within peptides. Such methods may involve coupling peptides to solid supports, contacting the peptides with chemically-reactive conjugates that bind to terminal amino acid residues, and iteratively cleaving amino acids while transferring sequence information to nucleic acid constructs. The nucleic acid constructs may encode positional information (such as cycle number and spatial location) along with amino acid identity information. These encoded constructs may be amplified, assembled, and sequenced to reconstruct the amino acid sequence and abundance of the original peptides. Various implementations may utilize different strategies for information transfer, including immobilized location tags, solution-phase indexing, affinity-based separation, or combinations thereof.

[0092] In some embodiments, methods and compositions are provided for creating spatially organized molecular territories on surfaces. Such methods may involve establishing discrete zones, each containing a central analyte binding site surrounded by multiple copies of a unique molecular identifier. The molecular territories may be generated through controlled amplification processes that create sharp boundaries between adjacent zones. In some implementations, the spatial relationships between territories may be determined by detecting junctions between adjacent zones, encoding these junction relationships into nucleic acid constructs, and computationally reconstructing the spatial arrangement from sequencing data. The resulting systems may provide enhanced detection efficiency for low-abundance analytes and efficient spatial mapping with reduced sequencing requirements.

[0093] The protein sequencing methods and spatial organization technologies described herein may be implemented independently or in combination. When integrated, the spatial organization methods may provide location-specific molecular identifiers that enhance the protein sequencing workflows by creating dense local concentrations of identifying barcodes around each protein analyte. This integration may address fundamental limitations in single-molecule detection while maintaining the ability to track individual molecules throughout analytical processes.Introduction

[0094] The sequences and concentrations of cellular and secreted proteins are important indicators of cell health. Abnormal sequences or concentrations can indicate disease. However, current tools and technologies for characterizing proteomes are lacking in sensitivity, accuracy, affordability, and unbiased analysis. Early detection of unusual protein sequences or concentrations is essential for diagnosing andAttorney Docket No: 062954-509001 WO treating diseases like cancer. For these reasons, improved tools are needed to evaluate protein and peptide sequences and concentrations in biological samples.

[0095] Next-generation sequencing (NGS) of DNA and RNA has revolutionized diagnostics, clinical approaches, and research by allowing the analysis of billions of DNA sequences with high throughput and low cost. However, the ability to detect and quantify proteins and peptides has not kept pace, mainly because there is no equivalent to polymerase chain reaction (PCR) for amino acids. New tools for sensitive protein quantification and sequence analysis, similar to NGS, could advance the understanding of cellular processes and continue transforming research, diagnostics, clinical practices, and precision medicine.

[0096] Current advanced proteomics toolkits generally include: 1) Edman degradation followed by chromatography, 2) fragmentation followed by advanced separation and mass spectrometry (MS) techniques, and 3) protein recognition via affinity molecules. These methods provide useful information but do not achieve the scale, throughput, reproducibility, accessibility, or affordability needed for transformative applications in research, diagnostics, or therapeutics.

[0097] Peptide sequencing via Edman degradation was first proposed and automated by Pehr Edman in the 1950s. The process is similar to Sanger sequencing. In short, the N-terminal amino acid of a peptide undergoes stepwise degradation through chemical reactions and high-performance liquid chromatography (HPLC) to reveal the peptide sequence. First, the N-terminal amino acid reacts with phenyl isothiocyanate (PITC) under basic conditions to form a phenylthiocarbamoyl (PTC) derivative. The PTC-modified amino group is then treated with acid (usually anhydrous trifluoroacetic acid, TFA) to yield an ATZ-modified amino acid, which is separated from the peptide, exposing the next N-terminus. The ATZ-amino acid is then converted to a PTH-amino acid and analyzed through chromatography. This process is repeated to determine the entire peptide sequence. While effective, this method requires large protein samples and lacks the throughput and cost-effectiveness for large-scale discovery.

[0098] More recent developments include multiplexed methods and devices for Edman degradation-based peptide sequencing of small protein quantities, such as described in Chhabra, U.S. Patent No. 7,611,834 B2. However, these methods are still unsuitable for highly parallelized, high-throughput proteomics.

[0099] In the past 20 years, peptide analysis through fragmentation and mass spectrometry (LC / MS) has been increasingly used to quantify protein abundance and determine sequences. In certain applications, recognition-based proteomics has also been employed, where affinity molecules (e.g., antibodies, aptamers, or modified proteins) recognize the analyte’s tertiary structure. These are often linked to molecular beacons that fluoresce or indicate binding, such as in ELISA assays. However, like other approaches, fragmentation and recognition-based methods do not offer the throughput and efficiency needed for large-scale discovery.

[0100] Thus, the current landscape for proteomic analysis includes the following general approaches: 1) Edman degradation followed by chromatography; 2) fragmentation followed by advanced separation andAttorney Docket No: 062954-509001 WO mass spectroscopy techniques; and 3) recognition of proteins via affinity molecules. While these (and other) approaches can provide useful information for researchers, they do not provide such information at the scale, throughput, or cost needed to unlock transformative applications in research, diagnostics, or therapeutics. Some more particular challenges associated with current approaches (e.g., Edman’s, LC / MS, and affinity approaches) include:(a) Protein folding is dynamic, and proteins can lose their characteristic shape. When they do, recognition-based methods become inaccurate. This can happen in the case of labile proteins, or uncontrolled sample treatment prior to analysis.(b) Recognition-based methods do not inform as to whether the protein sequence is a catalytically- ineffective variant, as often becomes the case in cancer biology.(c) Biomarkers of interest are likely to be present at fM or lower concentrations, beneath the detection limit of most available tools used to quantify the absolute abundance of multiple proteins.(d) The universe of protein molecules is extensive. It is much more complex than the RNA transcriptome, due to additional diversity introduced by post-translational modifications (PTMs).(e) Proteins within a cell dynamically change (in expression level and modification state) in response to the environment, physiological state, and disease state. Thus, proteins contain a vast amount of relevant information that is largely unexplored.

[0101] Generating an effective collection of affinity agents having low cross-reactivity to off-target macromolecules can be time-consuming.(a) Multiplexing the readout of a collection of affinity agents having minimizing cross-reactivity between the affinity agents and off-target macromolecules is challenging.(b) Existing methods and the automation around current approaches is slow, expensive, and for the case of Edman’s methods, have a limited throughput of only a few peptides per day.

[0102] LC / MS suffers from drawbacks including: high instrument cost, requirement for a sophisticated user, poor quantification ability, limited dynamic range, and limited repeatability / reproducibility. Since proteins ionize at different levels of efficiencies, absolute quantitation and even relative quantitation between samples is challenging.

[0103] LC / MS analyzes the more abundant species, so there is a need to employ complex upfront sample preparation, e.g., nanoparticle corona, making characterization of low abundance proteins challenging.(a) LC / MS sample throughput is typically limited to a few thousand peptides per run.

[0104] More recent attempts to develop single-molecule methodologies suffer from poor discrimination of n-terminal or c-terminal AA and are confounded by “neighbor effects”. Still other single molecule methodologies under development may require costly instrumentation because amplification of the analyte is not possible and small numbers of photons, electrons may impact detection elements.Attorney Docket No: 062954-509001 WO

[0105] The present disclosure addresses the above challenges as well as other needs by providing methods, systems, and compositions for analyzing polymeric macromolecules via recoding of their sequences into DNA polymers for subsequent DNA sequencing and analysis. The numerous applications of the present disclosure include peptide sequence and quantification determination in synthetically-derived and biologically-derived samples that include a plurality of protein complex, protein, and / or polypeptide components.

[0106] The present disclosure provides methods for analyzing polymeric macromolecules, such as peptides, polypeptides, and proteins. Accordingly, aspects of the present disclosure relate to the field of proteomics.Example Illustrations

[0107] FIG. 1 illustrates a simplified diagram of an exemplary workflow for analyzing polymeric macromolecules according to embodiments of the present disclosure. Any aspect or combination of aspects from FIG. 1 may be included in a method herein.

[0108] FIG. 2 schematically illustrates primary steps of Process 100 and / or Process 200, according to embodiments of the present disclosure. Any aspect or combination of aspects from FIG. 2 may be included in a method herein.

[0109] In a first stage, operations (a)-(e) in FIG. 2, location tags are immobilized to the same solid support as the macromolecular analyte. At operation (a) a macromolecular analyte is attached to a solid support. At operation (b) a chemically-reactive conjugate (CRC) having a location tag is contacted with the N- terminus of an immobilized macromolecular analyte. At operation (c) under basic conditions the conjugate reacts with the N-terminal amino acid to form a phenylthiocarbamoyl- amino acid (PTC) conjugate. A stringent wash removes non-coupled conjugate, and then, at operation (d), activation of an orthogonal chemistry is initiated to to tether the conjugate to the support, in proximity to the anchor point of the associated analyte. For example, by changing redox conditions to induce di-thiol formation, or adding Cu2+, stabilizer and redox components to induce a Click reaction, PTC-thiol conjugates or PTC-alkyne conjugates may be immobilized to a thiol-functionalized or an azide-functionalized solid support, respectively. Following immobilization of conjugates to the solid support, a conjugate-reactive scavenger may be added to cap the reactivity of any bound conjugate that was not washed away in the previous step(s), to render it inactive for future n-terminal amino acid reaction. At operation (e), peptide bond cleavage targeting the N- terminal amino acid of the peptide is induced. In examples employing Edman’s degradation chemistry, this is facilitated by a change in pH from basic to acidic conditions. Operations (a)-(e) may be repeated provide additional location tags on a solid support.Attorney Docket No: 062954-509001 WO

[0110] In a second stage, operations (f)-(k) in FIG. 2, recode units are released from the solid support of the macromolecular analyte. At operation (f) a chemically-reactive conjugate (CRC) having a cycle tag is contacted with the N-terminus of an immobilized macromolecular analyte. At operation (g) under basic conditions the conjugate reacts with the N-terminal amino acid to form a coupled phenylthiocarbamoylamino acid (PTC) conjugate. A stringent wash removes non-coupled conjugate, and then, at operation (h) the transfer complex between spatially localized cycle tag and location tags is formed. At operation (i) polymerase joins the information of the location tag to the coupled cycle tag conjugate. A more detailed illustration of operation (i) is provided in FIG. 3. At operation (j) the conjugate is liberated via cleavage of the terminal peptide bond under acidic conditions. A first iteration through operations (f)-(k) (e.g., first cycle) provides information related to a monomer of the immobilized polymeric analyte. A second cycle thereof provides information related to the next monomer of the immobilized polymeric analyte, and so on. Iterating through steps (f)-(k) for n cycles creates a pool of spatially localized conjugates each holding cycle and location information. With appropriate spacing between anchor points of immobilized macromolecular analytes, and associatd placements of location tags, conjugates associated with a single analyte are distinguishable from other analytes.

[0111] In a third stage, operations (l)-(n) in FIG. 2, information is assembly into amplicons for NGS readout. At operation (1) the PTH-AA moiety of the solution-phase conjugate provides a target for amino acid identity determination. An affinity molecule having a recode tag contacts the conjugate holding the PTH-amino acid, cycle information and location information, and a hydridiztion complex is formed. At operation (m) information of the amino acid identity is transferred to the conjugate holding cycle and location information to create a recode block conjugate. A more detailed illustration of operation (1) is provided in FIG. 4 Prior to operation (n), the immobilizing moiety of the CRC may be used to purify the recode block conjugates. At operation (n) recode blocks are assembled. More detailed illustrations of assembly of recode blocks is provided in FIG. 5 and FIG. 6.

[0112] FIG. 3 depicts an operation of the iterative process whereby transfer of information occurs from the CRC having a location tag to the CRC having a cycle tag. Extension of the 3’ end of a location tag may be prevented by incorporation of blocking agents, such as a dideoxy NT introduced during location tag nucleic acid synthesis. The cycle tag nucleic acids of the CRC may be designed to include universal hyb sequences that promote DNA duplex formation with nearly location tags designed to provide a complementary sequence. Polymerase extension of the 3’ end of the cycle tag provdes an example of joining cycle and location information. Temperature, oganic solvents, and or acidic conditions of the peptide cleavage reaction promote dehybridization and release of the recode unit conjugates.

[0113] FIG. 4 depicts an operation whereby recode blocks are built. In this operation, amino acid information is associated with cycle and location information. Briefly, a plurality of binding agents thatAttorney Docket No: 062954-509001 WO recognize the various immobilized conjugate-AA complexes are introduced, and bind to their cognate targets. The binding agents are engineered to preferentially recognize specific conjugates based on differences in the cognate amino acid of the immobilized conjugate. Those agents that possess both the cognate AA will thereby direct transfer, for example by ligation or extension, of AA information. Repeated binding, washing, and ligation allows multiple attempts to find cognate partners and transfer information to each immobilized conjugate-AA-cycle tag complex to build a recode block. Accordingly, multiple successive binding cycles can be used to drive the yield of information transfer from the binding agents to the immobilized conjugates to an arbitrarily high completion level. Exemplary elements of the completed recode block include: 1) a universal assembly sequence helpful to assembly recode blocks into amplicons in subsequenct steps of the Process 200; 2) cycle information; 3) location information; 4) amino acid identity information; 5) a universal assembly sequence helpful to assembly recode blocks into amplicons in subsequent steps of the Process 200.

[0114] FIG. 5 illustrates the assembly of recode blocks into a memory oligo (e.g. joining recode blocks into a long amplicon for efficient NGS analysis). A memory oligo is capable of being amplified, then analyzed using DNA sequencing methods to determine a sequence and / or abundance of the immobilized protein analyte(s). Briefly, at operation n.2, recode blocks are optionally separated from binding agents, CRC components, and incomplete extension or information transfer errors. One optional method apparent from the illustration of operation (n.2) is capture of the recode block oligo by a magnetic streptavidin coated bead. Following the clean-up (purification, concentration, buffer exchange or other manipulation) of operation (n.3) assembly of operation (n.4) is accomplished. More detailed illustrations of assembly processed of operation (n.4) provided in FIG. 6.

[0115] FIG. 6 schematically illustrates the assembly of a memory oligo at operation (n.4) of the recoding Process 200, according to embodiments of the present disclosure. As shown, the overlapping and complementary universal assembly sequences facilitate assembly. Several molecular biology methods and enzymes may be useful during assembly. In certain embodiments, a recode block comprises a sequence that facilitates assembly of a memory oligo, and / or that facilitates target enrichment, target depletion, and / or sequencing sample preparation (e.g. NGS sample preparation), such as a CRISPR PAM or spacer sequence.

[0116] More specifically, FIG. 6. illustrates the utilization of universal sequences to facilitate linking of recode blocks during memory oligo assembly without regard to any specific order, according to embodiments of the present disclosure. As described above, recode blocks may be linked in any sequential order to create a memory oligo. This is due to cycle, location, and amino acid information being immediately adjacent in an assembled recode blocks, regardless of whether the recode blocks are in sequential or non-sequential order within a memory oligo.Attorney Docket No: 062954-509001 WO

[0117] To facilitate the assembly of memory oligos without regard to any specific order, universal assembly sequences may be utilized during the recoding process. Such universal sequences may be attached to the 5’ and / or 3’ ends of cycle tags, location tags, and / or recode tags. Attaching complementary universal sequences to two or more cycle tags and / or recode tags facilitates the random linking (e.g., ligation) of resulting recode blocks during memory oligo assembly, without regard to sequential order, and a correct macromolecule analyte sequence may be assigned during post-sequencing analysis.

[0118] There are three primary operations required for efficient stepwise random assembly of memory oligos: Initiation, Extension, and Termination. Initiation: Initiation may be accomplished by providing a forward primer complementary to the recode block. Extension: U-AS is a unifying assembly sequence for memory oligo assembly. Various molecular biology approached to chew back and provide a single-stranded oligonucleotide that facilitates subsequent extension reactions are available. These include RNAseH, and various combinations of endonuclease and polymerase. One or more nicking endonuclease sites may be designed into the U-AS to provide a single-stranded oligonucleotide that facilitates subsequent extension reactions. During the initiation and extension steps, U-AS’ from the recode block interacts with U-AS of the primer. In some embodiments events are as follows: 1) hybridization of U-AS to U-AS’; 2) extension of U-AS by a polymerase; 3) RNAse-H action on the newly formed double-stranded RNA-DNA duplex to expose a new U-AS 3’ end; and 4) interaction of a U-AS’ sequence of a next recode block with the newly formed single-stranded U-AS’. Termination: Nucleic acid constructs used for assembly of recode blocks may be designed so as to terminate assembly, or limit assembly to a defined number, or a defined range of assembled recode blocks. Termination may facilitate subsequent PCR amplification of memory oligos by providing reverse primer sequence, UMI, PIN, sample index, or other information.

[0119] FIG. 7 provides an exemplary 2-dimensional perspective of a solid support on which a plurality of immobilized macromolecules (filled circles) and a plurality of immobilized recode unit complexes having location tags (open circles) reside, according to embodiments of the present disclosure. The overlapping areas present the opportunity for interaction between cycle tags and location nucleic acids during the operations of transferring information between cycle tags and location tags. In some locations the location tag completely overlaps with the location where one immobilized protein resides. In some locations the location tag partially overlaps with the location where one immobilized protein resides. In some locations the location tag either completely or partially overlaps with the location where more than one immobilized protein resides. In some locations more than one location tag completely overlaps with the location where 1 and only 1 immobilized protein resides. In some locations more than one location tag partially overlaps with the location where 1 and only 1 immobilized protein resides. In some locations more than one location tag either partially or completely overlaps with the location or one or more immobilized protein resides. In all cases, it may be possible to assign relative location information to macromolecule via in situ analysis.Attorney Docket No: 062954-509001 WO

[0120] FIG. 8 schematically illustrates primary steps of Process 300, according to embodiments of the present disclosure.

[0121] In a first stage of Process 400, operations (a)-(e) as described in FIG. 2, whereby location tags are immobilized to the same solid support as the macromolecular analyte are completed.

[0122] In a second stage, operations (f)-(k) as described in FIG. 2, whereby recode units are released from the solid support of the macromolecular analyte, are completed. Of particular usefulness for Process 300 is the introduction of modified nucleotide triphosphates having modifications, such as a biotin group, which are useful in downstream clean-up operations. Accordingly, a streptavin surface may be used to remove excess assembly sequences and / or excess assembly splints used to bring nucleic acids into proximity for ligation, or chemical assembly. Assembly of non-polymerizable long concatenated strucutres may provide a readable input to nanopore analysis.

[0123] FIG. 9 schematically illustrates primary steps of Process 400, according to embodiments of the present disclosure.

[0124] In a first stage of Process 400, operations (a)-(e) as described in FIG. 2, whereby location tags are immobilized to the same solid support as the macromolecular analyte are completed.

[0125] In a second stage, operations (f)-(j) as described in FIG. 2, whereby recode units are released from the solid support of the macromolecular analyte, are completed.

[0126] In solution separation, purification, and indexing operations are depicted in steps (k) through (s). These hold the advantage of solution-phase kinetics and efficiencies for joining operations. Joining recode units provide long constructs that are advantageous to improve NGS DNA sequencing efficiency.

[0127] At operation (k) though (1) serial application of binding elements, or groups of binding elements, allow for creation of sub-pools of recode unit complexes. Indexes may be efficiently joined to recode unit nucleic acids using appropriate methods and molecular biology tools to capture amino acid identity information. Methods to efficiently join recode unit nucleic acids for efficient indexing with cycle nucleic acid are described herein. Cycle information is applied only once per joined group of recode units, thereby providing cycle information to the entire group, and minimizing the number of nucloetides that must be sequenced in step (t) of Process 400.METHODS FOR DETERMINING PROTEIN SEQUENCE AND ABUNDANCE

[0128] Disclosed herein, in some embodiments, are methods for determining protein sequence or abundance. Some embodiments include determining amino acid identity and position (e.g. determining identity and positional information of an amino acid residue of a peptide). Some such methods are included in FIG. 1 or FIG. 2. A method may determine protein sequence. A method may determine protein abundance. A method may determine protein sequence and abundance. A method may includeAttorney Docket No: 062954-509001 WO cleavage of peptide bonds, generation or use of a recode tag, generation or use of a recode unit, generation or use of a recode block, transfer of information, joining, generation or use of a memory oligonucleotide, or a combination thereof. A method may include surface preparation. A method may include creation of a local ULI zone. A method may include ZIP code incorporation or use. A method may include spatial mapping. A method may include alternative ZIP code architecture. A method may include generation or use of a bridge oligonucleotide. A method may include junction detection. A method may include generation or purification of a junction-spanning amplicon. A method may includeintegration of protein sequencing with a ULI neighborhood, solution-phase or surface-based inegration, spatial biology integration, multiplexed sample analysis, use of multiple attachment points or in situ digestion, or a combination thereof. A method may be useful for diagnostic purposes. A method may be used with a microfluidic platform. A method may be used with an alternative molecular architecture. A method may be used with an automated or high-throughput process. A method may be made compatible with a process, or may be standardized. A method may make use of a chemically reactive conjugate (CRC) such as a trifunctional CRC, a bifuncitonal CRC, or both. A method may make use of a kit herein.

[0129] Some embodiments include providing a peptide. A peptide may be coupled to a solid support. Coupling to a solid suppot may be such that a N-terminal amino acid residue is not directly coupled to the solid support. Coupling to a solid suppot may be such that a N-terminal amino acid residue is exposed to reaction conditions. A peptide such as a peptide coupled to a solid support may be local to a nucleic acid barcode. A nucleic acid barcode may include a sequence individual to a peptide. A nucleic acid barcode may include deaza nucleotides. A nucleic acid barcode may be useful for identifying a peptide individually. Some embodiments include providing the peptide coupled to a solid support such that a N- terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions, and the peptide is local to a nucleic acid barcode for identifying the peptide individually.

[0130] Some embodiments include providing a chemically-reactive conjugate (CRC). A chemicallyreactive conjugate may include a cycle tag. A cycle tag may include a cycle nucleic acid. A cycle nucleic acid may include deaza nucleotides. A cycle nucleic acid may be associated with a cycle number A cycle nucleic acid may include a nucleic acid sequence associated with a cycle number. A chemically-reactive conjugate may include a location tag. A location tag may include a location nucleic acid. A location nucleic acid may include deaza nucleotides. A location nucleic acid may be associated with the peptide or a location of the peptide. A chemically-reactive conjugate may include a reactive moiety. A reactive moiety may bind an amino acid residue such as an N-terminal amino acid residue of a peptide. Some embodiments include providing a first CRC and a second CRC. Some embodiments include providing aAttorney Docket No: 062954-509001 WO first CRC and subsequent CRCs (e.g. for each subsequent amino acid of a peptide). A first CRC may include a location tag, while a second CRC or subsequent CRCs may include a cycle tag. A location tag may be used for identifying a protein, while a cycle tag may be usef for identifying an amino acid position within a peptide. Some embodiments include providing a chemically-reactive conjugate, the chemicallyreactive conjugate comprising: a cycle tag comprising a cycle nucleic acid associated with a cycle number, and a reactive moiety for binding the N-terminal amino acid residue of the peptide. Some embodiments include providing a first CRC and a second CRC. Some embodiments include providing a first CRC and subsequent CRCs (e.g. for each subsequent amino acid of a peptide). A first CRC may include a location tag, while a second CRC or subsequent CRCs may include a cycle tag. A location tag may be used for identifying a protein, while a cycle tag may be usef for identifying an amino acid position within a peptide.

[0131] Some embodiments include contacting a peptide with a chemically-reactive conjugate. Some embodiments include coupling a chemically-reactive conjugate to an amino acid. An amino acid may be an N-terminal amino acid of a peptide. Some embodiments include forming a conjugate complex. A conjugate complex may include an amino acid and a chemically-reactive conjugate, such as a CRC that includes a location tag. Some embodiments include contacting the peptide with the chemically-reactive conjugate, thereby coupling the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex. Some embodiments include immobilizing a conjugate complex, for example to a solid support. Immobilization may be via the immobilizing moiety. Some embodiments include contacting the second N-terminal amino acid residue of the peptide with the second chemically- reactive conjugate, thereby coupling the second chemically-reactive conjugate to the second N-terminal amino acid of the peptide to form a coupled-conjugate-complex.

[0132] Some embodiments include transferring information. Information of a nucleic acid barcode may be transferred. Information of a cycle tag may be transferred. Information may be transferred to a cycle tag. Information may be transferred to a nucleic acid barcode. A cycle tag may be of a chemically- reactive conjugate. Some embodiments include transferring information of the nucleic acid barcode to the cycle tag of the chemically-reactive conjugate.

[0133] Some embodiments include forming a transfer complex. A transfer complex may include a location complex (e.g. an immobilized location complex). A transfer complex may include a coupled- conjugate-complex. A transfer complex may include an immobilized location complex and a coupled- conjugate-complex. Some embodiments include bringing a location nucleic acid into proximity with a cycle nucleic acid. Some embodiments include forming a transfer complex comprising the immobilized location complex and the coupled-conjugate-complex, thereby bringing the location nucleic acid into proximity with the cycle nucleic acid.Attorney Docket No: 062954-509001 WO

[0134] Some embodiments include joining a location nucleic acid to a cycle nucleic acid. Some embodiments include joining a reverse complement of a location nucleic acid a cycle nucleic acid. Some embodiments include forming a recode unit. Some embodiments include joining information of the location nucleic acid and the cycle nucleic acid. A recode unit may be formed within a transfer complex. A recode unit may include a location nucleic acid. A recode unit may include a reverse complement of a location nucleic acid. A recode unit may include location tag sequence information, or may include location nucleic acid sequence information. A recode unit may include a cycle nucleic acid. A recode unit may include a reverse complement of a cycle nucleic acid. A recode unit may include cycle tag sequence information, or may include cycle nucleic acid sequence information. Some embodiments include within the transfer complex, joining the location nucleic acid or a reverse complement thereof to the cycle nucleic acid to form a recode unit, or otherwise joining information of the location nucleic acid and the cycle nucleic acid, thereby creating an immobilized recode unit comprising location and cycle information of the location nucleic acid and cycle nucleic acid.

[0135] Some embodiments include providing an amino acid complex. The amino acid complex may be in solution. Some embodiments include cleaving an amino acid residue. Some embodiments include cleaving an N-terminal amino acid residue. Cleaving may separate an amino acid residue from a peptide. Cleaving may separate a group of amino acids residues from a peptide. Cleaving may separate an N- terminal amino acid residue from a peptide. Cleaving may separate a group of N-terminal amino acid residues from a peptide. Such cleaving may provide an amino acid complex in solution. An amino acid complex may include a cleaved or separated N-terminal amino acid residue. An amino acid complex may include an amino acid coupled with a nucleic acid barcode. An amino acid complex may include an amino acid coupled with a cycle tag. An amino acid complex may include a nucleic acid barcode coupled with a cycle tag. An amino acid complex may include an amino acid coupled with a nucleic acid barcode and a cycle tag. Some embodiments include cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby providing an amino acid complex in solution, the amino acid complex comprising the cleaved and separated N-terminal amino acid residue coupled with the nucleic acid barcode and the cycle tag. Some embodiments include providing a location complex. Some embodiments include providing an immobilized location complex. A location complex (e.g. an immobilized location complex) can include a cleaved and separated N-terminal amino acid residue and a location nucleic acid. Some embodiments include cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby exposing a next amino acid residue as a second N-terminal amino acid residue on the cleaved peptide and providing an immobilized location complex, the immobilized location complex comprising the cleaved and separated N-terminal amino acid residue and the location nucleic acid. A second amino acid or a second N-terminal amino acid does not necessarily need to be an amino acid fromAttorney Docket No: 062954-509001 WO position 2 of a peptide. A second amino acid or a second N-terminal amino acid may merely be a subsequent amino acid or a subsequent N-terminal amino acid. While referring to a cleaved or liberated N-terminal amino acid, the amino acid can be disconnected from a peptide and may merely be referred to as “N-terminal” for sake of referring to an amino acid that was previously N-terminal.

[0136] Some embodiments include cleaving and thereby separating a second N-terminal amino acid residue from a peptide. Cleaving and separating an N-terminal amino acid such as a second N-terminal amino acid may expose another next amino acid residue as a third N-terminal amino acid residue on a cleaved peptide. Cleaving and separating a N-terminal amino acid such as a second N-terminal amino acid may liberate a recode unit complex from a solid support. Some embodiments include exposing another next amino acid residue as a third N-terminal amino. Some embodiments include liberating a recode unit complex from the solid support. A recode unit complex (e.g. liberated recode unit complex) may include a cleaved and separated N-terminal amino acid residue (e.g. a second N-terminal amino acid residue). A recode unit complex (e.g. liberated recode unit complex) may include a recode unit. A recode unit complex (e.g. liberated recode unit complex) may include a cleaved and separated N-terminal amino acid residue (e.g. a second N-terminal amino acid residue) and a recode unit. A liberated recode unit complex may include a liberated recode unit that is not immobilized, or is no longer immmobilized. Some embodiments include cleaving and thereby separating the second N-terminal amino acid residue from the peptide, thereby exposing another next amino acid residue as a third N-terminal amino acid residue on the cleaved peptide and liberating a recode unit complex from the solid support, the liberated recode unit complex comprising the cleaved and separated second N-terminal amino acid residue and the recode unit, wherein the recode unit is no longer immobilized.

[0137] Cleaving and separating an N-terminal amino acid may be performed in the presence of a Lewis acid. An example of a Lewis acid may include tris(pentafluorophenyl)borane (BCF). In some embodiments, the cleaving and separating an N-terminal amino acid is not performed in a presence of trifluoroacetic acid (TFA). Tags such as a cycle tag or location tag may include deazanucleotides.

[0138] Some embodiments include contacting an amino acid complex with a binding agent. A binding agent may include a binding moiety. A binding moiety may preferentially bind to an amino acid complex. A binding agent may include a recode tag. A recode tag may include a recode nucleic acid. A re nucleic acid may include deaza nucleotides. A recode nucleic acid may include a nucleic acid sequence, which may correspond with a binding agent. The binding agent may include: a binding moiety for preferentially binding to the amino acid complex, and a recode tag comprising a recode nucleic acid corresponding with the binding agent. Some embodiments inculde forming an affinity complex. An affinity complex may include an amino acid complex and a binding agent. Contacting the amino acid complex with a binding agent may form an affinity complex, the affinity complex comprising the amino acid complex and theAttorney Docket No: 062954-509001 WO binding agent. Some embodiments include contacting the amino acid complex with a binding agent, the binding agent comprising: a binding moiety for preferentially binding to the amino acid complex, and a recode tag comprising a recode nucleic acid corresponding with the binding agent, thereby forming an affinity complex, the affinity complex comprising the amino acid complex and the binding agent.

[0139] Some embodiments include contacting a recode unit complex with a binding agent. Such a binding agent may include a binding moiety for preferentially binding to an amino acid residue (e.g. a second N-terminal amino acid residue, or a subsequent N-terminal amino acid residue). The amino acid residue may be part of a recode unit complex. A binding agent may include a recode tag, for example comprising a recode nucleic acid corresponding with the binding agent. Som embodiments include forming an affinity complex. An affinity complex may include a recode unit complex. An affinity complex may include a binding agent. An affinity complex may include a recode unit complex and a binding agent. Forming an affinity complex may bring a recode nucleic acid into proximity with a recode unit. Some embodiments include bringing a recode nucleic acid into proximity with a recode unit. Some embodiments include contacting the recode unit complex with a binding agent, the binding agent comprising: a binding moiety for preferentially binding to the second N-terminal amino acid residue of the recode unit complex, and a recode tag comprising a recode nucleic acid corresponding with the binding agent, thereby forming an affinity complex, the affinity complex comprising the recode unit complex and the binding agent, and thereby bringing the recode nucleic acid into proximity with the recode unit.

[0140] Some embodiments include combining sequence information to generate a recode block. The sequence information may be of an afinity complex. Some embodiments include transferring information of the recode nucleic acid to the recode unit to generate a recode block. Some embodiments include obtaining sequence information of the recode block. Some embodiments include combining sequence information of a location nucleic acid and cycle nucleic acid. Some embodiments include combining sequence information of the nucleic acid barcode and cycle nucleic acid. Some embodiments include combining sequence information of the nucleic acid barcode and recode nucleic acid to generate a recode block. Some embodiments include combining sequence information of the cycle nucleic acid, and recode nucleic acid to generate a recode block. Some embodiments include combining sequence information of the nucleic acid barcode, cycle nucleic acid, and recode nucleic acid of the affinity complex to generate a recode block. Some embodiments include obtaining sequence information of the recode block. Some embodiments include based on the obtained sequence information, determining identity and positional information of an amino acid residue (e.g. a second amino acid residue or subsequent amino acid residues) of a peptide.Attorney Docket No: 062954-509001 WO

[0141] Some embodiments herein relate to a method for determining amino acid identity and positional information, including: providing a peptide coupled to a solid support such that a N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions, and the peptide is local to a nucleic acid barcode for identifying the peptide individually; providing a chemically-reactive conjugate, the chemically-reactive conjugate comprising: a cycle tag comprising a cycle nucleic acid associated with a cycle number, and a reactive moiety for binding the N-terminal amino acid residue of the peptide; contacting the peptide with the chemically-reactive conjugate, thereby coupling the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex; transferring information of the nucleic acid barcode to the cycle tag of the chemically-reactive conjugate; cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby providing an amino acid complex in solution, the amino acid complex comprising the cleaved and separated N-terminal amino acid residue coupled with the nucleic acid barcode and the cycle tag; contacting the amino acid complex with a binding agent, the binding agent comprising: a binding moiety for preferentially binding to the amino acid complex, and a recode tag comprising a recode nucleic acid corresponding with the binding agent, thereby forming an affinity complex, the affinity complex comprising the amino acid complex and the binding agent; combining sequence information of the nucleic acid barcode, cycle nucleic acid, and recode nucleic acid of the affinity complex to generate a recode block; and obtaining sequence information of the recode block.

[0142] Some embodiments include a method, comprising: providing a peptide, wherein the peptide is coupled to a solid support such that a N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions; providing a chemically-reactive conjugate, the chemically-reactive conjugate comprising: a location nucleic acid, a reactive moiety for binding the N- terminal amino acid residue of the peptide, and an immobilizing moiety for immobilization to the solid support; contacting the peptide with the chemically-reactive conjugate, thereby coupling the chemically- reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex; immobilizing the conjugate complex to the solid support via the immobilizing moiety; cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby exposing a next amino acid residue as a second N-terminal amino acid residue on the cleaved peptide and providing an immobilized location complex, the immobilized location complex comprising the cleaved and separated N-terminal amino acid residue and the location nucleic acid.

[0143] Some embodiments include a method, comprising: providing a chemically-reactive conjugate, the chemically-reactive conjugate comprising: a cycle tag comprising a cycle nucleic acid associated with a cycle number, and a reactive moiety for binding and cleaving an N-terminal amino acid residue of a peptide, and optionally a moiety for joining to an optional solid support; contacting the N-terminal aminoAttorney Docket No: 062954-509001 WO acid residue of the peptide with the chemically-reactive conjugate, thereby coupling the chemicallyreactive conjugate to the N-terminal amino acid of the peptide to form a coupled-conjugate-complex; forming a transfer complex comprising a location nucleic acid and the coupled-conjugate-complex, thereby bringing the location nucleic acid into proximity with the cycle nucleic acid; within the transfer complex, joining the location nucleic acid or a reverse complement thereof to the cycle nucleic acid to form a recode unit, or otherwise joining information of the location nucleic acid and the cycle nucleic acid, thereby creating an immobilized recode unit comprising location and cycle information of the location nucleic acid and cycle nucleic acid. Some embodiments further include cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby exposing another next amino acid residue as a subsequent N-terminal amino acid residue on the cleaved peptide and liberating a recode unit complex from the solid support, the liberated recode unit complex comprising the cleaved and separated N-terminal amino acid residue and the recode unit, wherein the recode unit is no longer immobilized. Some embodiments further include contacting the recode unit complex with a binding agent, the binding agent comprising: a binding moiety for preferentially binding to the amino acid residue of the recode unit complex, and a recode tag comprising a recode nucleic acid corresponding with the binding agent, thereby forming an affinity complex, the affinity complex comprising the recode unit complex and the binding agent, and thereby bringing the recode nucleic acid into proximity with the recode unit. Some embodiments further include transferring information of the recode nucleic acid to the recode unit to generate a recode block. Some embodiments further include obtaining sequence information of the recode block. Some embodiments further include based on the obtained sequence information, determining identity and positional information of the amino acid residue of the peptide.

[0144] Some embodiments include a method, comprising: contacting a recode unit complex with a binding agent, the recode unit complex comprising a recode unit and an amino acid, the recode unit comprising location and cycle information of a location nucleic acid and cycle nucleic acid, the binding agent comprising: a binding moiety for preferentially binding to the amino acid residue of the recode unit complex, and a recode tag comprising a recode nucleic acid corresponding with the binding agent, thereby forming an affinity complex, the affinity complex comprising the recode unit complex and the binding agent, and thereby bringing the recode nucleic acid into proximity with the recode unit. Some embodiments further include: transferring information of the recode nucleic acid to the recode unit to generate a recode block. Some embodiments further include: obtaining sequence information of the recode block. Some embodiments further include: based on the obtained sequence information, determining identity and positional information of the amino acid residue within a peptide. The location information may be used to identify the amino acid as being from the peptide,Attorney Docket No: 062954-509001 WO

[0145] Some embodiments include a method for determining amino acid identity and positional information, comprising: providing a peptide, wherein the peptide is coupled to a solid support such that a N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions; providing a chemically -reactive conjugate, the chemically-reactive conjugate comprising: a location nucleic acid, a reactive moiety for binding the N-terminal amino acid residue of the peptide, and an immobilizing moiety for immobilization to the solid support; contacting the peptide with the chemically-reactive conjugate, thereby coupling the chemically-reactive conjugate to the N- terminal amino acid of the peptide to form a conjugate complex; immobilizing the conjugate complex to the solid support via the immobilizing moiety; cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby exposing a next amino acid residue as a second N-terminal amino acid residue on the cleaved peptide and providing an immobilized location complex, the immobilized location complex comprising the cleaved and separated N-terminal amino acid residue and the location nucleic acid; providing a second chemically-reactive conjugate, the chemically-reactive conjugate comprising: a cycle tag comprising a cycle nucleic acid associated with a cycle number, and a reactive moiety for binding and cleaving the second N-terminal amino acid residue of the peptide; contacting the second N- terminal amino acid residue of the peptide with the second chemically-reactive conjugate, thereby coupling the second chemically-reactive conjugate to the second N-terminal amino acid of the peptide to form a coupled-conjugate-complex; forming a transfer complex comprising the immobilized location complex and the coupled-conjugate-complex, thereby bringing the location nucleic acid into proximity with the cycle nucleic acid; within the transfer complex, joining the location nucleic acid or a reverse complement thereof to the cycle nucleic acid to form a recode unit, or otherwise joining information of the location nucleic acid and the cycle nucleic acid, thereby creating an immobilized recode unit comprising location and cycle information of the location nucleic acid and cycle nucleic acid; cleaving and thereby separating the second N-terminal amino acid residue from the peptide, thereby exposing another next amino acid residue as a third N-terminal amino acid residue on the cleaved peptide and liberating a recode unit complex from the solid support, the liberated recode unit complex comprising the cleaved and separated second N-terminal amino acid residue and the recode unit, wherein the recode unit is no longer immobilized; contacting the recode unit complex with a binding agent, the binding agent comprising: a binding moiety for preferentially binding to the second N-terminal amino acid residue of the recode unit complex, and a recode tag comprising a recode nucleic acid corresponding with the binding agent, thereby forming an affinity complex, the affinity complex comprising the recode unit complex and the binding agent, and thereby bringing the recode nucleic acid into proximity with the recode unit; transferring information of the recode nucleic acid to the recode unit to generate a recode block; obtaining sequence information of the recode block; and based on the obtained sequenceAttorney Docket No: 062954-509001 WO information, determining identity and positional information of the second amino acid residue of the peptide. In some embodiments, a second chemically reactive conjugate includes a moiety for joining to a second solid support.Cleavage of peptide bonds

[0146] Some embodiments include cleaving a peptide bond. Disclosed herein, are methods for cleavage of peptide bonds with a methodology for determining the positional and identity information of amino acids within peptides.

[0147] The classic Edman procedure consists of three steps: 1) a coupling reaction of N-terminal amino acids with an isothiocyanate; 2) the cyclization / cleavage reaction, the formation of anilinothiazolinone (ATZ)-amino acids, with anhydrous trifluoroacetic acid (TFA); and 3) the conversion to phenylthiohydantoin (PTH)-amino acids with aqueous TFA. Other Bronstad acids, such as HC1 in methanol and triethylamine / acetic acid in acetonitrile have also been shown to produce a similar chemical transition. In step (2) anhydrous conditions are typically employed with TFA to minimize acid hydrolysis at non-target sites. To avoid unwanted reactions, mild alkaline conditions (G.C. Barrett, et.al., Tetrahedron Eett. (1985), 26, 36, 4375-4378), and Eewis acid conditions (Matsunaga, et. al., Biomedical Chromatography, vol. 10, 95 (1996)) have been studied for the cleavage of peptide bonds within the context of Edman degradation. These studies are instructive for the current invention, since strong Bronstad acid conditions are known to damage DNA via by protonation of nitrogens on the purine rings, followed by depurination and subsequent hydrolysis of the abasic phosphodiester bond.

[0148] Okiyama, et.al., (Analytica Chimica Acta 429 (2001) 293-300) report a modified Edman procedure using BF3-methanol as a cyclization / cleavage / conversion reagent. Matsunaga, et. al., Biomedical Chromatography, vol. 10, 95 (1996) reported on boron trifluoride-etherate as an efficient reagent for cyclization / cleavage reaction. Mayer et.al., cataloged the Eewis acid strength of several reagents. Boron trifluoride-etherate, one of the Lewis acids, was identified as an efficient acid in the Edman sequencing method, affording the retention of the N-terminal amino acid configuration at the cyclization / cleavage reaction from peptides. Other Lewis acids with lower acid strength (or higher acid strength) that may also be useful for cleavage while safe for nucleic acids include: mono-, di- or triarylboranes, where aryl substituents comprise the groups: -H, -F, -Cl, -Br, -I, -C6H5, -C6F5, -(4-Me2N-C6H4), -(4-MeO-C6H4), - 1(4-Me-C6H4), -(4-F-C6H4), -(4-C1-C6H4), -(2,4,6-F3-C6-H2), -(3,4,5-F3-C6-H2), -(2,4,6-Me3-C6H2), -(3-F-C6H4), -(2-F-C6H4), -(4-Me-2-F-C6H3), and these may be mixed in any possible combinations to affect the Boron electron pair acidity. Aryl substituents may comprise halo, H, alkyl, alkoxy, dialkylamino or additional aryl groups, at any or multiple positions on the ring. Aryl substituents may also comprise heteroatoms S, O, P, N, Se, B. Aryl substituents may also comprise multicyclic fused ring structuresAttorney Docket No: 062954-509001 WO including but not limited to: napthyl, anthracyl, indolyl, carbazole, benzofuran, dibenzofuran, quinoline, pyrene, benzopyrene, phenanthrene and partially or fully fluorinated versions. Other possible substituents on the boranes comprise alkyl groups: methyl, ethyl, propyl, butyl, pentyl, hexyl, heptyl, octyl, nonyl, decyl, undecyl, dodecyl, C13-C18, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cycloheptyl, cyclooctyl, and perfluorinated or partially fluorinated versions thereof, which may be mixed in any possible combination with other classes of substituents. These Lewis acids may be used alone, in solution or complexed with a variety of Lewis bases, comprising alcohols, methanol, ethanol, ethers, diethylether, tetrahydrofuran, pyridine, substituted pyridines, nitriles, acetonitrile, sulfides, dimethylsulfide.

[0149] Metal-containing Lewis acids that may be useful for cleavage while safe for nucleic acids comprise: Bi, Al, Fe, Cu, Ce, Co, Ni, In, Ir, Dy, Yb, La, Eu, Gd, Tm, Ho, Hf, Er, Nd, Sm, Lu, V, Ti, Zn, Sn, Sc, Zr, Si, Ga, Tl, Pr, Y, Pm, Mn, Ta, Nb, Tb, a transition metal, a lanthanide. Metal-containing Lewis acid catalysts may include counterions, ligands or substituents comprising: -F, -Cl, -Br, -I, -OAc, trifluoroacetate, -OTf, carboxylates, alkyl or aryl sulfonates, fluoroalkylsulfonates, nitriles, alkyl groups, aryl groups, sulfides, oxo groups, water, a heterocycle, a cyclooctadiene, an arylsulfonate, a fluoroarylsulfonate, hydroxyl groups, alkoxides, amines, phosphines, cyclopentadienyl, BF4-, PF6-, SbF6-. TMSOTf and bis(pertrifluoromethylcatecholato)silane are examples of silicon Lewis acids which may be useful.Cleavage efficiency requirements for protein sequencing.

[0150] Lewis acid-mediated cleavage reactions need not achieve quantitative conversion to be useful for protein sequencing applications. The iterative nature of Edman degradation, combined with the ability to align partial sequence reads to reference proteomes, provides robustness to incomplete cleavage at individual cycles. As demonstrated in WO 2024 / 015875, protein identification workflows can achieve high- confidence results even when per-cycle cleavage efficiencies range from approximately 40% to 98% depending on substrate complexity and reaction conditions, and where partial sequence information (e.g., 20 of 30 amino acids successfully sequenced, representing approximately 67% overall information capture) enables unambiguous protein identification.

[0151] Phosphorus containing Lewis acids as described in Bayne, J.M.; Stephan, D.W. Chem. Soc. Rev., 2016,45, 765-774 may also be useful for cleavage while safe for nucleic acids.

[0152] Frustrated Lewis Pairs (FLP) comprised of mixtures of Lewis acids and Lewis bases that are prevented from forming classical adducts due to steric hindrance may be useful for cleavage while safe for nucleic acids.

[0153] An exemplary method for Edman degradation using a commercially -viable Lewis acid is provided in the Examples herein.Attorney Docket No: 062954-509001 WORecode tags

[0154] Some embodiments include generation or use of a recode tag. Disclosed herein, in some embodiments, are recode tags. The recode tag may be a part of a binding agent. The recode tag may correspond with a binding agent. For example, the recode tag may convey information about a molecule (e.g. an amino acid or PTM) to which the binding agent binds. The recode tag may include a nucleic acid such as a recode nucleic acid. In some embodiments, the recode nucleic acid comprises DNA or RNA. In some embodiments, the recode tag is a DNA sequence.

[0155] In some embodiments, the recode tag may comprise additional sequence elements, such as PIN nucleotides, PCR primer nuclotides, SBS primer sequences, a sample index, a unique molecular identifier (UMI), a universal priming site, a CRISPR protospacer adjacent motif (PAM) sequence, a universal assembly sequence, a nuclease restriction sequence, a sequence, nucleotides, or modified nucleotides, a sequence helpful in assembling recode blocks into memory oligos, or any combination thereof.

[0156] In some embodiments, the recode tag is an RNA sequence. The recode nucleic acid may be useful to encode amino acid information in a nucleic acid. The recode tag may be used in a method described herein, such as a method for determining protein information such as amino acid identity.Recode units

[0157] Some embodiments include generation or use of a recode unit. Disclosed herein, in some embodiments, are recode units. In some embodiments, the recode unit comprises the joined information of a cycle tag and a location nucleic acid. The recode unit may be a part of a CRC. In some embodiments, a recode unit may comprise additional sequence elements, such as PIN nucleotides, PCR primer nuclotides, SBS primer sequences, a sample index, a unique molecular identifier (UMI), a universal priming site, a CRISPR protospacer adjacent motif (PAM) sequence, a universal assembly sequence, a nuclease restriction sequence, a sequence, nucleotides, or modified nucleotides helpful in assembling recode blocks into memory oligos, or any combination thereof. In some embodiments the recode unit may comprise the joined information of a sequence useful for joining nucleic acids and a location nucleic acid.Recode blocks

[0158] Some embodiments include generation or use of recode block. Disclosed herein, in some embodiments, are recode blocks. The recode block may include a cycle tag, a location tag, and a recode tag or any combination therof, or reverse complements thereof. For example, the recode block may include a cycle tag, a reverse complement of a recode tag, and a reverse complement of a location tag. The recode block may include information corresponding to the cycle tag, the location tag, and the recode tag. For example, the recode block may include a cycle nucleic acid, a cycle nucleic acid sequence, or a reverseAttorney Docket No: 062954-509001 WO complement thereof, may include a location nucleic acid, a location nucleic acid sequence, or a reverse complement thereof, and may include a recode nucleic acid, a recode nucleic acid sequence, or a reverse complement thereof. The recode block may be useful for joining into a memory oligonucleotide, either of which may convey information about relative amino acid position and identity within a protein. The recode block may be used in a method described herein, such as a method for determining protein information such as amino acid location or identity.

[0159] In some embodiments, a recode block comprises information of (a) relative position of an amino acid in a polymer chain, (b) adjacency of an amino acid of an immobilized peptide to a location on a solid support, and (c) identity of the amino acid. In some embodiments, a recode block comprises information of (a) relatative position of an amino acid in a polymer chain, (b) relative location of a protein of origin, and (c) identity of the amino acid. In some embodiments a recode block is created by interaction between a cycle tag of a coupled-conjugate-complex and the location nucleic acid of a spatially adjacent location nucleic acid. Typically, a recode block is a chimeric nucleic acid molecule that contains the information relating the workflow cycle, the amino acid identity or class of amino acid, and / or the relative loation information of the macromolecule of origin.

[0160] Further, the recode block may comprise information and / or nucleic acid sequence to direct assembly of a memory oligo, and / or amplify the recode block. A recode block may be formed by utilizing an extension-ligation method to transfer information from or between the recode tag, cycle tag, and location tag, or via a ligation reaction under appropriate conditions in the presence of ligase and ligation oligo(s). A recode block may be formed by joining seqenences or reverse complement sequences of cycle tags, location tags, and / or recode tags. The format of a recode block is not necessarily a nucleic acid. It may also take the form of mass tags that could be used to assign identity for cycle, amino acid, and relative location information, or other modalities that represent the information of cycle, amino acid, and location, and are amenable to group that information for analysis.Transfer of information

[0161] Disclosed herein, in some embodiments, are methods which include transferring information. For example, a method may include transferring information of the recode nucleic acid to the cycle nucleic acid of the immobilized conjugate complex to generate a recode block. The transfer of information may form a recode block, or may be used to form a memory oligonucleotide. The transfer of information may be included in a method described herein, such as a method for determining protein information such as amino acid location or identity.

[0162] In some embodiments, said transferring information comprises performing a nucleic acid sequencebased amplification, for example to generate the sequence of the recode nucleic acid or the sequence of theAttomey Docket No: 062954-509001 WO cycle nucleic acid. In some embodiments, said transferring information comprises performing polymerase chain reaction (PCR) to generate the sequence of the recode nucleic acid or the sequence of the cycle nucleic acid. In some embodiments, the PCR comprises real-time PCR, digital PCR, multiplex PCR, nested PCR, hot-start PCR, touchdown PCR, or quantitative PCR. In some embodiments, said transferring information comprises performing or conducting a ligase chain reaction, a helicase-dependent amplification, a strand displacement amplification, a loop-mediated isothermal amplification, a rolling circle amplification, a recombinase polymerase amplification, a nicking enzyme amplification reaction, a whole genome amplification, a transcription-mediated amplification, a multiple displacement amplification, or multiple annealing and looping-based amplification cycles, for example to generate the sequence of the recode nucleic acid or the sequence of the cycle nucleic acid. The amplification or other procedure may be to generate the sequence of the recode nucleic acid, the sequence of the cycle nucleic acid, a reverse complement, or a combination thereof. In some embodiments, the information of the recode nucleic acid comprises a sequence of the recode nucleic acid or a reverse complement of the sequence of the recode nucleic acid.

[0163] In some embodiments, the transfer of information involves a polymerase chain reaction. In some embodiments, the transfer of information involves a reverse transcription polymerase chain reaction. In some embodiments, the transfer of information involves a real-time polymerase chain reaction. In some embodiments, the transfer of information involves a digital polymerase chain reaction. In some embodiments, the transfer of information involves a multiplex polymerase chain reaction. In some embodiments, the transfer of information involves a nested polymerase chain reaction. In some embodiments, the transfer of information involves a hot-start polymerase chain reaction. In some embodiments, the transfer of information involves a touchdown polymerase chain reaction. In some embodiments, the transfer of information involves a quantitative polymerase chain reaction. In some embodiments, the transfer of information involves a ligase chain reaction. In some embodiments, the transfer of information involves a helicase-dependent amplification. In some embodiments, the transfer of information involves a strand displacement amplification. In some embodiments, the transfer of information involves a loop-mediated isothermal amplification. In some embodiments, the transfer of information involves a rolling circle amplification. In some embodiments, the transfer of information involves a recombinase polymerase amplification. In some embodiments, the transfer of information involves a nicking enzyme amplification reaction. In some embodiments, the transfer of information involves a whole genome amplification. In some embodiments, the transfer of information involves a transcription-mediated amplification. In some embodiments, the transfer of information involves a multiple displacement amplification. In some embodiments, the transfer of information involves a multiple annealing and looping-Attorney Docket No: 062954-509001 WO based amplification cycles. In some embodiments, the transfer of information involves a nucleic acid sequence-based amplification.

[0164] In some embodiments, said transferring information comprises joining the recode nucleic acid or a reverse complement of the recode nucleic acid with the cycle nucleic acid.Sequential Proximity Ligation for Enhanced Specificity

[0165] In some embodiments, two sequential proximity ligation reactions are performed to enhance amino acid identification specificity. In the first ligation, a first binding agent comprising a first splint oligonucleotide contacts the recode unit complex. The first splint oligonucleotide comprises a proximity hybridization sequence that overlaps the cycle tag, an amino acid identification sequence, and an amino- acid-specific second proximity hybridization sequence. A first free ligation oligonucleotide comprising the complement of the amino acid identification sequence and the second proximity hybridization sequence is ligated to form a first partial recode block.

[0166] In the second ligation, a second binding agent comprising a second splint oligonucleotide contacts the first partial recode block. The second splint oligonucleotide hybridizes to the amino-acid-specific second proximity hybridization sequence from the first ligation. If the first binding agent incorrectly identified the amino acid, this hybridization will not occur efficiently. A second free ligation oligonucleotide comprising a universal assembly sequence is ligated to form a full recode block.

[0167] This sequential approach with amino-acid-specific compatibility between steps reduces noise from non-specific binding and contaminating nucleic acids. Only complexes that successfully complete both ligations contain the universal assembly sequence required for downstream processing.Side Chain Derivatization for Amino Acid Identification

[0168] In some embodiments, amino acid side chains are derivatized with nucleic acid tags (AA tags) prior to the N-terminal sequencing process. The AA tags may comprise an amino acid identification sequence encoding the side chain identity and a joining sequence for downstream ligation. Derivatization may target lysine, cysteine, aspartic acid, glutamic acid, tyrosine, phosphoserine, phosphothreonine, or combinations thereof using chemistry appropriate to each side chain functional group.

[0169] Following the iterative sequencing cycles, amino acids that were derivatized comprise both a recode unit complex and an AA tag. These components may be joined in solution through ligation, creating recode blocks that contain both positional information from the cycle tag and identity information from the AA tag. Selective amplification using primers that hybridize to both the cycle tag and AA tag regions enriches for successfully joined products.Attorney Docket No: 062954-509001 WO

[0170] This approach provides an amino acid fingerprint that can aid in protein identification and enables detection of specific post-translational modifications through selective derivatization of modified residues.Joinins

[0171] Disclosed herein, in some embodiments, are methods which include joining. For example a recode nucleic acid or a reverse complement thereof may be joined with a cycle nucleic acid or a reverse complement thereof. The joining may form a recode block, or may be used to form a memory oligonucleotide. The joining may be included in a method described herein, such as a method for determining protein information such as amino acid location or identity.

[0172] In some embodiments, joining comprises enzymatic ligation. In some embodiments, joining comprises splint ligation. In some embodiments, joining comprises chemical ligation. In some embodiments, joining comprises template-assisted ligation. In some embodiments, joining comprises the use of a ligase enzyme. In some embodiments, joining comprises the use of a splint oligonucleotide. In some embodiments, joining comprises the use of a catalyst. In some embodiments, joining comprises the use of a bridging molecule. In some embodiments, joining comprises the use of a condensation agent. In some embodiments, joining comprises the use of a coupling reagent. In some embodiments, joining comprises the use of a polymerase enzyme. In some embodiments, joining comprises the use of a complementary nucleic acid sequence. In some embodiments, joining comprises the use of a nicking enzyme. In some embodiments, joining comprises the use of a nucleic acid modifying enzyme. In some embodiments, joining comprises the use of a recombinase. In some embodiments, joining comprises the use of a strand-displacing polymerase. In some embodiments, joining comprises the use of a single-strand binding protein. In some embodiments, joining comprises a click chemistry reaction. In some embodiments, joining comprises a phosphodiester bond formation. In some embodiments, joining comprises a peptide nucleic acid-mediated ligation. In some embodiments, each binding agent comprises recode tags with a unique nucleic acid sequence. In some embodiments, a plurality of binding agents comprises recode tags with the same nucleic acid sequence. In some embodiments, binding agents comprises recode tags which may have a unique sequence portion and a common sequence portion.

[0173] In some embodiments, joining the recode nucleic acid or a sequence of the recode nucleic acid with the cycle nucleic acid or a sequence of the cycle nucleic acid to generate a recode block comprises: (i) joining the recode nucleic acid with the cycle nucleic acid, (ii) joining the recode nucleic acid with a sequence of the cycle nucleic acid, (iii) joining a sequence of the recode nucleic acid with the cycle nucleic acid, or (iv) joining a sequence of the recode nucleic acid with a sequence of the cycle nucleic acid. Some embodiments include performing a nucleic acid sequence-based amplification to generate the sequence of the recode nucleic acid or the sequence of the cycle nucleic acid. Some embodiments include performingAttorney Docket No: 062954-509001 WO polymerase chain reaction (PCR) to generate the sequence of the recode nucleic acid or the sequence of the cycle nucleic acid. In some embodiments, the PCR comprises real-time PCR, digital PCR, multiplex PCR, nested PCR, hot-start PCR, touchdown PCR, or quantitative PCR. Some embodiments include performing or conducting a ligase chain reaction, a helicase-dependent amplification, a strand displacement amplification, a loop-mediated isothermal amplification, a rolling circle amplification, a recombinase polymerase amplification, a nicking enzyme amplification reaction, a whole genome amplification, a transcription-mediated amplification, a multiple displacement amplification, or multiple annealing and looping-based amplification cycles, to generate the sequence of the recode nucleic acid or the sequence of the cycle nucleic acid.

[0174] In some embodiments, the joining comprises enzymatic ligation, splint ligation, chemical ligation, template-assisted ligation, use of a ligase enzyme, use of a splint oligonucleotide, use of a catalyst, use of a bridging molecule, use of a condensation agent, use of a coupling reagent, use of a polymerase enzyme, use of a complementary nucleic acid sequence, use of a nicking enzyme, use of a nucleic acid modifying enzyme, use of a recombinase, use of a strand-displacing polymerase, use of a single-strand binding protein, a click chemistry reaction, a phosphodiester bond formation, or a peptide nucleic acid-mediated ligation.

[0175] Some embodiments include contacting an additional immobilized amino acid complex with a second binding agent. In some embodiments, the binding agent and the second binding agent comprise distinct recode tags having different recode nucleic acids from each other. In some embodiments, the binding agent and the second binding agent comprise recode tags having identical recode nucleic acids as each other. In some embodiments, the binding agent and the second binding agent comprise distinct recode tags having recode nucleic acids that have different sequences from each other, and that have a portion of the recode nucleic acids that are identical.

[0176] In some embodiments, said transferring information comprises joining or combining the recode nucleic acid, a sequence of the recode nucleic acid, or a reverse complement of the sequence of the recode nucleic acid with the cycle nucleic acid, a sequence of the cycle nucleic acid, or a reverse complement of the sequence of the cycle nucleic acid, to generate a recode block.

[0177] In some embodiments, the joining comprises modified nucleotides, and / or modified nucleotide triposphates. In some embodimnets, said modified nucleotides or nucleotide triphosphates are incorporated by an enzyme or other means as described herein during the process of joining information. In some embodiments the enzyme is a polymerase, or a ligase. In some embodiments the modification comprises a functional moiety. In some embodiments functional moiety is used for purification, separation, concentration, washing, buffer exchange, or other beneficial action related to the determination of protein information.Attorney Docket No: 062954-509001 WO

[0178] In some embodiments, binding agents are covalently stabilized to amino acid complexes after binding but prior to information transfer. Covalent stabilization enables stringent washing to remove non- specifically bound species while retaining specifically bound binding agents.

[0179] Several chemistries may be employed for covalent stabilization. Photo-crosslinking may be performed using photoreactive groups incorporated into the binding agent or target, such as p-benzoyl-L- phenylalanine, diazirine derivatives, or benzophenone derivatives, which form covalent bonds upon UV irradiation. UV exposure dose may be optimized to minimize damage to nucleic acid components. Click chemistry between alkyne and azide groups on the binding agent and target may be performed with or without copper catalysis. Chemical crosslinkers such as di(sulfosuccinimidyl)suberate may react with amine groups on both the binding agent and target to form covalent linkages.

[0180] The covalent stabilization step is performed after binding agent binding and before information transfer operations such as ligation or other joining reactions. Following stabilization, washing steps may be performed under more stringent conditions than would be compatible with non-covalent binding alone.Assembly of Memory Oligos from solution-phase Recode Blocks

[0181] Some embodiments include generation or use of a memory oligonucleotide. A method for in assembly of information from solution-phase recode blocks is disclosed herein. In some embodiments, purification, separation, washing, buffer exchange, and / or concentration of recode blocks is performed prior to assembly and / or amplification of recode blocks, using commonly employed methods and according to methods described herein.

[0182] Enzyme-combined isothermal replication and / or amplification methods have been described previously (Tan, et. al., Biochemistry. 2008 September 23; 47(38): 9987-9999. doi:10.1021 / bi800746p; Maples et.al. US 9,689,031). Nucleic acid assembly methods utilizing combinations of enzymes, such as nickases, endonucleases, polymerases, exonucleases, RNAaseH, have also been described for immobilized nucleic acids, including recode block nucleic acids.

[0183] Assembly of solution-phase nucleic acids can be designed to occur in a stepwise or parallel manner, and in each case, can be designed to produce ordered or random assemblies of information. Methods for stepwise and parallel assembly of ordered information are shown in figures of references included herein. In some embodiments, recode blocks share a common sequence that allows enzymatic ligation via a common splint. The issue with this simple parallel method for random assembly is that cyclization of the memory oligo constructs may occur, thereby limiting throughput. Ordered methods avoid this pitfall, but may be kinetically inefficient. Stepwise random assembly avoids throughput limitations and increases efficiency over ordered assembly methods. Therefore, several examples of efficient stepwise methods to assemble co-localized nucleic acids in a random fashion have been disclosed in references included herein.Attorney Docket No: 062954-509001 WOA person skilled in the art will recognize that these are applicable to assemble solution-phase nucleic acids in a stepwise random assembly, as well.

[0184] There are three primary operations required for efficient stepwise random assembly of memory oligos: Initiation, Extension, and Termination.

[0185] Initiation. Initiation may be accomplished by providing a forward primer complementary to the recode block (FIG. 5, n.3, n.4). To avoid priming all recode blocks, a mixture of binding agents may be employed, some having sequence complementary to the primer sequence, and some not. The ratio of the binding agents with complementray primer sequence to those without complementary primer sequence will influence the distribution of assembled memory oligos. Optionally, the order of operations shown in FIG. 5 may be varied. For example, the initiation extension step can be completed while the recode tags are immobilized to a purification resin, to facilitate washing excess primer, prior to initiating the first extension step. Additional elements may be included in the assembly initiation oligo. These include but are not limited to PIN nucleotides, SBS primer sequence, a sample index, a spacer, a unique molecular identifier (UMI), a universal priming site, a CRISPR protospacer adjacent motif (PAM) sequence, or any combination thereof.

[0186] Extension. U-AS is a unifying assembly sequence for memory oligo assembly. The recode block nucleic acid may be fully or partially RNA to allow RNAase chew back to provide a single-stranded oligonucleotide that facilitates subsequent extension reactions. Alternately, one or more nicking endonuclease sites may be designed into the hyb tag and / or U-AS to provide a single-stranded oligonucleotide that facilitates subsequent extension reactions. FIG. 6 shows the interface of the memory oligo, and illustrates the operations of step (n) of Process 200. In FIG. 6, U-AS is a unifying assembly sequence for memory oligo assembly and a preferred nucleic acid composition is depicted. The assembly is designed such that there is a limited number of U-AS sequences per analyte. This number could range from about 0.00001 to about 2. During the initiation and extension steps, U-AS’ from the recode block interacts with U-AS of the primer. In some embodiments events are as follows: 1) hybridization of U-AS to U-AS’ ; 2) extension of U-AS by a polymerase; 3) RNAse-H action on the newly formed double-stranded RNA-DNA duplex to expose a new U-AS 3’ end; and 4) interaction of a U-AS’ sequence of a next recode block with the newly formed single-stranded U-AS’.

[0187] PNA is not effective with existing enzymes to serve as a template for PCR or ligation, but is highly capable to create strong specific heteroduplexes with various nucleic acids including DNA. Fouz et.al., Molecules 2020, 25, 786; Pezo, et.al., Angew. Chem., Int. Ed., 2013, 52, 8139-8143; Duffy et al. BMC Biology (2020) 18:112. Therefore, transfer of information from PNA into DNA is possible.

[0188] Termination. Nucleic acid constructs used for assembly of recode blocks may be designed so as to terminate assembly, or limit assembly to a defined number, or a defined range of assembled recodeAttorney Docket No: 062954-509001 WO blocks. Termination may facilitate subsequent PCR amplification of memory oligos by providing reverse primer sequence, UMI, PIN, sample index, or other information.

[0189] There are several favorable features of the disclosed method. Note that the events may be completed with or without washing or solution exchange between events. Temperature may be varied or held constant. Temperature may be used to modulate specificity. Designed mismatches of nucleobases may be used to optimize relative Tm and specificity. The method is highly efficient to assemble nucleic acid information blocks in random order, avoiding circularization and termination events.

[0190] In some embodiments, a combination of exonuclease and polymerase with U-AS facilitate memory oligo assembly.

[0191] In some embodiments the memory oligo or a representative oligonucleotide containing location, amino acid identity and positional information is assembled in solution after the recode blocks from one or more cycles of recode block assembly are completed, and prior to subsequent cycles. This may be accomplished through contacting the recode blocks with primer.

[0192] In some embodiments, more than one U-AS sequence may be employed. Different U-AS sequences may be applied in any order, and may be useful to normalize memory oligo length.

[0193] In some embodiments, Extension occurs in serial steps that may include temperature modulation, and in others embodiments the process may be run in a single isothermal step.

[0194] In some aspects, base pair mismatches may be used to adjust binding energies (Tm) and / or specificity of nucleic acids of the assembly process.

[0195] In some embodiments, endonuclease nickase or similar enzymes may replace RNAse-H to provide a single-stranded U-AS.

[0196] A method for assembly of information from solution-phase recode blocks is disclosed herein. In some embodiments, purification, separation, washing, buffer exchange, and / or concentration operations are performed prior to assembly, concatenation, and / or amplification of recode units or memory oligos, using commonly employed methods and according to methods described herein.

[0197] A potential issue with simple parallel methods for concatentating short DNA contracts using universal assembly sequences is that termination of assembly may occur via early cyclization of the growing chain, thereby limiting the length of the product. Ordered methods avoid this pitfall, but may be kinetically inefficient. Semi-random assembly avoids throughput limitations and increases efficiency over ordered assembly methods, while addressing early cyclization. Assembly of short solution-phase nucleic acids, specifically in this context recode units, to provide kb long chimeric sequence of DNA having cycle, location, and amino acid information by appropriate sequence design and processes avoids circularization of constructs. Longer lengths may be advantageous to improve efficiency and reduce the cost of DNAAttorney Docket No: 062954-509001 WO sequencing associated with determination of peptide sequences. Methods for assembly of short nuleic acids has been described (Horspool et al. BMC Research Notes 2010, 3:291).

[0198] Appropriate oligo design is a primary requirement for efficient semi-random assembly of memory oligos. Recode units as disclosed herein may be configured for semi-random assembly as follows:RAS-UAS-Cycle-Location-UAS-RASFormula IV

[0199] where UAS represents Univeral Assembly Sequence and RAS represents Random Assembly Sequence.

[0200] RAS segments of the recode unit denote random or biased random sequence. The length of each RAS section affects the efficiency of the assembly by modulating the concetration of any given pair of recode units and the probability of cyclization. Random sequence may be synthesized by applying equimolar mixtures of nucleotide bases within a single step of oligonucleotide synthesis, for example, equal molar ratios of the 4 natural DNA bases A,G,C, or T could be incorporated in a growing oligonucleotide strand. "Wobbles" are equimolar mixtures of two or more different bases at a given position within an oligonucleotide sequence. "Biased Wobbles" are non-equimolar mixtures of two or more different bases at a given position within an oligonucleotide sequence. Biased wobbles may be synthesized by applying specific non-equimolar mixtures of nucleotide bases within a single step of oligonucleotide synthesis. For example, a molar ratio of 10:1:1:0.5 of the 4 natural DNA bases A,G,C, or T could be incorporated in a growing oligonucleotide strand. This may be useful to tune the diversity distribution of the pool of recode unit olginuceotides. For example, the distribution of the sequences from a 10:1:1:0.5 ration is biased to those that have more A’s and fewer T’s, providing a higher concentration of each biased random sequence in an assembly operation, and therefore faster kinetics - but higher probability for circularization. A ratio of 1:10:10:1 for A, T, G, C may affect the Tm and the combinatorial complexity of resulting nucleotide sequences, providing additional value and / or flexibility in the assembly operation.

[0201] In some embodiments the molar ratio of each nucleotide base during synthesis may include any appropriate number between 0 and 1.

[0202] In some embodiments, the wobbles or biased wobbles are contiguous sequence. In some embodiments, the wobbles or biased wobbles are not contiguous.

[0203] In some embodiments, the RAS length at each end of the recode unit is equal. In some embodiments, the RAS length at each end of the recode unit is not equal. In typical embodiments, the length of RAS segements is between 0 and 10 nucleotides, but may be as many as 50 nucleotides.

[0204] In some embodiments, the assembly operation(s) comprise enzymatic ligation using a set of splints having random, or biased random, sequence composition complementary to the RAS segments of recode units.Attorney Docket No: 062954-509001 WO

[0205] In some embodiments, Universal Assembly Sequence flanks, or is interspersed with, RAS. This may be advantageous to provide additional hybridization energy and helical structure useful in ligation reactions.

[0206] In some embodiments, single stranded DNA may be cleavage by restriction endonucleases as described by Nishigaki, et.al. (Nucleic Acids Res. 1985 Aug 26; 13(16): 5747-5760). In some embodiments this provides a free 3’ end that can be ligated.

[0207] In some embodiments, the 5-‘ end of an oligo nucleotide may be phosphorylated in preparation for ligation by using a phosphatase, for example, the commercially-available T4 Polynucleotide Kinase from NEB, (PNK, cat#M0201).

[0208] An exemplary method to efficiently create memory oligos comprising ligation operations is provided in the Examples herein.Process 100 - Identity & Positional Information

[0209] Process 100 can relate to identifying amino acid identity and positional information. Any aspect or combination of aspects from Process 100 may be included in a method herein.

[0210] Disclosed herein, in some embodiments, are methods for determining identity and positional information of an amino acid residue of a peptide coupled to a solid support, the method comprising:(a) providing the peptide to the solid support, the peptide coupled to the solid support such that a N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions;(b) providing a chemically -reactive conjugate, the chemically-reactive conjugate comprising: (x) a location nucleic acid, (y) a reactive moiety for binding the N-terminal amino acid residue of the peptide, and (z) an immobilizing moiety for immobilization to the solid support;(c) contacting the peptide with the chemically-reactive conjugate, thereby coupling the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex;(d) immobilizing the conjugate complex to the solid support via the immobilizing moiety;(e) cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby exposing the next amino acid residue as a N-terminal amino acid residue on the cleaved peptide and providing an immobilized location complex, the immobilized location complex comprising the cleaved and separated N-terminal amino acid residue and a location nucleic acid;(f) providing a second chemically-reactive conjugate, the chemically-reactive conjugate comprising: (xx) a cycle tag comprising a cycle nucleic acid associated with a cycle number, and (yy) a reactive moiety for binding and cleaving the N-terminal amino acid residue of the peptide, and optionally (zz) a moiety for joining to an optional second solid support;Attorney Docket No: 062954-509001 WO(g) contacting the N-terminal amino acid residue of the peptide with the chemically-reactive conjugate, thereby coupling the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a coupled-conjugate-complex;(h) forming a transfer complex comprising an immobilized location complex and a coupled- conjugate-complex, thereby bringing a location nucleic acid into proximity with a cycle tag nucleic acid within the formed transfer complex;(i) within each formed transfer complex, joining a location nucleic acid or a reverse complement thereof to a cycle tag nucleic acid to form a recode unit, or otherwise joining information of the location nucleic acid and the cycle tag nucleic acid, thereby creating an immobilized recode unit, each recode unit corresponding with a formed transfer complex;(j) cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby exposing the next amino acid residue as a N-terminal amino acid residue on the cleaved peptide and liberating the immobilized recode unit complex from the solid support, the liberated recode unit complex comprising the cleaved and separated N-terminal amino acid residue, cycle and location information.(k) contacting the recode unit complex with a binding agent, the binding agent comprising: a binding moiety for preferentially binding to the PTH-amino acid moiety of the recode unit complex, and a recode tag comprising a recode nucleic acid corresponding with the binding agent, thereby forming an affinity complex, the affinity complex comprising a recode unit complex and the binding agent, and thereby bringing the recode unit nucleic acid into proximity with the recode tag within the affinity complex;(l) transferring information of the recode tag nucleic acid to the recode unit to generate a recode block;(m) obtaining sequence information of the recode block; and(n) based on the obtained sequence information, determining identity and positional information of an amino acid residue of the peptide.

[0211] In some embodiments, methods for determining identity and positional information of an amino acid residue of a peptide are disclosed herein. The peptide may be coupled to a solid support. In some embodiments, the method comprises providing the peptide to the solid support. In some embodiments, the peptide is coupled to the solid support such that a N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions.

[0212] In some embodiments, the method additionally comprises providing a chemically-reactive conjugate. In some embodiments, the chemically-reactive conjugate comprises: (x) a location nucleic acid, (y) a reactive moiety for binding the N-terminal amino acid residue of the peptide, and (z) an immobilizing moiety for immobilization to the solid support.Attorney Docket No: 062954-509001 WO

[0213] In some embodiments, the method additionally comprises contacting the peptide with the chemically-reactive conjugate, thereby coupling the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex.

[0214] In some embodiments, the method additionally comprises immobilizing the conjugate complex to the solid support via the immobilizing moiety.

[0215] In some embodiments, the method additionally comprises cleaving and thereby separating the N- terminal amino acid residue from the peptide. In some embodiments, cleaving and thereby separating the N-terminal amino acid residue from the peptide exposes the next amino acid residue as a N-terminal amino acid residue on the cleaved peptide and provides an immobilized location complex. In some embodiments, the immobilized location complex comprises the cleaved and separated N-terminal amino acid residue and a location nucleic acid.

[0216] In some embodiments, the method additionally comprises providing a second chemically-reactive conjugate. In some embodiments, the chemically-reactive conjugate comprises (xx) a cycle tag comprising a cycle nucleic acid associated with a cycle number. In some embodiments, the chemically-reactive conjugate comprises (yy) a reactive moiety for binding and cleaving the N-terminal amino acid residue of the peptide. In some embodiments, the chemically-reactive conjugate additionally comprises (zz) a moiety for joining to an optional second solid support.

[0217] In some embodiments, the method additionally comprises contacting the N-terminal amino acid residue of the peptide with the chemically-reactive conjugate, thereby coupling the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a coupled-conjugate-complex.

[0218] In some embodiments, the method additionally comprises forming a transfer complex comprising an immobilized location complex and a coupled-conjugate-complex, thereby bringing a location nucleic acid into proximity with a cycle tag nucleic acid within the formed transfer complex.

[0219] In some embodiments, the method additionally comprises, within each formed transfer complex, joining a location nucleic acid or a reverse complement thereof to a cycle tag nucleic acid to form a recode unit, or otherwise joining information of the location nucleic acid and the cycle tag nucleic acid, thereby creating an immobilized recode unit. In some embodiments, each recode unit corresponds with a formed transfer complex.

[0220] In some embodiments, the method additionally comprises cleaving and thereby separating the N- terminal amino acid residue from the peptide, thereby exposing the next amino acid residue as a N-terminal amino acid residue on the cleaved peptide and liberating the immobilized recode unit complex from the solid support. In some embodiments, the liberated recode unit complex comprises the cleaved and separated N-terminal amino acid residue, cycle and location information.Attorney Docket No: 062954-509001 WO

[0221] In some embodiments, the method additionally comprises contacting the recode unit complex with a binding agent. In some embodiments, the binding agent comprises: a binding moiety for preferentially binding to the PTH-amino acid moiety of the recode unit complex, and a recode tag comprising a recode nucleic acid corresponding with the binding agent, thereby forming an affinity complex, the affinity complex comprising a recode unit complex and the binding agent, and thereby bringing the recode unit nucleic acid into proximity with the recode tag within the affinity complex.

[0222] In some embodiments, the method additionally comprises transferring information of the recode tag nucleic acid to the recode unit to generate a recode block.

[0223] In some embodiments, the method additionally comprises obtaining sequence information of the recode block.

[0224] In some embodiments, the method additionally comprises, based on the obtained sequence information, determining identity and positional information of an amino acid residue of the peptide.

[0225] Disclosed herein, in some embodiments, cleaving the N-terminal amino acid residue from the peptide exposes a next amino acid residue as a N-terminal amino acid residue on the cleaved peptide. In some embodiments, the reactive moiety of the chemically-reactive conjugate cleaves the N-terminal amino acid residue from the peptide.

[0226] In some embodiments, steps (b) through (k) are repeated for each subsequent amino acid of the peptide.

[0227] In some embodiments, the method comprises washing the immobilized amino acid complex before said contacting the immobilized amino acid complex with a binding agent.

[0228] In some embodiments, a likely three-dimensional structure of the peptide based on the sequence information is determined.

[0229] In some embodiments, the recode nucleic acid comprises DNA or RNA or modified or unmodified DNA or RNA. In some embodiments, the recode nucleic acid comprises a hybridization-capable code. In some embodiments, the recode nucleic acid comprises a non-colliding code in an additive vector space. In some embodiments, the cycle nucleic acid comprises DNA or RNA.

[0230] In some embodiments, the binding moiety comprises a peptide, antibody, antibody fragment, antibody derivative, or aptamer. In some embodiments, the binding moiety binds to a natural amino acid, a post-translationally modified amino acid, a derivatized version of an amino acid, a derivatized or stabilized version of a post-translationally modified amino acid, a synthetic amino acid, an amino acid with a specific side chain, an amino acid with a phosphorylated side chain, an amino acid with a glycosylated side chain, an amino acid with a methylation modification, or a D-amino acid, phenylthiohydantoin (PTH) derivative or anilinothiazolinone (ATZ) derivative of an amino acid, or binds to a combination thereof. In someAttorney Docket No: 062954-509001 WO embodiments, the binding moiety binds covalently or non-covalently to the immobilized amino acid complex.

[0231] In some embodiments, the solid support comprises a bead, a plate, or a chip.In some embodiments, the solid support comprises a glass slide, silica, a resin, a gel, a hydrogel, a membrane, polystyrene, a metal, nitrocellulose, a mineral, plastic, polyacrylamide, latex, or ceramic.

[0232] In some embodiments, the peptide comprises a hormone, neurotransmitter, enzyme, antibody, viral protein, bacterial protein, synthetic peptide, bioactive peptide, peptide hormone, oligopeptide, polypeptide, fusion protein, cyclic peptide, branched peptide, recombinant protein, tumor marker, therapeutic peptide, antigenic peptide, or signaling peptide. In some embodiments, the peptide is derived from a cell lysate, blood sample, plasma sample, serum sample, tissue biopsy, saliva sample, urine sample, cerebrospinal fluid sample, effusion sample, synovial fluid sample, fecal sample, gut microbiome sample, environmental water sample, soil sample, bacterial culture, viral culture, organoid, tumor biopsy, sputum sample, or hair sample. In some embodiments, the peptide is associated with a disease.

[0233] In some embodiments, said transferring information comprises performing nucleic acid amplification, extension, enzymatic ligation, splint ligation, chemical ligation, template-assisted ligation, use of a ligase enzyme, use of a splint oligonucleotide, use of a catalyst, use of a bridging molecule, use of a condensation agent, use of a coupling reagent, use of a polymerase enzyme, use of a complementary nucleic acid sequence, use of a nicking enzyme, use of a nucleic acid modifying enzyme, use of a recombinase, use of a strand-displacing polymerase, use of a single-strand binding protein, a click chemistry reaction, a phosphodiester bond formation, or a peptide nucleic acid-mediated ligation. In some embodiments, said transferring information comprises joining the recode nucleic acid or a reverse complement of the recode nucleic acid with the cycle nucleic acid.

[0234] In some embodiments, the information of the recode nucleic acid comprises a sequence of the recode nucleic acid or a reverse complement of the sequence of the recode nucleic acid.

[0235] In some embodiments, obtaining the sequence information for the recode block comprises performing sequencingProcess 200 - Plurality, Identity & Positional Information

[0236] Process 200 can relate to identifying amino acid identity and positional information of multiple aminio acids. Any aspect or combination of aspects from Process 200 may be included in a method herein.

[0237] Disclosed herein, in some embodiments, is a method for determining identity and positional information of a plurality of amino acid residues of a peptide, the peptide comprising n amino acid residues, the method comprising:Attorney Docket No: 062954-509001 WO(a) coupling the peptide to a solid support such that a N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions;(b) providing a chemically -reactive conjugate, the chemically-reactive conjugate comprising: (x) a location nucleic acid, (y) a reactive moiety for binding and cleaving the N-terminal amino acid residue of the peptide and exposing a next amino acid residue as a N-terminal amino acid residue on the cleaved peptide, and (z) an immobilizing moiety for immobilization to the solid support;(c) contacting the peptide with the chemically-reactive conjugate, thereby coupling the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex;(d) immobilizing the conjugate complex to the solid support via the immobilizing moiety;(e) cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby exposing the next amino acid residue as a N-terminal amino acid residue on the cleaved peptide and providing an immobilized location complex, the immobilized location complex comprising the cleaved and separated N-terminal amino acid residue and the location nucleic acid;(f) providing a second chemically-reactive conjugate, the chemically-reactive conjugate comprising: (xx) a cycle tag comprising a cycle nucleic acid associated with a cycle number, and (yy) a reactive moiety for binding and cleaving the N-terminal amino acid residue of the peptide and exposing a next amino acid residue as a N-terminal amino acid residue on the cleaved peptide, and optionally (zz) a moiety for joining to an optional second solid support.(g) contacting the peptide with the chemically-reactive conjugate, thereby coupling the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a coupled-conjugate- complex;(h) forming one or more transfer complexes, each transfer complex comprising an immobilized location complex and a coupled-conjugate-complex, thereby bringing a location nucleic acid into proximity with a cycle tag nucleic acid within each formed transfer complex;(i) within each formed transfer complex, joining a location nucleic acid or a reverse complement thereof to a cycle tag nucleic acid to form an immobilized recode unit, or otherwise joining information of the location nucleic acid and the cycle tag nucleic acid, thereby creating one or a plurality of immobilized recode units, each recode unit corresponding with a formed transfer complex;(j) cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby exposing the next amino acid residue as a N-terminal amino acid residue on the cleaved peptide and liberating the immobilized recode unit complex from the solid support, the liberated recode unit complex comprising the cleaved and separated N-terminal amino acid residue, cycle and location information.Attorney Docket No: 062954-509001 WO(k) repeating (f) through (j) n-1 times to liberate and collect pools having a plurality of recode unit complexes, each additional plurality of recode unit complexes comprising information associated with cycle 2 to n, accordingly;(l) contacting the collection of recode units complexes with binding agents, the binding agents comprising: a binding moiety for preferentially binding to one, or to a subset, of the recode unit complexes, and a recode tag comprising a recode nucleic acid corresponding with the binding agent, thereby forming one or more affinity complexes, each affinity complex comprising an recode unit complex and the binding agent, thereby bringing a recode unit nucleic acid into proximity with a recode tag within each formed affinity complex;(m) within each formed affinity complex, joining a recode unit or a reverse complement thereof to a recode tag to form a recode block, or otherwise transferring information of the recode tag to the recode unit complex, thereby creating a plurality of recode blocks, each recode block corresponding with a formed affinity complex;(n) optionally, joining two or more members of the plurality of recode blocks to form a memory oligonucleotide;(o) obtaining sequence information for the recode blocks or memory oligonucleotides; and(p) based on the obtained sequence information, determining identity and positional information of a plurality of amino acid residues of the peptide.

[0238] In some embodiments, a method for determining identity and positional information of a plurality of amino acid residues of a peptide is disclosed. In some embodiments, the peptide comprises n amino acid residues. In some embodiments, the method comprises coupling the peptide to a solid support such that a N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions.

[0239] In some embodiments, the method additionally comprises providing a chemically-reactive conjugate. In some embodiments, the chemically-reactive conjugate comprises: (x) a location nucleic acid, (y) a reactive moiety for binding and cleaving the N-terminal amino acid residue of the peptide and exposing a next amino acid residue as a N-terminal amino acid residue on the cleaved peptide, and (z) an immobilizing moiety for immobilization to the solid support.

[0240] In some embodiments, the method additionally comprises contacting the peptide with the chemically-reactive conjugate, thereby coupling the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex.

[0241] In some embodiments, the method additionally comprises immobilizing the conjugate complex to the solid support via the immobilizing moiety.Attorney Docket No: 062954-509001 WO

[0242] In some embodiments, the method additionally comprises cleaving and thereby separating the N- terminal amino acid residue from the peptide, thereby exposing the next amino acid residue as a N-terminal amino acid residue on the cleaved peptide and providing an immobilized location complex. In some embodiments, the immobilized location complex comprises the cleaved and separated N-terminal amino acid residue and the location nucleic acid.

[0243] In some embodiments, the method additionally comprises providing a second chemically -reactive conjugate. In some embodiments, the chemically -reactive conjugate comprises: (xx) a cycle tag comprising a cycle nucleic acid associated with a cycle number, and (yy) a reactive moiety for binding and cleaving the N-terminal amino acid residue of the peptide and exposing a next amino acid residue as a N-terminal amino acid residue on the cleaved peptide. In some embodiments, the chemically-reactive conjugate additionally comprises (zz) a moiety for joining to an optional second solid support. In some embodiments, the chemically-reactive conjugate optionally additionally comprises (zz) a moiety for joining to an optional second solid support.

[0244] In some embodiments, the method additionally comprises contacting the peptide with the chemically-reactive conjugate. In some embodiments contacting the peptide with the chemically-reactive conjugate couples the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a coupled-conjugate-complex.

[0245] In some embodiments, the method additionally comprises forming one or more transfer complexes. In some embodiments, each transfer complex comprises an immobilized location complex and a coupled- conjugate-complex, thereby bringing a location nucleic acid into proximity with a cycle tag nucleic acid within each formed transfer complex.

[0246] In some embodiments, the method additionally comprises, within each formed transfer complex, joining a location nucleic acid or a reverse complement thereof to a cycle tag nucleic acid to form an immobilized recode unit, or otherwise joining information of the location nucleic acid and the cycle tag nucleic acid, thereby creating one or a plurality of immobilized recode units. In some embodiments each recode unit corresponds with a formed transfer complex.

[0247] In some embodiments, the method additionally comprises cleaving and thereby separating the N- terminal amino acid residue from the peptide, thereby exposing the next amino acid residue as a N-terminal amino acid residue on the cleaved peptide and liberating the immobilized recode unit complex from the solid support. In some embodiments, the liberated recode unit complex comprises the cleaved and separated N-terminal amino acid residue, cycle and location information.

[0248] In some embodiments, the method additionally comprises repeating one or more steps n-1 times to liberate and collect pools having a plurality of recode unit complexes, each additional plurality of recodeAttorney Docket No: 062954-509001 WO unit complexes comprising information associated with cycle 2 to n, accordingly. In some embodiments, the repeated steps are (f) through (j).

[0249] In some embodiments, the method additionally comprises contacting the collection of recode units complexes with binding agents. In some embodiments, the binding agents comprise: a binding moiety for preferentially binding to one, or to a subset, of the recode unit complexes, and a recode tag comprising a recode nucleic acid corresponding with the binding agent, thereby forming one or more affinity complexes, each affinity complex comprising an recode unit complex and the binding agent, thereby bringing a recode unit nucleic acid into proximity with a recode tag within each formed affinity complex.

[0250] In some embodiments, the method additionally comprises within each formed affinity complex, joining a recode unit or a reverse complement thereof to a recode tag to form a recode block, or otherwise transferring information of the recode tag to the recode unit complex, thereby creating a plurality of recode blocks, each recode block corresponding with a formed affinity complex.

[0251] In some embodiments, the method additionally comprises joining two or more members of the plurality of recode blocks to form a memory oligonucleotide.

[0252] In some embodiments, the method additionally comprises obtaining sequence information for the recode blocks or memory oligonucleotides.

[0253] In some embodiments, the method additionally comprises, based on the obtained sequence information, determining identity and positional information of a plurality of amino acid residues of the peptide.

[0254] In some embodiments, steps (f)-(k) are repeated 2, 3, 4, or more times.

[0255] In some embodiments, n is an integer greater than or equal to 2.

[0256] In some embodiments, each binding agent comprises recode tags with a unique nucleic acid sequence. In some embodiments, a plurality of binding agents comprises recode tags with the same nucleic acid sequence.

[0257] In some embodiments, the binding agents are contacted with combined pool(s) of recode unit complexes.

[0258] In some embodiments, the method comprises washing the immobilized amino acid complex before said contacting the immobilized amino acid complex with a coupled-conjugate-complex. In some embodiments, the method comprises washing the coupled-conjugate-complex before said contacting the immobilized amino acid complex with a coupled-conjugate-complex.

[0259] In some embodiments, a plurality of peptides are immobilized to the solid support and analyzed in concert.

[0260] In some embodiments, obtaining the sequence information for the memory oligonucleotide comprises performing sequencing.Attorney Docket No: 062954-509001 WO

[0261] In some embodiments, a method comprises protecting or deprotecting the location tag at any appropriate step of the workflow between (b) and (k).

[0262] In some embodiments, said transferring information comprises performing nucleic acid amplification, extension, enzymatic ligation, splint ligation, chemical ligation, template-assisted ligation, use of a ligase enzyme, use of a splint oligonucleotide, use of a catalyst, use of a bridging molecule, use of a condensation agent, use of a coupling reagent, use of a polymerase enzyme, use of a complementary nucleic acid sequence, use of a nicking enzyme, use of a nucleic acid modifying enzyme, use of a recombinase, use of a strand-displacing polymerase, use of a single-strand binding protein, a click chemistry reaction, a phosphodiester bond formation, or a peptide nucleic acid-mediated ligation.

[0263] In some embodiments, the information of the recode nucleic acid comprises a sequence of the recode nucleic acid or a reverse complement of the sequence of the recode nucleic acid.

[0264] In some embodiments, said transferring information comprises joining the recode nucleic acid or a reverse complement of the recode nucleic acid with the cycle nucleic acid.

[0265] In some embodiments, determining the identity and positional information of the plurality of amino acid residues of the peptide comprises determining the identity and positional information of all of the amino acid residues of the peptide. In some embodiments, determining the identity and positional information of the plurality of amino acid residues of the peptide comprises determining the identity and positional information of only a subset of the amino acid residues of the peptide.

[0266] In some embodiments, a method comprises identifying the peptide by comparing the identity and positional information of the plurality of amino acid residues to a database.

[0267] In some embodiments, following step (f) metadata may be immobilized proximal to a macromolecule, and the metadata information may be transferred between the metadata nucleic acid and the location nucleic acid of the immobilized location complex. In some embodiments, metadata may be immobilized proximal to a macromolecule prior to step (f), and the metadata information may be transferred between the metadata nucleic acid and the location nucleic acid of the immobilized location complex.

[0268] In some embodiments washes between steps may improve the specificity and accuracy of the assays. For example, washing away unbound chemically-reactive conjugates following step (f) may improve the efficiency of associating a bound chemically-reactive conjugate with the immobilized location complex. Similarly, temperature modulation and other stimuli may be employed to improve specificity and accuracy of the assay.

[0269] In some embodiments, the information of the cycle tag is joined to the information of the location tag via the actions of polymerase, ligase, or chemically to form the recode unit.Attorney Docket No: 062954-509001 WO

[0270] In some embodiments, a collection of recode unit complexes across multiple cycles, e.g., repetitions of steps (f)-(k) , are pooled prior to contacting with binding agents. In some embodiments the collection of recode unit complexes are not pooled across cycles prior to contacting with binding agents.

[0271] In some embodiments, cleaving the N-terminal amino acid residue from the peptide exposes a next amino acid residue as a N-terminal amino acid residue on the cleaved peptide.

[0272] In some embodiments, the reactive moiety of the chemically-reactive conjugate cleaves the N- terminal amino acid residue from the peptide.

[0273] Some embodiments include repeating steps (b) through (k) for each subsequent amino acid of the peptide. Some embodiments include washing the immobilized amino acid complex before said contacting the immobilized amino acid complex with a binding agent.

[0274] In some embodiments, the binding moiety comprises a peptide, antibody, antibody fragment, or antibody derivative.

[0275] In some embodiments, the binding moiety comprises an aptamer. In some embodiments, the binding moiety binds to a natural amino acid, a post-translationally modified amino acid, a derivatized version of an amino acid, a derivatized or stabilized version of a post-translationally modified amino acid, a synthetic amino acid, an amino acid with a specific side chain, an amino acid with a phosphorylated side chain, an amino acid with a glycosylated side chain, an amino acid with a methylation modification, or a D-amino acid, or binds to a combination thereof.

[0276] In some embodiments, the solid support comprises a bead, a plate, or a chip.

[0277] In some embodiments, the solid support comprises glass slide, silica, a resin, a gel, a membrane, polystyrene, a metal, nitrocellulose, a mineral, plastic, polyacrylamide, latex, or ceramic.

[0278] In some embodiments, the peptide comprises a hormone, neurotransmitter, enzyme, antibody, viral protein, bacterial protein, synthetic peptide, bioactive peptide, peptide hormone, oligopeptide, polypeptide, fusion protein, cyclic peptide, branched peptide, recombinant protein, tumor marker, therapeutic peptide, antigenic peptide, or signaling peptide.

[0279] In some embodiments, the peptide is derived from a cell lysate, blood sample, plasma sample, serum sample, tissue biopsy, saliva sample, urine sample, cerebrospinal fluid sample, effusion sample, synovial fluid sample, fecal sample, gut microbiome sample, environmental water sample, soil sample, bacterial culture, viral culture, organoid, tumor biopsy, sputum sample, or hair sample. In some embodiments, the peptide is associated with a disease.

[0280] In some embodiments, said transferring information comprises performing nucleic acid amplification, extension, enzymatic ligation, splint ligation, chemical ligation, template-assisted ligation, use of a ligase enzyme, use of a splint oligonucleotide, use of a catalyst, use of a bridging molecule, use of a condensation agent, use of a coupling reagent, use of a polymerase enzyme, use of a complementaryAttorney Docket No: 062954-509001 WO nucleic acid sequence, use of a nicking enzyme, use of a nucleic acid modifying enzyme, use of a recombinase, use of a strand-displacing polymerase, use of a single-strand binding protein, a click chemistry reaction, a phosphodiester bond formation, or a peptide nucleic acid-mediated ligation. In some embodiments, the information of the recode nucleic acid comprises a sequence of the recode nucleic acid or a reverse complement of the sequence of the recode nucleic acid. In some embodiments, said transferring information comprises joining the recode nucleic acid or a reverse complement of the recode nucleic acid with the cycle nucleic acid. In some embodiments, the recode nucleic acid comprises a hybridization- capable code.

[0281] In some embodiments, the reactive moiety binds covalently to the immobilized amino acid complex (e.g. to the amino acid of the immobilized amino acid complex).

[0282] In some embodiments, the recode nucleic acid comprises a hybridization-capable code.

[0283] Some embodiments include obtaining sequence information for a memory oligonucleotide. In some embodiments, obtaining the sequence information for the memory oligonucleotide comprises performing sequencing.

[0284] In some embodiments, determining the identity and positional information of the plurality of amino acid residues of the peptide comprises the use of serial hybridization events using one or more fluorophores. In some embodiments, the fluorophores are conjugated to oligonucleotides.

[0285] In some embodiments, steps (b)-(e) are repeated one or more times to immobilize additional immobilized amino acid complexes comprising the cleaved and separated N-terminal amino acid residue and the location nucleic acid.

[0286] In some embodiments, randomly located location tags are joined to a solid support in leiu of steps (b)-(e), and prior to steps (f)-(k).

[0287] In some embodiments, the solid surface is a polymer. In some embodiments the solid surface may be cyclically adsorbed to, desorbed from a bead, slide, or any other solid surface to facilite efficienct solution phase reactions alternating with efficienct heterogenous wash steps. In some embodiments, the solid surface is a polymer which is cyclically precipitated, colloidally-aggregated, or solubilized by applying various solvents or combinations of solvents.

[0288] In some embodiments, metadata information may be joined with location information to provide a recode block using any sutiable methods, for example, amplification, enzymatic ligation, splint ligation, chemical ligation, template-assisted ligation, use of a ligase enzyme, use of a splint oligonucleotide, use of a catalyst, use of a bridging molecule, use of a condensation agent, use of a coupling reagent, use of a polymerase enzyme, use of a complementary nucleic acid sequence, use of a nicking enzyme, use of a nucleic acid modifying enzyme, use of a recombinase, use of a strand-displacing polymerase, use of a single-strand binding protein, a click chemistry reaction, a phosphodiester bond formation, or a peptideAttorney Docket No: 062954-509001 WO nucleic acid-mediated ligation. In some embodiments, the information of the metadata nucleic acid comprises a sequence of the metadata nucleic acid or a reverse complement of the sequence of the metadata nucleic acid. In some embodiments, said transferring information comprises joining the metadata nucleic acid or a reverse complement of the recode nucleic acid with the location nucleic acid. In some embodiments, the metadata nucleic acid comprises a hybridization-capable code.Process 300 - Affinity Recognition-Free Sequencing

[0289] Any aspect or combination of aspects from Process 300 may be included in a method herein: disclosed herein, in some embodiments, is a method for determining identity and positional information of a plurality of amino acid residues of a peptide, the peptide comprising n amino acid residues, the method comprising:(a) coupling the peptide to a solid support such that a N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions;(b) providing a chemically -reactive conjugate, the chemically-reactive conjugate comprising: (x) a location nucleic acid, (y) a reactive moiety for binding and cleaving the N-terminal amino acid residue of the peptide and exposing a next amino acid residue as a N-terminal amino acid residue on the cleaved peptide, and (z) an immobilizing moiety for immobilization to the solid support;(c) contacting the peptide with the chemically-reactive conjugate, thereby coupling the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex;(d) immobilizing the conjugate complex to the solid support via the immobilizing moiety;(e) cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby exposing the next amino acid residue as a N-terminal amino acid residue on the cleaved peptide and providing an immobilized location complex, the immobilized location complex comprising the cleaved and separated N-terminal amino acid residue and the location nucleic acid;(f) providing a second chemically-reactive conjugate, the chemically-reactive conjugate comprising: (xx) a cycle tag, and (yy) a reactive moiety for binding and cleaving the N-terminal amino acid residue of the peptide and exposing a next amino acid residue as a N-terminal amino acid residue on the cleaved peptide, and (zz) a nucleic acid useful for nucleic acid joining operations, or an immobilizing moiety.(g) contacting the peptide with the chemically-reactive conjugate, thereby coupling the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a coupled-conjugate- complex;Attorney Docket No: 062954-509001 WO(h) forming one or more transfer complexes, each transfer complex comprising an immobilized location complex and a coupled-conjugate-complex, thereby bringing a location nucleic acid into proximity with a cycle tag nucleic acid within each formed transfer complex;(i) within each formed transfer complex, joining a location nucleic acid or a reverse complement thereof to a cycle tag nucleic acid to form an immobilized recode unit, or otherwise joining information of the location nucleic acid and the cycle tag nucleic acid, thereby creating one or a plurality of immobilized recode units, each recode unit corresponding with a formed transfer complex;(j) cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby exposing the next amino acid residue as a N-terminal amino acid residue on the cleaved peptide and liberating the immobilized recode unit complex from the solid support, the liberated recode unit complex comprising the cleaved and separated N-terminal amino acid residue, cycle and location information.(k) repeating (f) through (j) n-1 times to liberate and collect pools having a plurality of recode unit complexes, each additional plurality of recode unit complexes comprising information associated with cycle 2 to n, accordingly;(m) optionally, attaching a nucleic acid comprising nucleic acid sequence useful in nucleic acid joining operations, to the recode unit via the immobilizing moiety of the CRC(n) optionally, joining two or more members of the plurality of recode blocks to form a memory oligonucleotide;(o) obtaining sequence information for the recode units or memory oligonucleotides; and(p) based on the obtained sequence information, determining identity and positional information of a plurality of amino acid residues of a plurality of peptides.

[0290] In some embodiments, a method for determining identity and positional information of a plurality of amino acid residues of a peptide is disclosed herein. In some embodiments, the peptide comprises n amino acid residues. In some embodiments, the method comprises (a) coupling the peptide to a solid support such that a N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions.

[0291] In some embodiments, the method additionally comprises (b) providing a chemically-reactive conjugate, the chemically-reactive conjugate comprising: (x) a location nucleic acid, (y) a reactive moiety for binding and cleaving the N-terminal amino acid residue of the peptide and exposing a next amino acid residue as a N-terminal amino acid residue on the cleaved peptide, and (z) an immobilizing moiety for immobilization to the solid support.

[0292] In some embodiments, the method additionally comprises (c) contacting the peptide with the chemically-reactive conjugate, thereby coupling the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex.Attorney Docket No: 062954-509001 WO

[0293] In some embodiments, the method additionally comprises (d) immobilizing the conjugate complex to the solid support via the immobilizing moiety.

[0294] In some embodiments, the method additionally comprises (e) cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby exposing the next amino acid residue as a N- terminal amino acid residue on the cleaved peptide and providing an immobilized location complex, the immobilized location complex comprising the cleaved and separated N-terminal amino acid residue and the location nucleic acid.

[0295] In some embodiments, the method additionally comprises (f) providing a second chemicallyreactive conjugate, the chemically-reactive conjugate comprising: (xx) a cycle tag, and (yy) a reactive moiety for binding and cleaving the N-terminal amino acid residue of the peptide and exposing a next amino acid residue as a N-terminal amino acid residue on the cleaved peptide, and (zz) a nucleic acid useful for nucleic acid joining operations, or an immobilizing moiety.

[0296] In some embodiments, the method additionally comprises (g) contacting the peptide with the chemically-reactive conjugate, thereby coupling the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a coupled-conjugate-complex.

[0297] In some embodiments, the method additionally comprises (h) forming one or more transfer complexes, each transfer complex comprising an immobilized location complex and a coupled-conjugate- complex, thereby bringing a location nucleic acid into proximity with a cycle tag nucleic acid within each formed transfer complex.

[0298] In some embodiments, the method additionally comprises (i) within each formed transfer complex, joining a location nucleic acid or a reverse complement thereof to a cycle tag nucleic acid to form an immobilized recode unit, or otherwise joining information of the location nucleic acid and the cycle tag nucleic acid, thereby creating one or a plurality of immobilized recode units, each recode unit corresponding with a formed transfer complex.

[0299] In some embodiments, the method additionally comprises (j) cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby exposing the next amino acid residue as a N- terminal amino acid residue on the cleaved peptide and liberating the immobilized recode unit complex from the solid support, the liberated recode unit complex comprising the cleaved and separated N-terminal amino acid residue, cycle and location information.

[0300] In some embodiments, the method additionally comprises (k) repeating (f) through (j) n-1 times to liberate and collect pools having a plurality of recode unit complexes, each additional plurality of recode unit complexes comprising information associated with cycle 2 to n, accordingly.Attorney Docket No: 062954-509001 WO

[0301] In some embodiments, the method additionally comprises (m) optionally, attaching a nucleic acid comprising nucleic acid sequence useful in nucleic acid joining operations, to the recode unit via the immobilizing moiety of the CRC.

[0302] In some embodiments, the method additionally comprises (n) optionally, joining two or more members of the plurality of recode blocks to form a memory oligonucleotide.

[0303] In some embodiments, the method additionally comprises (o) obtaining sequence information for the recode units or memory oligonucleotides.

[0304] In some embodiments, the method additionally comprises (p) based on the obtained sequence information, determining identity and positional information of a plurality of amino acid residues of a plurality of peptides.

[0305] In some embodiments, the immobilization moiety of the CRC of step (f) may associate a nucleic acid useful for nucleic acid joining operations.

[0306] In some embodiments, a nucleic acid useful for nucleic acid joining operations is associated before step (j) or after step (j), or between any appropriate steps of Proess 300 to provide a recode unit comprising nuclic acid useful for nucleic acid joining operations.

[0307] In some embodiments, associating nucleic acid useful for nucleic acid joining operations may occur through a binding interaction, such as biotin to streptavidin, by formation of a chemical bond, by establishing non-covalent interactions, such as hydrogen bonds, pi-stacking interactions, hydrophobic interactions, vanderWaals interactions, entanglement interactions, and / or electrostatic interactions.

[0308] In some embodiments, obtaining sequence information comprises nanopore sequencing, whereby the identity of the amino acid is ascertained based on the unique current signature of an amino acid within the constant context of the CRC component of the recode unit complex as it passes through the pore, while the location and cycle information of the associated amino acid is ascertained as the DNA bases of the recode unit complex pass through the pore.

[0309] In some embodiments, the constant context of the CRC component includes flanking known base pairs of nucleic acid, such that the size of the constant region exceeds the dimensions of the nanopore

[0310] In some embodiments, the constant context of the CRC component includes flanking known base pairs of nucleic acid, such that the size of the constant region exceeds the dimensions of the nanopore by a certain amount of base pairs on either side, including 1,2, 3, 4 or more base pairs

[0311] In some embodiments, the CRC components are aggregated into multiple units prior to nanopore sequencing, comprising N units (N from 1 to 100, 1000, or more depending on the nanopore technology and read length desired)

[0312] In some embodiments, obtaining sequence information comprises nanopore sequencing, whereby sequence information of concatenated recode unit complexes (memory oligos) is obtained.Attorney Docket No: 062954-509001 WO

[0313] In some embodiments, steps (f)-(k) are repeated 2, 3, 4, or more times.

[0314] In some embodiments, n is an integer greater than or equal to 2.

[0315] In some embodiments after step (j), a liberated recode unit complex is captured on a solid support by interaction of the immobilizing moiety of the recode unit complex with a complementary feature of the solid support. In some embodiments, after step (j) a liberated recode unit complex is captured on a solid support by covalent, non-covalent, and / or non-specific interactions with a solid support. In some embodiments, after step (j) liberated recode unit complexes are captured on a solid support. In some embodiments, the solid support is a magnetic bead, a bead coated with streptavidin, a SPRI bead, a bead having a functional group or groups, or any combination of solid surface modifications thereof.

[0316] In some embodiments, modified NTPs incorporated into the recode unit oligo at step (i) provide functional groups useful for purification.

[0317] In some aspects, washing operations purify and concentrate recode unit complexes prior to joining recode unit complexes and / or obtaining sequence information of the recode units.

[0318] In some embodiments, a plurality of peptides are immobilized to the solid support and analyzed in concert.

[0319] In some embodiments, a method comprises protecting or deprotecting the location tag at any appropriate step of the workflow between (b) and (k).

[0320] In some embodiments, said transferring information comprises performing nucleic acid amplification, extension, enzymatic ligation, splint ligation, chemical ligation, template-assisted ligation, use of a ligase enzyme, use of a splint oligonucleotide, use of a catalyst, use of a bridging molecule, use of a condensation agent, use of a coupling reagent, use of a polymerase enzyme, use of a complementary nucleic acid sequence, use of a nicking enzyme, use of a nucleic acid modifying enzyme, use of a recombinase, use of a strand-displacing polymerase, use of a single-strand binding protein, a click chemistry reaction, a phosphodiester bond formation, or a peptide nucleic acid-mediated ligation.

[0321] In some embodiments, the information of the recode nucleic acid comprises a sequence of the recode nucleic acid or a reverse complement of the sequence of the recode nucleic acid.

[0322] In some embodiments, said transferring information comprises joining the recode nucleic acid or a reverse complement of the recode nucleic acid with the cycle nucleic acid.

[0323] In some embodiments, the solid support comprises glass slide, silica, a resin, a gel, a membrane, polystyrene, a metal, nitrocellulose, a mineral, plastic, polyacrylamide, latex, or ceramic.

[0324] In some embodiments, said transferring information comprises performing nucleic acid amplification, extension, enzymatic ligation, splint ligation, chemical ligation, template-assisted ligation, use of a ligase enzyme, use of a splint oligonucleotide, use of a catalyst, use of a bridging molecule, use of a condensation agent, use of a coupling reagent, use of a polymerase enzyme, use of a complementaryAttorney Docket No: 062954-509001 WO nucleic acid sequence, use of a nicking enzyme, use of a nucleic acid modifying enzyme, use of a recombinase, use of a strand-displacing polymerase, use of a single-strand binding protein, a click chemistry reaction, a phosphodiester bond formation, or a peptide nucleic acid-mediated ligation. In some embodiments, the information of the recode nucleic acid comprises a sequence of the recode nucleic acid or a reverse complement of the sequence of the recode nucleic acid. In some embodiments, said transferring information comprises joining the recode nucleic acid or a reverse complement of the recode nucleic acid with the cycle nucleic acid. In some embodiments, the recode nucleic acid comprises a hybridization- capable code.

[0325] In some embodiments the 3’ end of the recode block is capped to prevent 3’ extension by a polymerase, or enzyme. In some embodiments, said cap comprises a dideoxy nucleotide. In some embodiments the dideoxy nucleotide is provided by action of a terminal transferase following step (1) of Process 100 or step (m) of Process 200. In some embodiments, said capping comprises a chemical ligation. In some embodiments the dideoxy nucleotides are incorporated during step (1) of Process 100 or during step (m) of Process 200. In some embodiments, modified nucleic acids, such as LNA, are applied during step (1) of Process 100 or during step (m) of Process 200. In some embodiments, combination of modified nucleotidetriphosphates are employed during step (1) of Process 100 or during step (m) of Process 200. In some embodiments, a combination of modified nucleotidetriphosphates are employed during step (i) of Process 300.

[0326] In some embodiments, nucleotide triphosphates having modified bases may be incorporated into a recode unit or recode block during polymerase extention steps of the Process 100, 200, or 300. In some embodiments, nucleotide modifications comprise, a biotin, a dinitrophenyl, a dioxane group, or other small molecule hapten, a peptide, an amine group, an oxyamine group, an aldehyde, a hydrazine, a tetrazine, an azide, an alkyne, an alkene, a trans -cyclooctene, a DBCO, a bicyclononyne, a norbornene, a strained alkyne, or a strained alkene, or a derivative thereof. In some embodiments, the modified nucleic acid comprises a protected functional group, for example a protected oxyamine group, a protected thiol, a protected amine, a protected hydrazine, or other, or a derivative thereof. In some embodiments, said modifications may be useful for purification, separation, or functionalization of nucleic acid sequences. Examples of commercially-available modified nucleotide triphosphates include: Trilink Biotechnologies Cat # N-5004, Millipore Sigma Cat# 1093070910, Biotium cat# 40035, Jena Biosciences cat#NU-976. In some embodiments, said nucleic acid functionalization may be used to couple a moiety that may be used to purify, separate, or further functionalize a nucleic acid, such as a His tag, flag tag,, biotin, a dinitrophenyl, a dioxane group, or other small molecule hapten, a peptide, an amine group, an oxyamine group, an aldehyde, a hydrazine, a tetrazine, an azide, an alkyne, an alkene, a trans-cyclooctene, a DBCO, a bicyclononyne, a norbornene, a strained alkyne, or a strained alkene, or a derivative thereof, etc. In some embodiments theAttorney Docket No: 062954-509001 WO recode block and / or recode unit complex contains a group useful for purification of the recode block and / or recode unit.

[0327] Alternative Embodiments Using Enzymatic Polymerization

[0328] In some embodiments, step (i) of Process 300 comprises enzymatic incorporation of amino acid information into a growing oligonucleotide through template-independent polymerization. In these embodiments, modified nucleotide triphosphates serve as both substrates for polymerase extension and carriers of amino acid residue information, providing an alternative to assembly-based approaches for joining cycle and location information.

[0329] In some embodiments, an amino acid polymer (protein, peptide) is linked to an oligonucleotide sequence such that the 3' end is unblocked and available for extension and the N-terminus of the amine is unblocked and available for reaction. The oligonucleotide may comprise a location tag or UMI and may comprise acid-resistant bases including 7-deazaadenine and 7-deazaguanine. The complex may be linked to a solid support to enable reagent exchange.

[0330] In some embodiments, a modified nucleotide triphosphate possessing a group capable of binding and cleaving the N-terminal amino acid is introduced. Such groups include isothiocyanate groups (e.g., phenylisothiocyanate, naphthylisothiocyanate), aldehyde groups, guanidinylating agents, or other aminereactive moieties. The modified nucleotide triphosphate reacts with the N-terminal amine group to form a conjugate complex.

[0331] In some embodiments, an enzyme capable of incorporating nucleotide triphosphates is introduced. The enzyme may be a terminal deoxynucleotidyl transferase (TdT) or a polymerase. The enzyme adds the modified nucleotide triphosphate to the 3' end of the oligonucleotide, thereby covalently incorporating the N-terminal amino acid-modified nucleotide into the growing oligonucleotide strand. In some embodiments, other nucleotide triphosphates are introduced at this step to create a segment of DNA containing the amino acid-linked modified nucleotide. These other nucleotides may be selected to confer acid stability, such as dCTP and dTTP, which are resistant to strong acids including trifluoroacetic acid.

[0332] After washing away the enzyme and unincorporated nucleotides, an acid is introduced to cleave the N-terminal amino acid from the protein or peptide. Suitable acids include trifluoroacetic acid, Lewis acids, or other cleavage reagents. The acid is then washed away and the cycle may be repeated with the modified nucleotide binding to the next amino acid in the sequence. Through iterative cycles, an oligonucleotide is generated containing amino acid residues attached at positions that sequentially correspond to their sequence positions in the original peptide or protein.

[0333] In some embodiments, a blocking group is used on the 3'OH of the modified nucleotide, such as an azidomethyl group, such that polymerization terminates after the modified nucleotide triphosphate is incorporated in each cycle. A deblocking step is performed after removing the enzyme and nucleotides inAttorney Docket No: 062954-509001 WO order to allow the next cycle to proceed. The other nucleotides may be introduced in any order or mixture and may be introduced in an alternating fashion (e.g., polyCs in one cycle, polyTs in the next cycle, polyCs in the next cycle) to enable improved detection by highlighting the junctions between homopolymer regions.

[0334] In some embodiments, translocation control elements are introduced in the mixture of other nucleotides or may coincide with the modified nucleotide to modulate nanopore translocation velocity and signal detection.

[0335] In alternative embodiments, the initial oligonucleotide is attached to the solid support in proximity to the protein or peptide, rather than directly linked to the protein or peptide.

[0336] In alternative embodiments, a DNA template is used to template the growth of the oligonucleotide. The template may be used to control length and composition of the sequences intervening between each amino acid-modified nucleotide. By controlling which free nucleotides are added (for example, adding only dTTP in one cycle), the length and composition of sequences intervening between each amino acid- modified nucleotide may be precisely controlled. In some embodiments, the template comprises a poly-T sequence, which provides superior resistance to acid degradation during the cleavage steps. The use of a poly-T template allows for robust oligonucleotide synthesis that can withstand repeated exposure to strong acids such as trifluoroacetic acid without significant degradation of the nucleic acid backbone.

[0337] In some embodiments, the oligonucleotides generated through these enzymatic polymerization approaches are released from the solid support and subjected to nanopore analysis. The oligonucleotide construct, comprising amino acid residues covalently incorporated at defined positions within the nucleic acid backbone, provides a readable input for nanopore sequencing that encodes both amino acid identity and positional information.

[0338] These polymerization-based embodiments differ from assembly-based approaches in that the oligonucleotide carrying amino acid information is synthesized de novo during the workflow through enzymatic incorporation of modified nucleotide substrates, rather than assembled from pre-synthesized oligonucleotide fragments. The enzymatic polymerization enables continuous strand synthesis with amino acid information encoded through the incorporation pattern of modified nucleotides and intervening standard nucleotides.Process 400 - Pooled Indexing of Recode Information

[0339] Any aspect or combination of aspects from Process 400 may be included in a method herein: disclosed herein, in some embodiments, is a method for determining identity and positional information of a plurality of amino acid residues of a peptide, the peptide comprising n amino acid residues, the method comprising:Attorney Docket No: 062954-509001 WO(a) coupling the peptide to a solid support such that a N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions;(b) providing a chemically -reactive conjugate, the chemically-reactive conjugate comprising: (x) a location nucleic acid, (y) a reactive moiety for binding and cleaving the N-terminal amino acid residue of the peptide and exposing a next amino acid residue as a N-terminal amino acid residue on the cleaved peptide, and (z) an immobilizing moiety for immobilization to the solid support;(c) contacting the peptide with the chemically-reactive conjugate, thereby coupling the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex;(d) immobilizing the conjugate complex to the solid support via the immobilizing moiety;(e) cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby exposing the next amino acid residue as a N-terminal amino acid residue on the cleaved peptide and providing an immobilized location complex, the immobilized location complex comprising the cleaved and separated N-terminal amino acid residue and the location nucleic acid;(f) providing a second chemically-reactive conjugate, the chemically-reactive conjugate comprising: (xx) a nucleic acid useful for nucleic acid joining operations, and (yy) a reactive moiety for binding and cleaving the N-terminal amino acid residue of the peptide and exposing a next amino acid residue as a N-terminal amino acid residue on the cleaved peptide, and optionally (zz) a moiety for joining to an optional second solid support.(g) contacting the peptide with the chemically-reactive conjugate, thereby coupling the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a coupled-conjugate- complex;(h) forming one or more transfer complexes, each transfer complex comprising an immobilized location complex and a coupled-conjugate-complex, thereby bringing a location nucleic acid into proximity with nucleic acid useful for nucleic acid joining operations;(i) within each formed transfer complex, joining a location nucleic acid or a reverse complement thereof to a coupled-conjugate-complex to form an immobilized recode unit complex, or otherwise joining information of the location nucleic acid and the coupled-conjugate-complex, thereby creating an immobilized recode unit complex, each recode unit complex corresponding with a formed transfer complex;(j) cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby exposing the next amino acid residue as a N-terminal amino acid residue on the cleaved peptide and liberating the immobilized recode unit complex from the solid support, the liberated recode unit complex comprising the cleaved and separated N-terminal amino acid and location information.Attorney Docket No: 062954-509001 WO(k) contacting a recode unit complex with an affinity recognition agent, the affinity recognition agent comprising: (a) a binding moiety for preferentially binding to one, or to a subset, of possible recode unit complexes, prescribed by the identity of the amino acid of the recode unit complex, and (b) a moiety useful for immobilizing the affinity recognition agent to a solid support, thereby forming an affinity complex comprising a recode unit complex and the affinity recognition agent.(l) Separating the affinity complex, and excess affinity recognition agents, from non-cognate recode unit complexes by immobilizing the affinity complex, and affinity recognition agents, to a solid support.(m) optionally, purifying the recode unit complex(n) joining a nucleic acid, or otherwise transferring information, representing the identity of the affinity recognition agent to the recode unit nucleic acid(o) repeating steps (k)-(n) one or more times to separate affinity complexes, and excess affinity recognition agents, from non-cognate recode unit complexes, thereby creating a plurality of recode units comprising location and amino acid information of an amino acid of a peptide.(p) optionally, combining the plurality of recode units(q) optionally, concatenating, or otherwise joining, the recode units(r) joining a nucleic acid, or otherwise transferring information, representing the cycle at which the recode unit complex was liberated from the immobilized peptide with the recode unit information to provide a memory oligo comprising cycle, amino acid, and location information for one or more amino acids of the immobilized peptide.(s) Repeating steps (f)-(r) one or more times, and combining the memory oligos from each repetition(t) obtaining sequence information for the memory oligonucleotides; and(u) based on the obtained sequence information, determining identity and positional information of a plurality of amino acid residues of a peptide.

[0340] In some embodiments, a method for determining identity and positional information of a plurality of amino acid residues of a peptide is disclosed. In some embodiments the peptide comprises n amino acid residues. In some embodiments, the method comprises (a) coupling the peptide to a solid support such that a N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions.

[0341] In some embodiments, the method additionally comprises (b) providing a chemically-reactive conjugate, the chemically-reactive conjugate comprising: (x) a location nucleic acid, (y) a reactive moiety for binding and cleaving the N-terminal amino acid residue of the peptide and exposing a next amino acidAttorney Docket No: 062954-509001 WO residue as a N-terminal amino acid residue on the cleaved peptide, and (z) an immobilizing moiety for immobilization to the solid support.

[0342] In some embodiments, the method additionally comprises (c) contacting the peptide with the chemically-reactive conjugate, thereby coupling the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex.

[0343] In some embodiments, the method additionally comprises (d) immobilizing the conjugate complex to the solid support via the immobilizing moiety.

[0344] In some embodiments, the method additionally comprises (e) cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby exposing the next amino acid residue as a N- terminal amino acid residue on the cleaved peptide and providing an immobilized location complex, the immobilized location complex comprising the cleaved and separated N-terminal amino acid residue and the location nucleic acid.

[0345] In some embodiments, the method additionally comprises (f) providing a second chemically- reactive conjugate, the chemically-reactive conjugate comprising: (xx) a nucleic acid useful for nucleic acid joining operations, and (yy) a reactive moiety for binding and cleaving the N-terminal amino acid residue of the peptide and exposing a next amino acid residue as a N-terminal amino acid residue on the cleaved peptide, and optionally (zz) a moiety for joining to an optional second solid support.

[0346] In some embodiments, the method additionally comprises (g) contacting the peptide with the chemically-reactive conjugate, thereby coupling the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a coupled-conjugate-complex.

[0347] In some embodiments, the method additionally comprises (h) forming one or more transfer complexes, each transfer complex comprising an immobilized location complex and a coupled-conjugate- complex, thereby bringing a location nucleic acid into proximity with nucleic acid useful for nucleic acid joining operations.

[0348] In some embodiments, the method additionally comprises (i) within each formed transfer complex, joining a location nucleic acid or a reverse complement thereof to a coupled-conjugate-complex to form an immobilized recode unit complex, or otherwise joining information of the location nucleic acid and the coupled-conjugate-complex, thereby creating an immobilized recode unit complex, each recode unit complex corresponding with a formed transfer complex.

[0349] In some embodiments, the method additionally comprises (j) cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby exposing the next amino acid residue as a N- terminal amino acid residue on the cleaved peptide and liberating the immobilized recode unit complex from the solid support, the liberated recode unit complex comprising the cleaved and separated N-terminal amino acid and location information.Attorney Docket No: 062954-509001 WO

[0350] In some embodiments, the method additionally comprises (k) contacting a recode unit complex with an affinity recognition agent, the affinity recognition agent comprising: (a) a binding moiety for preferentially binding to one, or to a subset, of possible recode unit complexes, prescribed by the identity of the amino acid of the recode unit complex, and (b) a moiety useful for immobilizing the affinity recognition agent to a solid support, thereby forming an affinity complex comprising a recode unit complex and the affinity recognition agent.

[0351] In some embodiments, the method additionally comprises (1) Separating the affinity complex, and excess affinity recognition agents, from non-cognate recode unit complexes by immobilizing the affinity complex, and affinity recognition agents, to a solid support.

[0352] In some embodiments, the method additionally comprises (m) optionally, purifying the recode unit complex.

[0353] In some embodiments, the method additionally comprises (n) joining a nucleic acid, or otherwise transferring information, representing the identity of the affinity recognition agent to the recode unit nucleic acid.

[0354] In some embodiments, the method additionally comprises (o) repeating steps (k)-(n) one or more times to separate affinity complexes, and excess affinity recognition agents, from non-cognate recode unit complexes, thereby creating a plurality of recode units comprising location and amino acid information of an amino acid of a peptide.

[0355] In some embodiments, the method additionally comprises (p) optionally, combining the plurality of recode units.

[0356] In some embodiments, the method additionally comprises (q) optionally, concatenating, or otherwise joining, the recode units.

[0357] In some embodiments, the method additionally comprises (r) joining a nucleic acid, or otherwise transferring information, representing the cycle at which the recode unit complex was liberated from the immobilized peptide with the recode unit information to provide a memory oligo comprising cycle, amino acid, and location information for one or more amino acids of the immobilized peptide.

[0358] In some embodiments, the method additionally comprises (s) Repeating steps (f)-(r) one or more times, and combining the memory oligos from each repetition.

[0359] In some embodiments, the method additionally comprises (t) obtaining sequence information for the memory oligonucleotides.

[0360] In some embodiments, the method additionally comprises (u) based on the obtained sequence information, determining identity and positional information of a plurality of amino acid residues of a peptide.

[0361] In some embodiments, n is an integer greater than or equal to 2.Attorney Docket No: 062954-509001 WO

[0362] In some embodiments, a plurality of peptides are immobilized to the solid support and analyzed in concert. In some embodiments, sequence information is obtained, determining identity and positional information of a plurality of amino acid residues of a plurality of peptides

[0363] In some embodiments, the CRC of Process 400, step (f) is a bifunctional CRC, as described herein.

[0364] In some embodiments, an affinity recognition reagent comprises a binding moiety, as defined herein.

[0365] In some embodiments, step (k) comprises contacting a recode unit complex with a plurality of affinity recognition agents, each comprising: (a) a binding moiety for preferentially binding to one, or to a subset, of possible recode unit complexes, prescribed by the identity of the amino acid of the recode unit complex, and (b) a moiety useful for immobilizing the affinity recognition agent to a solid support, thereby forming an affinity complex comprising a recode unit complex and the affinity recognition agent.

[0366] In some embodiments, a plurality of affinity recognition agents are combined to provide a pool of affinity recognition agents. In some embodiments, affinity recognition agent pools are serially contacted with recode units. In some embodiments, a heat, chemical, or other stimulus is applied to dissociate affinity complexes between applications of pools of affinity recognition agents. In some embodiments the affinity recognition agents are denatured between serial applications of pools of affinity recognition agents. In some embodiments, the composition of the plurality of affinitiy recognition agents provides error checking, redundancy, or efficiency benefits in the assignment of amino acid identity. In some embodiments, 5 serial applications of pools of affinity recognition agents can distinguish 32 different recode units.

[0367] In some embodiments, at step (n) amino acid identity is assigned due to the lack of formation of a recode unit with an affinity recognition agent. In some embodiments, at step (n) one or more amino acid identities are associated with a recode unit due to the lack of formation of an affinity complex.

[0368] In some embodiments, an affinity recognition reagent comprises one or more moieties that can be useful for immobilizing the affinity recognition reagent to a solid support. Immobilization through the immobilization moiety to a solid support may occur through a binding interaction, such as biotin to streptavidin, by formation of a chemical bond, by establishing non-covalent interactions, such as hydrogen bonds, pi-stacking interactions, hydrophobic interactions, vanderWaals interactions, entanglement interactions, and / or electrostatic interactions. The solid support may be any solid support useful within a method for determining protein information such as relative amino acid position within a peptide, identity, and / or relative location. Specific examples include: a biotin, a click chemistry moiety, an aliphatic molecule, a his-tag, snap-tag, spy-tag, oligonucleotide, linker, or any combination thereof.

[0369] In some embodimnets, purifying comprises, cleaving the nucleic acid useful for nucleic acid joining operations to separate the PTH-AA moiety of the recode unit complex from recode unit nucleic acid comprising location information using an endonuclease, immobilization of the recode unit nucleic acid viaAttorney Docket No: 062954-509001 WO the immobilization moiety of the recode unit complex and performing washing and / or buffer exchange, immobilization of the affinity complex via the immobilization moiety of the affinity recognition agent and performing washing and / or buffer exchange, serial dilution of non-target or non-cognate materials, or any combination thereof.

[0370] In some embodiments, obtaining sequence information comprises any short-read or long read NGS sequencing technology.

[0371] In some embodiments, steps (f)-(k) are repeated 2, 3, 4, or more times.

[0372] In some embodiments, step (j) of process 100, 200, 300, or 400 comprises a Lewis acid.

[0373] In some embodiments step (j) of process 100, 200, 300, or 400 comprises nucleic acids that comprise 7-Deazaadenosine and / or 7-Deaza-2'-deoxy-guanosine and / or FRNA, and / or other modified nucleo-bases. In some embodiments, step (i) of process 100, 200, 300, or 400 comprises 7-Deazaadenosine- 5'-O-triphosphate and / or 7-Deaza-2'-deoxy-guanosine-5'-triphosphate, or other modified nucleotide triphosphates.

[0374] In some embodiments, separating affinity complexes and excess affinity recognition agents comprises use of solid supports. In some embodiments, the solid support comprises glass slide, silica, a resin, a gel, a membrane, polystyrene, a metal, nitrocellulose, a mineral, plastic, polyacrylamide, latex, or ceramic. In some embodiments, the solid support is a magnetic bead, a bead coated with streptavidin, a SPRI bead, a bead having a functional group or groups, or any combination of solid surface modifications thereof.

[0375] In some embodiments after step (j), a liberated recode unit complex is captured on a solid support by interaction of the immobilizing moiety of the recode unit complex with a complementary feature of the solid support. In some embodiments, after step (j) a liberated recode unit complex is captured on a solid support by covalent, non-covalent, and / or non-specific interactions with a solid support. In some embodiments, after step (j) liberated recode unit complexes are captured on a solid support.

[0376] In some aspects, washing operations purify and concentrate recode unit complexes prior to joining recode unit complexes and / or obtaining sequence information of the recode units.

[0377] In some embodiments, steps of Process 100, 200, 300, and 400 are performed in any appropriate order to obtain sequence information, determining identity and positional information of a plurality of amino acid residues of a plurality of peptides

[0378] In some embodiments, a method comprises protecting or deprotecting the location tag at any appropriate step of the workflow between (b) and (k).

[0379] In some embodiments, said transferring information comprises performing nucleic acid amplification, extension, enzymatic ligation, splint ligation, chemical ligation, template-assisted ligation, use of a ligase enzyme, use of a splint oligonucleotide, use of a catalyst, use of a bridging molecule, use ofAttorney Docket No: 062954-509001 WO a condensation agent, use of a coupling reagent, use of a polymerase enzyme, use of a complementary nucleic acid sequence, use of a nicking enzyme, use of a nucleic acid modifying enzyme, use of a recombinase, use of a strand-displacing polymerase, use of a single-strand binding protein, a click chemistry reaction, a phosphodiester bond formation, or a peptide nucleic acid-mediated ligation.

[0380] In some embodiments the 3’ end of the recode unit is capped to prevent 3’ extension by a polymerase, or enzyme. In some embodiments, said cap comprises a dideoxy nucleotide. In some embodiments the dideoxy nucleotide is provided by action of a terminal transferase. In some embodiments, said capping comprises a chemical ligation. In some embodiments, a combination of modified nucleotidetriphosphates are employed during step of Process 400.

[0381] In some embodiments, the CRC of step(f) comprises a cycle tag, and information of the cycle tag and location oligo are joined, as described in Process 100, 200, and 300

[0382] In some embodiments, pooled recode unit complexes may be indexed. In some embodiments, one or more strings of concatentated recode units in a pool, or stirngs of joined recode units in a pool may be indexed. The index may include information related to cycle, amino acid identity, or assembly sequence.

[0383] It is recognized that steps of Process 100, 200, 300, and 400 may be combined in appropriate ways to accomplish the objective of creating a memory oligo from which polymeric macromolecules, including polymeric macromolecules such as peptides, polypeptides, and proteins may be analyzed. As one example, cycle information may be joined to location information at step (i) of Process 400, making indexing with cycle information at step (r) of Process 400 redundant.Process 500 - Creation of Local ULI Zones with Analyte Binding SitesSurface Preparation and Materials

[0384] Process 500 can relate to creating a local ULI zone with an analyte binding site, and can include use of development of: surface preparation, oligonucleotide attachment, a local ULI zone, an extension reaction, bridge amplification, a restriction enzyme, ULI propegation, a USER enzyme, secondary extension, analyte attachment, or multiple analyte binding sites. Any aspect or combination of aspects from Process 500 may be included in a method herein.

[0385] In some embodiments, a method is provided for creating spatially defined molecular zones on a surface. The surface may comprise a planar substrate, bead, nanoparticle, hydrogel, or other material that is chemically modified to support oligonucleotide attachment. Suitable chemical groups for attachment include carboxyl, amine, aldehyde, thiol, or azide groups, and surfaces may be functionalized using standard coupling chemistries such as EDC / NHS activation, click chemistry, or silanization. Surface functionalization chemistries may include carbodiimide coupling (EDC / NHS) for carboxyl-to-amine conjugation, thiol-maleimide coupling, copper-catalyzed or strain-promoted azide-alkyne cycloaddition,Attorney Docket No: 062954-509001 WO amino-silanization followed by succinimidyl ester or epoxy activation, biotin-streptavidin capture systems, oligonucleotide hybridization-based capture, photolithographic patterning with UV-activatable crosslinkers, plasma or corona discharge activation followed by coupling chemistry, or combinations thereof. In some embodiments, the surface may comprise or be associated with a hydrogel matrix. Hydrogels may be formed from synthetic polymers, natural polymers, or hybrid polymers, including but not limited to acrylamide, acrylate, vinyl pyridine, dihydroxy methacrylates, methacrylates, HEMA, PHEMA, PVA, HPMC, PLGA, PEG, or combinations thereof, in linear, branched, crosslinked, or block co-polymer configurations. A hydrogel may be associated with a solid support through covalent or non- covalent interactions and may comprise orthogonal conjugation chemistry modalities. Surface passivation strategies may employ polyethylene glycol (PEG), bovine serum albumin (BSA), casein, synthetic block copolymers, or other blocking agents to minimize non-specific binding. The surface density of immobilized oligonucleotides may be controlled through titration of reactive groups, blocking steps, or dilution with non-reactive spacer molecules.In certain embodiments, the surface comprises glass, silicon, polystyrene, polypropylene, polyethylene, polyacrylamide, agarose, or other biocompatible materials. The surface may be flat, curved, porous, or structured with microfeatures to enhance performance. In some embodiments, the surface comprises paramagnetic beads with diameters ranging from 0.1 to 100 micrometers, which facilitate separation and manipulation during processing steps.

[0386] In alternative embodiments, the surface may comprise hydrogel matrices composed of polyacrylamide, polyethylene glycol, or other hydrophilic polymers. Such hydrogels may provide three- dimensional scaffolds for ULI distribution while maintaining analyte accessibility. The hydrogel concentration, crosslinking density, and porosity may be adjusted to optimize ULI cluster formation and analyte binding.

[0387] This disclosure contemplates various surface densities of initial oligonucleotide attachment, ranging from approximately 10A6 to 10A12 molecules per square centimeter, with different densities being optimal for different applications.

[0388] This disclosure contemplates various 2 and 3 dimensional densities wherein the distance between initial oligonucleotides range from approximately 1-10000 nm, with distances being optimal for different applications.

[0389] In some embodiments, the template that provides a site for analyte placement is provided at a defined ratio to the template(s) that provide for ULI amplification. Is some aspects, the ratio is 1:1, 1:10, l;1000, 1:100000, or 1 : 10A8 or any ratio in between.

[0390] In some embodiments, the template(s) that provide a site for analyte placement are provided at a defined ratio to the template(s) that provide for ULI amplification, and are seeded (brought into proximity for hybridization or chemical or enzymatic attachment) onto a surface that is patterned. In someAttorney Docket No: 062954-509001 WO embodiments, the patterns comprise areas that support amplification of ULI territories and areas that do not support amplification of ULI territories.

[0391] In some embodiments, the areas that support ULI amplification and / or oligonucleotide immobilization via chemical or other means are configured to maximize the density of ULI territories. In some embodiments, the areas these areas are configured to optimize the density of ULI territories for optical imaging.

[0392] Some embodiments may include any or all of the following, related to ULI neighborhood creation: (1) surface functionalization with distinct oligonucleotide primers, one of which may include a cleavable terminus; (2) a selective extension reaction that introduces a sparse population of ULI- containing strands to define discrete analyte binding sites; (3) bridge amplification that copies ULI sequences to adjacent primers, which can create localized molecular territories with sharp boundaries defined by selective nucleotide exclusion; (4) enzymatic cleavage used to simplify molecular architecture surrounding each neighborhood; or (5) reconstitution of analyte binding sites with chemical modifications suitable for downstream conjugation, which may complete formation of spatially defined, barcode-dense molecular zones.Initial Oligonucleotide Attachment

[0393] In some embodiments, the surface is functionalized with two distinct oligonucleotide types (type 1 and type 2) that serve as anchoring primers for later extension. A cytosine-thymine-only (CT) alphabet can be advantageous, for example, when downstream steps require resistance to low-pH or acid-cleavage conditions, because purine-free backbones exhibit greater hydrolytic stability under acidic environments. Where such acid resistance is not needed, any mixed four-base sequence (A / C / G / T) may be substituted, provided the two strands remain sufficiently orthogonal.

[0394] To minimize cross-hybridization between anchors, CT1 and CT2 can be designed as partially homopolymeric strands. For example, each may be 30 - 60 nt long, with -15 - 25 consecutive C’s at the 5' end of CT1 and a matching run of T’s in CT2. This strategy is useful for shortening a portion that in some embodiments must be uniquely specified for orthogonality while preserving the C / T alphabet. The exact strand length, the size and placement of the homopolymer blocks, and the composition of the residual “unique zone” may be varied case-by-case; the sole constraint is that the unique segment be short and sequence-diverse enough to keep predicted self- or cross-hybridization energies above roughly - 5 kcal moL1under the intended reaction conditions. Alternative embodiments may distribute homopolymer tracts non-contiguously, use longer or shorter runs, or replace them with other low-complexity motifs, so long as the orthogonality goal is met.Attorney Docket No: 062954-509001 WO

[0395] Attachment to the surface may occur through the 5' terminus, a moiety internal to the oligonucleotide, or though the 3' terminus using standard chemistries such as carbodiimide / NHS activation, copper- or strain-promoted azide-alkyne cycloaddition, thiol-maleimide coupling, or photo-activatable crosslinkers for spatial patterning (e.g., masked illumination or direct-write photolithography), or other suitable conjugation chemistry. Blocking and washing steps may be included to reduce non-specific adsorption and remove excess reagents.

[0396] In some embodiments, a third orthogonal CT only oligo, CT3 may be used to serve as the initiator of the ULI neighborhood and the analyte binding site. The length of the variable CT portion may be adjusted to 20, 18, or less to maintain three orthogonal CT sections.Creation of local ULI zones with analyte bindins sitesExtension Reactions for Analyte Binding Site Creation

[0397] Following initial oligonucleotide attachment, a first extension or ligation reaction may be performed to create sparse analyte binding sites surrounded by amplifiable anchors. In some embodiments, this extension reaction utilizes three distinct oligonucleotide components with specific structures and functions.

[0398] The first component may comprise a sequence complementary to the type 1 oligonucleotide, extended with two consecutive iso-guanine nucleotides and a defined primer sequence (H5 or equivalent). This component may be present at a concentration of 0.5 to 10 pM in the extension reaction mixture.

[0399] The second component may comprise a defined primer sequence (H7 or equivalent), followed by two consecutive iso-cytosine nucleotides and a sequence complementary to the type 2 oligonucleotide. This component may also be present at a concentration of 0.5 to 10 pM.

[0400] The third component may comprise a defined primer sequence (H7 or equivalent), followed by two consecutive iso-cytosine nucleotides, a universal hybridization site, a unique location identifier, a second defined primer sequence (H5 or equivalent), a restriction site complement, and / or a sequence complementary to the type 2 oligonucleotide. This third component may be present at substantially lower concentration than the first two components, typically 0.1% to 1% of their concentration, or between 0.5 to 100 nM.

[0401] In some embodiments, ULIs incorporate sample indexing sequences to enable multiplexed workflows. Such sample-indexed ULIs facilitate pooling of analytes from different samples while maintaining sample identity throughout downstream processing steps. The sample index may comprise any suitable nucleotide sequence of sufficient length and composition to provide discrimination between different input samples. Sample indexing within ULIs may be particularly advantageous in high-throughput applications, enabling parallel processing of multiple samples with reduced resource requirements. WhenAttorney Docket No: 062954-509001 WO implementing sample-indexed ULIs, the index sequence may be positioned at a consistent location within the ULI architecture to facilitate computational identification and sample deconvolution during data analysis. The sample index and territory-specific portions of the ULI may be separated by defined spacer sequences or may be contiguous, depending on the specific implementation requirements.

[0402] These three exemplary oligonucleotides are depicted in FIG. 3, which illustrates their structures and roles in the initial extension reaction.

[0403] The extension reaction may be performed using a DNA polymerase with strand displacement activity, such as Bst, Phi29, or their engineered variants. The reaction conditions may include temperatures of 25 °C to 95°C, incubation times of 0.1 minute or longer, and standard buffer compositions appropriate for the selected polymerase.

[0404] In some embodiments, the unique location identifiers (ULIs) comprise random oligonucleotide sequences of 15 to 50 nucleotides in length, generated through combinatorial synthesis or enzymatic incorporation of random nucleotides. In other embodiments, the ULIs comprise structured sequences with defined segments for improved identification and error correction.

[0405] The universal hybridization site may comprise a sequence of 15 to 30 nucleotides or any length that provides efficient and specific hybridization to a universal primer. This site facilitates subsequent analysis steps by providing a common primer binding site across all ULIs.

[0406] The resulting distribution of some exemplary extended products on the surface is schematically shown in FIG. 4.Initial Bridge Amplification

[0407] Following the extension reaction, a single cycle of bridge amplification is performed to copy the ULI information to a proximal anchor. This bridge amplification utilizes the H5 and H7 primer sequences as anchors for hybridization and extension. A schematic of an exemplary initial bridge amplification process is provided in FIG. 5.

[0408] In some embodiments, the bridge amplification may be performed using a polymerase with stranddisplacement activity under isothermal conditions of 2550°C to 65 °C for 0.1 minute or longer. Alternative embodiments may utilize thermal cycling between 25 °C to 95°C to facilitate denaturation and annealing.

[0409] The extension during bridge amplification may incorporate standard deoxynucleotides (dATP, dTTP, dCTP, dGTP) or modified nucleotides such as deaza-dATP, deaza-dGTP, deaza-dCTP, deaza-dTTP, deaza-dUTP, iso-dCTP, iso-dGTP, 2'-deoxyinosine triphosphate, 5-methyl-dCTP, 5-hydroxymethyl-dCTP, dUTP, 7-deaza-7-substituted purine nucleotides, phosphorothioate-modified nucleotides, locked nucleic acid (LNA) nucleotides, or 2'-fluoro-modified nucleotides.. Some embodiments include a deaza nucleotide. Some embodiments include a deaza adenosine nucleotide. Some embodiments include a deaza cytidineAttorney Docket No: 062954-509001 WO nucleotide. Some embodiments include a deaza guanosine nucleotide. Some embodiments include a deaza thymidine nucleotide. Some embodiments include a deaza uridine nucleotide. Deaza-adenine and deazaguanine nucleotides may be used to generate amplicons with improved acid stability, facilitating compatibility with workflows that involve low pH conditions, such as single-molecule protein sequencing. The reaction mixture may include betaine, DMSO, or other additives to enhance polymerase performance and specificity.

[0410] This initial bridge amplification may create a single copy of the ULI at a location adjacent to the analyte binding site, establishing the foundation for subsequent ULI cluster formation.Restriction Enzyme Cleavage

[0411] Following initial bridge amplification, an endonuclease may be applied to temporarily remove the analyte binding site from participation in subsequent amplification steps. In some embodiments, the endonuclease comprises BbsLHF, BsmBI, Bsal, or other Type IIS restriction enzymes that cleave outside their recognition sequences. An exemplary cleavage step is illustrated in FIG. 6, which shows removal of the analyte binding strand via type IIS restriction enzyme.

[0412] The restriction enzyme reaction may be performed at temperatures of 25 °C to 65 °C for 1 minute or longer, with buffer compositions optimized for the specific enzyme. In some embodiments, bovine serum albumin may be included to stabilize the enzyme and reduce adsorption to surfaces.

[0413] The restriction site may be designed to leave specific overhangs that facilitate subsequent modification of the analyte binding site. For example, adenine-guanine-rich overhangs may be designed to facilitate specific extension with complementary cytosine-thymine-rich oligonucleotides in later steps.

[0414] In alternative embodiments, CRISPR-Cas systems may be employed instead of restriction enzymes, utilizing guide RNAs directed to sequences adjacent to the analyte binding site. Such approaches may offer increased specificity and programmability.Bridge Amplification for ULI Propagation

[0415] Following restriction enzyme cleavage, bridge amplification may be continued to propagate the ULI sequence to adjacent primers immobilized on the surface, forming a local molecular neighborhood or "territory" in which all replicated oligonucleotides carry the same ULI.

[0416] In some embodiments, this amplification is performed under controlled nucleotide conditions. The reaction mixture may comprise standard cytosine and thymine nucleotides, iso-guanine nucleotides, and deaza-adenine and deaza-guanine nucleotides to confer acid stability to the resulting amplicons, an important consideration for downstream workflows such as protein sequencing. Iso-cytosine may be deliberately excluded to ensure that extension halts at iso-guanine-containing motifs (e.g., iGiG), therebyAttorney Docket No: 062954-509001 WO chemically defining the outer limits of each amplification zone. In some cases, particularly under high magnesium or osmolyte conditions, certain polymerases may exhibit limited bypass of iGiG barriers. Careful selection of polymerase and reaction conditions may minimize this risk and maintain sharp territory boundaries. An exemplary process is illustrated in FIG. 7, which depicts continued bridge amplification between H5' and H7 anchors using the selective nucleotide mixture to generate discrete ULI territories.

[0417] During bridge amplification, ULI-labeled extension products may propagate outward from their origin, forming radially expanding wavefronts. In densely seeded systems, these wavefronts may encounter adjacent wavefronts, resulting in physical boundaries where the surface becomes saturated with extended products and no additional primer sites remain.

[0418] In implementations where analyte binding sites are introduced at lower density via Poisson-based loading, not all wavefronts may encounter neighbors before amplification is halted. Amplification in these cases may be terminated according to a predetermined reaction time, number of amplification cycles, or other user-defined parameters. However, due to the stochastic nature of surface seeding, some territories may still partially overlap, or wavefronts may contact one another by chance, resulting in imperfect spatial segregation. Such overlap may be tolerable for detection-focused applications that do not rely on strict spatial resolution and may still yield valid results in non-mapping implementations or mapping implementations which are tolerant of overlap.

[0419] The amplification process may be carried out using thermal cycling or isothermal incubation, and the extent of ULI propagation may be tuned by adjusting reaction time, polymerase concentration, or primer density, and / or number of thermal cycles.USER Enzyme Cleavage

[0420] Following bridge amplification, Uracil-Specific Excision Reagent (USER) enzyme may be applied to cleave at uracil sites, removing non-target strands. This cleavage step may simplify the molecular architecture and prepares the surface for subsequent modification of the analyte binding site.

[0421] In some embodiments, a complementary oligonucleotide may be annealed to the uracil-containing strand to ensure efficient USER enzyme cleavage. While uracil-DNA glycosylase can act on both single- and double-stranded DNA, cleavage efficiency is often enhanced in the context of a local duplex, particularly when paired with an AP-site-cleaving enzyme such as Endonuclease VIII. The choice of enzyme formulation and reaction conditions may influence a possible requirement for duplex formation at the cleavage site.

[0422] The USER enzyme reaction may be performed at temperatures of 25°C to 42°C for 1 or more minutes, using buffer compositions recommended by the enzyme manufacturer. Alternative embodimentsAttorney Docket No: 062954-509001 WO may utilize other uracil-selective cleavage approaches, such as UNG (uracil-N-glycosylase) followed by alkaline treatment or heat-induced cleavage at abasic sites.

[0423] Following USER enzyme treatment, each ULI territory may comprise:

[0424] A central analyte binding site with a specific terminal sequence derived from the restriction cleavage

[0425] Multiple copies of ULI-containing oligonucleotides surrounding the binding site

[0426] FIG. 8 illustrates an exemplary version of step 14D, in which the uracil-containing strand is removed after amplification.

[0427] An exemplary molecular configuration is shown in FIG. 9, which depicts a final composition of the ULI neighborhood following USER cleavage.

[0428] In some embodiments, a washing step is performed after U SER enzyme cleavage to remove cleaved fragments and other reaction components. This washing may utilize buffers, acids, and / or bases of varying stringency, potentially including detergents or chaotropic agents to ensure thorough removal of non- covalently bound material.Secondary Extension for Analyte Binding Site Modification

[0429] A secondary extension step may be performed to modify the analyte binding site for analyte attachment. In some embodiments, this extension may introduce an alkyne modification suitable for subsequent conjugation to peptides, proteins, nucleic acids, or other analytes through copper-catalyzed azide-alkyne cycloaddition (CuAAC) or strain-promoted azide-alkyne cycloaddition (SPAAC). An example final extension process is illustrated in FIG. 25, which shows the incorporation of alkyne-modified cytosines for downstream analyte conjugation.

[0430] The extension may utilize an oligonucleotide primer complementary to the terminal sequence of the analyte binding site. This primer may comprise a stretch of guanine nucleotides followed by a polyadenine sequence, creating multiple opportunities for incorporation of alkyne-modified cytosine nucleotides during extension.

[0431] The extension reaction may utilize a DNA polymerase lacking strand-displacement activity, such as T4 DNA polymerase or T7 DNA polymerase, to ensure specific extension without disruption of adjacent structures. The reaction mixture includes alkyne-modified cytosine nucleotide triphosphate, which can become incorporated into the extending strand.

[0432] In alternative embodiments, other functional groups may be incorporated instead of or in addition to alkynes. These may include amines, azides, dibenzocyclooctynes, tetrazines, trans-cyclooctenes, norbornenes, thiols, maleimides, bicyclononyne, aldehyde, hydrazine, oxyamine, strained alkene, orAttorney Docket No: 062954-509001 WO photocrosslinkable groups, or other bioorthogonal reactive groups. The choice of functional group may depend on the specific chemistry planned for subsequent analyte attachment.

[0433] The extension reaction may be performed at temperatures of 20°C to 95°C for 0.1 minutes or longer, with buffer compositions optimized for the specific polymerase. Following extension, the surface may be washed to remove unreacted nucleotides and other reaction components.Analyte Attachment

[0434] In some embodiments, analytes may be attached to the modified binding sites through bioorthogonal click chemistry reactions. For alkyne-modified binding sites, this attachment may utilize azide-modified analytes in a copper-catalyzed azide-alkyne cycloaddition (CuAAC) reaction.

[0435] The CuAAC reaction may be performed using copper sulfate (0.1 to 1 mM) with sodium ascorbate (1 to 10 mM) as a reducing agent, potentially with acceleration by tris(benzyltriazolylmethyl)amine (TBTA) or similar ligands (0.1 to 1 mM). The reaction may proceed at room temperature for 0.1 to 24 hours in aqueous buffer with or without organic co-solvents such as DMSO or tert-butanol (up to 50%).

[0436] In alternative embodiments, the attachment may utilize strain-promoted azide-alkyne cycloaddition, avoiding a possible need for copper catalysis. This approach may be particularly suitable for sensitive biological analytes that might be damaged by copper ions.

[0437] The analytes may comprise peptides, proteins, nucleic acids, carbohydrates, polymers, and / or small molecules. They may be modified with reactive moieties. For example, analytes may be modified with azide groups through various strategies, including:(a) Direct chemical synthesis incorporating azide-modified amino acids, nucleotides, or building blocks(b) Post-synthetic modification using azide-containing crosslinkers targeting lysine, cysteine, or other reactive residues(c) Enzymatic incorporation of azide-containing substrates(d) Metabolic labeling approaches for in vivo incorporation of azide-containing precursors

[0438] Following attachment, unreacted analytes may be removed through washing procedures of varying stringency, potentially including detergents, chaotropic agents, or competitive blocking agents to minimize non-specific binding.

[0439] In some embodiments, the analyte attachment step may be performed either before or after ULI incorporation, or before or after ULI bridge amplification, which may be particularly useful in applications where the analyte's presence might interfere with ULI amplification or mapping procedures.

[0440] In some embodiments the defined number and density of analyte binding sites on the solid support simplifies the preparation of samples by removing the requirement to quantify analyte concentration priorAttorney Docket No: 062954-509001 WO to introducing the sample to the solid support. The defined territories impose a limitation on the number and density of analyte molecules on the solid support, independent of analyte concentration in the sample solution. This avoids “over clustering” analytes, which is a common challenge, for example, across multiple DNA sequencing technologies.Multiple Analyte Binding Sites Within ULI Neighborhoods

[0441] While the embodiments described above detail the creation of ULI neighborhoods with a single analyte binding site, the present disclosure also encompasses various approaches for incorporating multiple analyte binding sites within a single ULI territory. This capability may be useful for multiplexed detection, detection of analyte co-localization, complexes, analyte-analyte interactions, increased analyte density, and enhanced functionality while maintaining the association with a single unique location identifier.

[0442] In some embodiments, the primary alkyne-modified binding site created through the processes described above may be further modified with multi-functional linkers to provide multiple attachment points having the same chemical moiety and / or multiple attachment points having multiple orthogonal chemical moieties. For example, multi-azide compounds such as octa-azide dendrimers, multi-arm PEG- azides, or branched azide-containing structures may be conjugated to the alkyne-modified binding site via copper-catalyzed azide-alkyne cycloaddition (CuAAC). In such reactions, one azide group of the multiazide compound participates in the reaction with the alkyne-modified binding site, while the remaining azide groups remain available for subsequent conjugation to alkyne-modified analytes. A single octa-azide compound, for instance, may provide seven additional azide moieties after attachment to the primary binding site, providing for capture of up to seven additional analytes within the same ULI neighborhood. In implementations utilizing multivalent linkers or dendrimeric extensions, steric hindrance at the binding site may reduce the efficiency of subsequent analyte conjugation steps. Linker length and branching density may be optimized to balance capture efficiency with spatial accessibility.

[0443] Alternative approaches may include the use of branched nucleic acid structures to create multiple binding sites within a single ULI territory. For example, Y-shaped or multi-armed DNA structures may be attached to the primary binding site, with each arm functionalized to capture specific analytes. In some embodiments, structured DNA assemblies such as branched or multi-arm nucleic acids may be used to increase binding valency, though integration with ULI zoning may require specialized design.

[0444] In certain embodiments, analytes are captures via hybridization, action of enzymes, protein-protein interaction, protein-nucleic acid (DNA, RNA, PNA, BNA, XNA...) and / or non-covalent interactions.

[0445] In certain applications, alternative scaffolds such as proteins or hybrid materials may offer additional valency or structural complexity, though their integration with ULI zoning remains to be demonstrated. Other embodiments may utilize protein scaffolds with multiple reactive sites, such asAttorney Docket No: 062954-509001 WO modified streptavidin molecules with additional chemical functionalities, or engineered antibody fragments with multiple conjugation sites for analyte attachment.

[0446] Regardless of the specific approach employed, the association between multiple analyte binding sites and a single ULI territory is maintained, preserving the core advantage of the ULI neighborhood concept: the transition from "one molecule finding one molecule" to "many identical molecules finding multiple related molecules." This capability further enhances the versatility of some embodiments across different experimental contexts and analytical needs, including multiplex diagnostics, simultaneous detection of interacting biomolecules, and high-density molecular profiling applications.

[0447] The methods described herein for creating multiple binding sites per ULI territory may be applied to any of the surface types, chemistries, and configurations detailed elsewhere in this disclosure, providing a flexible framework for customizing the ULI neighborhood architecture according to specific application requirements.

[0448] In some embodiments, ULI territories may be applied to the analysis of protein-drug interactions through physical separation techniques. The ULI neighborhoods provide a framework for assigning samplespecific indices to proteins separated through various fractionation methods, enabling multiplexed analysis while maintaining sample origin information. Sample indexing may be applied to proteins separated by various techniques including Size Exclusion Chromatography (SEC), Ion Exchange Chromatography (IEX), Hydrophobic Interaction Chromatography (HIC), Affinity Chromatography, and Thermal Shift Assays with Fractionation where proteins stabilized by ligand binding remain soluble at elevated temperatures. The sample-indexed ULI approach may be particularly valuable for multiplexed protein interaction studies. For example, proteins remaining in the soluble fraction at various temperature points in a thermal shift assay may be captured on beads with different sample indices, then pooled for downstream processing while maintaining the ability to determine which temperature condition each protein originated from. This approach may provide a direct indication of protein-ligand interactions, as the shift of a protein to a different temperature-indexed fraction may demonstrate stabilization by ligand binding. Fractions eluted from chromatography columns at different retention times may be directed to distinctly indexed territories. The shift in elution profile caused by protein-ligand or protein-protein interactions may be captured and preserved through the sample indexing system. This approach may enable high-throughput screening of interaction-induced shifts across thousands of proteins simultaneously, offering advantages over existing gel-based techniques that are typically limited to analyzing smaller numbers of proteins. The indexed ULI system may address challenges associated with mass spectrometry analysis of complex protein mixtures by preserving information about interaction-induced mobility shifts while facilitating pooled analysis.Attorney Docket No: 062954-509001 WOProcess 600 - Spatial Mapping of ULI NeighborhoodsMathematical Basis for Junction-Based Spatial Mappins

[0449] Process 600 can relate to spatial mapping of a ULI neighborhood, and can include use of development of: junction-based spatial mapping, ZIP code incorporation, an alternative ZIP code architecture, a ZIP pool size, a bridge oligonucleotide, junction detection, a junction-spanning amplicon, sequencing, data analysis, or alternative methods provided herein. Any aspect or combination of aspects from Process 600 may be included in a method herein.

[0450] The spatial mapping approach described herein benefits from concepts of computational geometry, particularly Delaunay triangulation and Voronoi diagrams. These mathematical frameworks may provide a theoretical foundation for detecting junctions between ULI territories to recover spatial organization.

[0451] When ULI neighborhoods are distributed across a surface, each forms a discrete territory containing multiple copies of a unique molecular identifier surrounding a central analyte binding site. These territories, under certain conditions, may naturally create a Voronoi-like pattern, where each region contains all points closer to its central binding site than to any other (FIG. 26A). The boundaries between these regions represent the junctions where different ULI sequences meet.

[0452] The junction detection system described herein, using orthogonal ZIP code pairs and directional bridge oligonucleotides, effectively samples these boundaries. By identifying which ULI territories are adjacent to one another, the system constructs a graph that represents the Delaunay triangulation of the ULI territories (FIG. 26B). In computational geometry, a Delaunay triangulation connects points that share Voronoi cell boundaries, which precisely corresponds to our junction-spanning approach.

[0453] This adjacency information is then processed using spring layout algorithms (such as Kamada- Kawai) to reconstruct the spatial arrangement. These algorithms position nodes (ULI territories) such that connected nodes (those with detected junctions) are placed at appropriate distances from each other, while disconnected nodes are positioned further apart (FIG. 26C). This transformation effectively converts sequence data from junction-spanning amplicons back into spatial coordinates.

[0454] Unlike pixel-based DNA microscopy approaches (Weinstein et al., 2019; WO2017044893A1) that treat each oligonucleotide as an individual spatial coordinate requiring exhaustive pairwise distance measurements through diffusive concatenation, the junction-focused method described herein selectively targets only the information-rich boundaries between molecular territories. In pixel-based methods, every molecule is uniquely barcoded and spatial reconstruction requires measuring or inferring proximities among potentially 107-109individual molecules. In contrast, the territorial approach organizes molecules into discrete clusters defined by shared ULI sequences, with junction detection performed selectively at territorial boundaries via designed ZIP code complementarity. This architecture reduces sequencing requirements to scale with territory count (potentially 103-106territories) rather than total molecule count,Attorney Docket No: 062954-509001 WO while maintaining spatial resolution at the territory scale. The adjacency information obtained from junction detection is then processed using spring layout algorithms (such as Kamada-Kawai) to reconstruct the spatial arrangement, effectively converting sequence data from junction-spanning amplicons back into spatial coordinates without requiring optical imaging or external positional references. This territorial junction mapping approach provides a more efficient method for reconstructing spatial arrangements from sequence data alone, as the number of informative junctions scales with territory count rather than requiring exhaustive measurement of all molecular proximities.

[0455] The accuracy of this reconstruction depends on the density of ULI territories and the efficiency of junction detection. With sufficient territory density and junction sampling, positional errors can be reduced to the scale of individual ULI territory sizes, providing spatial mapping of molecular distributions over large areas.

[0456] This approach transforms DNA sequence information into spatial coordinates through a mathematical framework that exploits the natural adjacency relationships between molecular territories, providing a powerful new tool for spatial biology and molecular diagnostics.ZIP Code Incorporation for Spatial Mapping

[0457] To provide spatial mapping of ULI territories, short hybridization regions termed ZIP codes (Zone Interaction Pairing codes) may be incorporated at specific positions within ULI-containing oligonucleotides. These ZIP codes act as addressable pairing handles for bridge oligonucleotides, providing precise and sequence-resolved detection of junctions between neighboring molecular neighborhoods.

[0458] Each ZIP code may comprise a 15 to 30 nucleotide sequence selected from a pool of N distinct ZIP pairs, where each may consist of two unique, orthogonal sequences:

[0459] ZIPna: the upstream or “a-position” ZIP code for pair n

[0460] ZIPnb: the downstream or “b-position” ZIP code for the same pair n

[0461] This creates the set of 2N total ZIP codes, {ZIPla, ZIPlb, ZIP2a, ZIP2b, ..., ZIPNa, ZIPNb}. Sequences may be designed for minimal cross-hybridization and balanced thermodynamic properties, with exemplary melting temperatures between 60°C and 75°C. Computational optimization may be used to ensure both specificity and compatibility across all ZIP codes.

[0462] In this embodiment, ZIP codes may be incorporated during the initial synthesis of ULI-containing strands. Each ULI is assigned a single ZIP pair, which may be embedded into a defined sequence layout as follows:

[0463] CT2 - H7 - ZIPna - iCiC - UHS - ULI - H5a - ZIPnb - H5b• ZIPna may be placed between the H7 primer region and the iCiC terminatorAttorney Docket No: 062954-509001 WO• ZIPnb may be inserted between two adjacent primer segments, H5a and H5b, which together constitute the H5 priming region

[0464] Importantly, in this exemplary design the H5 primer is split into two regions, H5a and H5b, to flank the ZIPnb site. Either H5a or H5b may independently serve as a generalized primer hybrdization site in downstream reactions, providing flexible PCR strategies.

[0465] Each ULI-containing oligo thus carries a unique ZIPna-ZIPnb pair, and adjacent territories are likely to have differing ZIP pairs due to random or pseudo-random assignment across the pool.

[0466] Because the bridge oligonucleotides are designed to hybridize directionally from ZIPma (5' end of a neighbor) to ZIPnb (3' end of the target), the total number of valid, non-redundant bridge combinations across an N-pair ZIP pool is N x (N - 1). This includes all possible directional pairings between ZIP pairs

[0467] In alternative embodiments, ZIP codes may be introduced through post-synthetic ligation, strand displacement, or ZIP-containing adaptors, rather than encoded during initial ULI synthesis. In all cases, the ability to distinguish ULI-ULI junctions via short, orthogonal, directionally assigned ZIP codes is critical to the spatial reconstruction strategy described herein.

[0468] The bridge oligonucleotide approach described herein represents one implementation of the junction-selective amplification principle. Alternative implementations using inherent complementarity within ZIP code architectures are described below:Alternative ZIP Code Architecture Using Complementary K-mer Elements

[0469] In alternative embodiments, ZIP codes may be designed with inherent complementarity that provides for direct hybridization between adjacent ULI territories without requiring exogenous bridge oligonucleotides. This approach utilizes a structured k-mer architecture in which each ZIP code comprises two functional components: a first component containing a defined subset of shared k-mer sequences, and a second component containing the reverse complement of a k-mer absent from the first component.

[0470] In some embodiments, a system of N ZIP codes may be constructed using N+l distinct k-mer sequences, where each k-mer may comprise 4 to 12 nucleotides, 5 to 10 nucleotides, or 6 to 8 nucleotides. Each ZIP code may be uniquely defined by the absence of exactly one k-mer from the shared pool. The missing k-mer for each ZIP code may then be provided in its reverse complement form as the second component of that ZIP code.

[0471] Table 1 illustrates an exemplary construction scheme for a system of 4 ZIP codes using 5 k-mers (designated k_A through k_E): Table 1: ZIP code construction using complementary k-mer architectureAttorney Docket No: 062954-509001 WOwhere k_X' denotes the reverse complement of k-mer X.

[0472] This architecture extends to systems of any size N, with the general relationship requiring N+l k- mers to create N unique ZIP codes. For example, a system of 10 ZIP codes may utilize 11 k-mers, a system of 20 ZIP codes may utilize 21 k-mers, and so forth. The k-mers may be selected to minimize secondary structure formation and cross-hybridization, with melting temperatures balanced across the pool. In some embodiments, k-mers may be designed with melting temperatures between 45°C and 75°C, between 50°C and 70°C, or between 55°C and 65°C.

[0473] In some embodiments, the complementary k-mer component may be incorporated during initial oligonucleotide synthesis on the 5' terminus of a ligation oligonucleotide, downstream of an internal spacer sequence. During bridge amplification, this complementary sequence may be copied to generate the complete ZIP code on surface-bound oligonucleotides. Following bridge amplification and strand processing, oligonucleotides within each ULI cluster may contain the full ZIP code sequence at their 5' termini.

[0474] The mechanism of junction detection in this embodiment relies on selective hybridization between complementary k-mer elements at territorial boundaries. Within a single ULI cluster, all oligonucleotides carry identical ZIP codes and therefore contain no complementary sequences capable of driving stable hybridization to one another. However, at the interface between two different ULI clusters that have been assigned different ZIP codes, the complementary k-mer present in one ZIP code may hybridize to the corresponding k-mer in the adjacent ZIP code. This design creates a molecular logic gate: amplification occurs if and only if two different ZIP codes are present, enabling detection of territorial boundaries while suppressing background signal from homogeneous regions.

[0475] For example, ZIP 1 (containing k_D' as its second component) may hybridize to ZIP 2 (containing k_D in its first component) via the k_D / k_D' complementary pair. Similarly, ZIP 2 (containing k_C as its second component) may hybridize to ZIP 1 (containing k_C in its first component) via the k_C / k_C complementary pair. This bidirectional complementarity between any two different ZIP codes provides for robust junction detection at territorial boundaries.

[0476] Following hybridization at a junction between territories, polymerase extension from the hybridized complementary k-mer may copy the upstream portion of the adjacent strand, including its distinct ULI sequence. The extended product thus contains ULI sequence information from both adjacent territories. These junction-spanning products may be selectively cleaved from surface-bound strands and isolated for sequencing, providing a direct readout of territorial adjacency relationships.Attorney Docket No: 062954-509001 WO

[0477] This k-mer complement approach offers several advantages over bridge oligonucleotide-based methods. First, it eliminates the need to synthesize and introduce a complex library of N x (N - 1) bridge oligonucleotides, reducing reagent cost and simplifying workflow implementation. Second, the inherent complementarity between ZIP codes provides for spontaneous junction formation at all heterotypic boundaries without requiring precise stoichiometric control of multiple bridge species. Third, the modular k-mer design facilitates computational optimization of sequence pools for minimal cross-reactivity and balanced thermodynamic properties.

[0478] The probability that any two adjacent ULI territories will successfully form a detectable junction depends on the diversity of ZIP code assignment across the population. If ZIP codes are randomly or pseudo-randomly distributed, the probability that two neighboring territories carry different ZIP codes may be approximately (N - 1) / N. For a system with N = 10, this probability may be approximately 90%; for N = 20, approximately 95%; and for N = 50, approximately 98%. Higher values of N thus increase the likelihood of junction detection across the entire population of territories.

[0479] In some embodiments, the k-mers in the first component of each ZIP code may be arranged in a defined order, while in other embodiments they may be concatenated in any order that maintains consistent relative positioning across the ZIP code pool. The order of k-mers may be optimized to minimize formation of stable secondary structures or unintended hybridization between non-complementary elements.

[0480] In some embodiments, modified nucleotides may be incorporated into k-mer sequences to enhance specificity or stability. Such modifications may include locked nucleic acids (LNAs), peptide nucleic acids (PNAs), 2'-O-methyl RNA, 2'-fluoro RNA, phosphorothioate linkages, or other chemical modifications that increase duplex stability or nuclease resistance.

[0481] Following junction formation and polymerase extension, the products may be amplified using primers that hybridize to conserved adapter sequences flanking the ZIP code and ULI regions. Amplification may be performed on-surface or following release of junction-spanning products into solution. The resulting amplicons may be purified using size-selection methods, sequence-specific capture, or a combination thereof, as described elsewhere herein.

[0482] In some embodiments, junction detection using complementary k-mer ZIP codes may be performed in solution phase rather than on a solid support. In such implementations, ULI-containing oligonucleotides with different ZIP codes may be combined under conditions that promote hybridization between complementary k-mer elements. Polymerase extension and ligation steps may then proceed to generate junction-spanning products that encode adjacency information. This solution-phase approach may be useful for validating ZIP code designs, optimizing reaction conditions, or performing specialized mapping applications that do not require surface-bound molecular territories.Attorney Docket No: 062954-509001 WO

[0483] Detection of junction-spanning products may be performed by gel electrophoresis, quantitative PCR, or direct sequencing. In gel-based detection, successful junction formation and extension produces amplicons of predictable size that may be distinguished from unamplified starting material or background products. The presence or absence of such amplicons provides a binary readout of whether complementary ZIP codes were present in the reaction, enabling validation of the junction-selectivity principle.

[0484] The k-mer complement approach may be combined with other features described herein, including iso-nucleotide termination motifs (iCiC / iGiG), restriction site cassettes, USER-mediated strand cleavage, and universal hybridization sites. These elements may be integrated into oligonucleotide designs to provide additional control over extension termination, strand processing, and amplification specificity.

[0485] In some embodiments, each ULI-containing oligonucleotide may be assigned a ZIP code comprising a first component with 3 to 10 k-mers from the shared pool, 4 to 8 k-mers, or 5 to 6 k-mers, and a second component comprising the reverse complement of the single absent k-mer. The total length of each complete ZIP code may range from 25 to 100 nucleotides, 30 to 80 nucleotides, or 35 to 60 nucleotides, depending on k-mer length and number.

[0486] Junction-spanning products generated by this method may be sequenced using standard nextgeneration sequencing platforms, including sequencing-by-synthesis, nanopore sequencing, or singlemolecule real-time sequencing. Sequence reads containing ULI pairs may be computationally identified and used to construct adjacency graphs for spatial reconstruction, as described elsewhere herein.

[0487] This complementary k-mer architecture represents a distinct approach to encoding and detecting territorial boundaries in spatially organized molecular systems. By embedding complementarity directly into the ZIP code design, this method provides an efficient, scalable alternative to bridge oligonucleotide- mediated junction detection while maintaining the fundamental principle of selective boundary amplification that distinguishes ULI-based spatial mapping from pixel-based approaches.Considerations for Choosing ZIP Pool Size (N)

[0488] The number of ZIP pairs (N) used in a given implementation may be tuned based on the desired spatial resolution and density of ULI territories. Because spatial mapping relies on the presence of distinct ZIP pairs at adjacent territories, any two neighboring ULI zones assigned the same ZIP pair (ZIPna, ZIPnb) may produce no detectable bridge between them, resulting in a local dropout in the junction graph. To minimize the frequency of these undetectable junctions, N may be chosen such that the probability of neighboring territories sharing the same ZIP pair is acceptably low.

[0489] The optimal value of N depends on several factors, including:1) Average number of neighbors per ULI zone (local topology)2) Total number of ULI zones being mappedAttorney Docket No: 062954-509001 WO3) Tolerable dropout rate in the final spatial graph4) Redundancy or overlap built into the mapping algorithm5) Bridge library synthesis constraints

[0490] In typical applications, N values between 5 and 50 may be used, with higher values reducing the risk of adjacent ZIP pair collisions. For example, with N = 20 and under the simplifying assumption of uniform ZIP assignment, the probability of adjacent zones sharing the same ZIP pair may be approximately 1 in 20, though real-world topologies may differ. Larger ZIP pools reduce the probability of junction collisions between adjacent territories and may be particularly useful when comprehensive spatial reconstruction is desired.Bridge Oligonucleotide Design and Synthesis

[0491] Bridge oligonucleotides may be designed to span the junctions between adjacent ULI territories by hybridizing to the unique ZIP codes present on neighboring strands. Each ULLcontaining oligonucleotide may be assigned a ZIP pair, denoted ZIPna and ZIPnb, where n is an integer from 1 to N (the total number of ZIP pairs). The ZIPna sequence may be positioned between the H7 region and the iCiC motif, while ZIPnb may be embedded between the H5a and H5b segments of the H5 primer region.

[0492] To detect spatial boundaries between territories, bridge oligonucleotides may be constructed to hybridize directionally across ZIP codes from two different ULI neighborhoods. Each bridge oligo may comprise the general structure:

[0493] 3'-ZIPnb'-[templated linker]-ZIPma'-5'

[0494] Where: (m n)• ZIPnb' is the reverse complement of the ZIPnb sequence from ULI pair n• ZIPma' is the reverse complement of the ZIPma sequence from ULI pair m• The templated linker may be a defined sequence that ensures correct spatial separation between the ZIP-binding regions and promotes effective ligation and / or extension across the junction

[0495] These bridge oligonucleotides may create unidirectional connections from ZIPma (the "a" position of one ULI) to ZIPnb (the "b" position of a neighboring ULI). To enforce this directionality and avoid redundancy, the bridge set includes only one orientation of each possible ZIP pair combination. For a ZIP pool of size N, this yields a total of N x (N - 1) distinct directional bridges.

[0496] Bridges may be synthesized via standard solid-phase oligonucleotide synthesis, with optional postsynthetic incorporation of modified bases, linkers, or labels. These directional bridges are useful to provide detection of physical proximity between ULI neighborhoods, forming the backbone of the sequence-based spatial reconstruction approach. Depending on the implementation, bridge oligonucleotides may be designed for unidirectional or bidirectional pairing across ZIP codes.Attorney Docket No: 062954-509001 WO.Junction Detection Through Extension and Ligation

[0497] The process for detecting spatial relationships between ULI territories involves a series of enzymatic steps designed to create junction-spanning amplicons, DNA constructs that encode physical adjacency between two molecular neighborhoods. These junctions are only detectable when the adjacent ULI territories possess distinct ZIP code pairs, providing for hybridization of a directional bridge oligonucleotide between them.

[0498] In some embodiments, this process begins with an initial extension from the H5a region, using a DNA polymerase to generate complementary strands that include the ULI and flanking sequences. Depending on the strand architecture, the polymerase may or may not require strand-displacement activity.

[0499] Following this extension, bridge oligonucleotides may be introduced to hybridize across ZIP regions between adjacent ULI-containing strands. Each bridge may be designed to pair with ZIPma (from one ULI) and ZIPnb (from its neighbor), forming a directional connection only if their ZIP pairs differ. The annealing conditions may range from 40°C to 65°C for 15 minutes to 2 or more hours, with optional temperature ramping to improve specificity.

[0500] Once annealed, the 3' end of the ULI complement strand may align adjacent to the 5' end of the bridge oligonucleotide. These ends may be ligated using DNA ligases such as T4, T7, or thermostable variants, with ligation performed at 16°C to 65 °C for 1 to 24 hours. Reaction buffers may include crowding agents such as PEG to improve ligation efficiency, especially for low-yield or sterically hindered complexes.

[0501] After ligation, a non-displacing DNA polymerase may be used to fill the gap between the ligated bridge and the adjacent iCiC site. This step incorporates iso-guanine nucleotides to pair with iso-cytosine bases in the ULI template strand. A second ligation step may then connect the 3' end of the extended product to the 5' end of the neighboring ULI complement, completing a junction-spanning amplicon that encodes a unique ZIPma-ZIPnb connection.Alternative Junction Detection Strategies

[0502] In some embodiments, alternative enzymatic strategies may be employed for junction detection, including:• Gibson assembly or other overlap-extension protocols• Proximity ligation techniques as used in chromosome conformation capture (e.g., Hi-C)• Strand invasion or branch migration mechanisms inspired by homologous recombinationAttorney Docket No: 062954-509001 WOPurification of Junction-Spanning Amplicons

[0503] Following formation of the junction-spanning products, a purification step enriches for full-length constructs over intermediate or incomplete ones.

[0504] Size-based selection methods may be employed, including:• Gel extraction (agarose or polyacrylamide)• Size-exclusion chromatography• Bead-based separation using paramagnetic beads

[0505] The target size range (typically 100-300 nt) distinguishes complete junction-spanning constructs from partial products (30-100 nt). Alternatively, sequence-specific enrichment strategies may be used:• Biotinylated capture oligonucleotides targeting junction motifs• Hybrid-capture protocols similar to those in target-enriched sequencing• Junction-specific PCR, using primers that span the ZIPma-ZIPnb boundary

[0506] In some embodiments, combined size-selection and sequence-specific capture may be used for maximum specificity. Importantly, perfect purification is not required; downstream sequencing and analysis can computationally filter out non-junction reads based on amplicon structure.Sequencing and Data Analysis

[0507] Purified junction-spanning amplicons are then sequenced using standard next-generation platforms such as Illumina, PacBio, or Oxford Nanopore. Library preparation may involve adapter ligation and indexing for multiplexing, with sequencing depth tailored to the complexity of the spatial landscape (typically 105to 109reads).

[0508] The sequencing data may then be processed using computational pipelines that may:• Identify and extract ULI sequences and associated ZIP codes from each read, and / or• Filter out improperly structured reads or incomplete junctions, and / or• Construct a graph where:(a) Each ULI (or ZIP pair) represents a node(b) Each detected junction (ZIPma-ZIPnb pair) represents a directed edge

[0509] A spring-layout algorithm or other graph-embedding technique may then be applied to reconstruct the spatial arrangement of ULI territories. In some embodiments, the layout algorithm may be adapted from DNA microscopy methods (e.g., Boulgakov et al., 2018), but modified for discrete and non-continuous node distributions. These modifications may include some or all of:• Adjusted force models to account for expected ULI territory spacing• Constraints to enforce minimum spacing between non-neighbors• Weighted edges based on junction detection frequencyAttorney Docket No: 062954-509001 WO• Hierarchical or tiled layout for large-scale samples (thousands of ULIs)

[0510] The result may be a spatial reconstruction of the original molecular layout, derived from sequence data.Machine Learning Enhancements

[0511] In some embodiments, machine learning models may be used to improve spatial inference, including some or all of:• Neural networks trained on simulated or reference spatial maps• Probabilistic models that incorporate priors on ULI distribution• Ensemble methods combining graph-based and learned models for optimal layout reconstruction

[0512] The junction-spanning spatial mapping approach may be combined with imaging for cross- validation or multi-modal analysis, in situ sequencing approaches for enhanced spatial resolution, and computational deconvolution methods for complex tissue analysis. The spatial mapping method is compatible with, but independent of, ULI neighborhoods and may be applied to alternative spatial systems such as random oligo deposition, lithographically patterned arrays, or surface-encoded molecular zones.Alternative Applications of Junction-Selective Amplification

[0513] While the methods and compositions described above have particular utility in spatial molecular mapping applications, the fundamental principle of junction-selective amplification, wherein amplification occurs at boundaries between different molecular identities but not within homogeneous populations, may be applied to diverse contexts where detection of molecular co-localization, proximity, or co-occurrence is desired. The ZIP code architectures described herein, including both bridge oligonucleotide-mediated designs and complementary k-mer-based designs, are not limited to surface-based implementations or ULI- containing sequences, and may be adapted for use in proximity detection assays, DNA computing applications, and other analytical or diagnostic contexts.Proximity Detection in Cellular and Molecular Assays

[0514] In some embodiments, oligonucleotides containing ZIP codes may be conjugated to binding reagents such as antibodies, aptamers, nanobodies, or other affinity molecules that recognize distinct cellular markers or biomolecules. When multiple markers are co-expressed on the same cell or within proximity distance allowing molecular interaction, the conjugated oligonucleotides may hybridize via their complementary elements and undergo polymerase extension to generate a detectable amplification product. In the absence of co-localization, no amplification product is formed. For example, a first antibody recognizing cell surface marker A may be conjugated to an oligonucleotide containing a first ZIP code, andAttorney Docket No: 062954-509001 WO a second antibody recognizing cell surface marker B may be conjugated to a second ZIP code. In a heterogeneous cell population, amplification signal is generated only from cells expressing both markers, providing AND-gate logic for detection of cellular phenotypes. This approach may be scaled to larger marker panels by assigning different ZIP codes to each antibody and sequencing the resulting junction products to determine which pairwise combinations of markers were co-localized.

[0515] This approach differs from existing proximity ligation assays (PLA) in fundamental architecture. In PLA, each probe oligonucleotide is designed to pair with one specific complementary probe, such that each antibody conjugate can only participate in detection of one predetermined marker combination. To detect multiple pairwise combinations among N markers, separate populations of antibody conjugates may be prepared, each designed for a specific pairing. In contrast, the ZIP code architectures described herein enable each antibody conjugate to participate in junction formation with any other different marker in the panel. A system of N markers requires only N distinct ZIP-coded antibody conjugates to enable detection of all pairwise combinations, as any two different ZIP codes may form functional junctions via their complementary elements without requiring pre-specification of pairing relationships.

[0516] In some embodiments, the binding reagents may be aptamers that recognize small molecules, metabolites, or other non-protein targets, enabling detection of molecular co-occurrence or co-localization. The oligonucleotides may contain unique identifier sequences such as ULIs or molecular barcodes to enable quantification or highly multiplexed detection of many marker combinations simultaneously. Following amplification, the products may be detected by gel electrophoresis, quantitative PCR, digital droplet PCR, or sequencing.DNA-Based Logical Computation

[0517] In some embodiments, the ZIP code architectures described herein may be used to implement logical operations in DNA computing applications. Because amplification occurs if and only if two different ZIP codes are present in the same reaction volume, the system functions as a molecular AND gate where output is generated only when multiple distinct inputs are present together. More complex logical circuits may be constructed by combining multiple ZIP code pairs with distinct compositions, enabling implementation of multi-input logical functions. Input signals may be represented by the presence or absence of specific ZIP-coded oligonucleotides, and output signals may be detected by amplification, fluorescence, colorimetric readout, or other detection modalities. The ZIP-coded oligonucleotides used in DNA computing applications need not contain ULI sequences, and may be designed with any functional elements appropriate for the computational task, including recognition sequences for enzymatic processing or reporter sequences for signal detection. The computational operations may be performed entirely in solution phase without requiring surface immobilization.Attorney Docket No: 062954-509001 WO

[0518] This approach differs from existing DNA computing implementations based on toehold-mediated strand displacement in that toehold systems require precise tuning of domain lengths and binding strengths, and may exhibit leak reactions where output is generated in the absence of correct inputs. The ZIP code architectures described herein provide intrinsic selectivity: amplification is thermodynamically unfavorable within homogeneous ZIP code populations due to the absence of complementary sequences capable of forming stable duplexes, creating a high thermodynamic barrier to background amplification without requiring kinetic discrimination. Additionally, the modular design of ZIP code pools facilitates systematic expansion of gate complexity. In k-mer complement systems, adding one additional k-mer to the pool enables creation of one additional ZIP code that may participate in junction formation with all existing ZIP codes, without redesigning existing sequences.Orthogonal ZIP Code Pools for Independent Amplification Networks

[0519] In some embodiments, multiple independent sets of ZIP codes may be designed to operate orthogonally within the same reaction volume or spatial context. Orthogonality is achieved by constructing each set from a distinct pool of k-mers that lack complementarity to k-mers in other sets, such that ZIP codes from different sets cannot form productive junctions with one another. For example, a first ZIP code set may be constructed from k-mer pool A comprising k-mers Al through A(N+1), and a second ZIP code set may be constructed from k-mer pool B comprising k-mers B 1 through B(M+ 1). If k-mer pools A and B are designed such that no k-mer from pool A shares significant sequence complementarity with any k-mer from pool B, then ZIP codes from set A may form junctions only with other ZIP codes from set A, and ZIP codes from set B may form junctions only with other ZIP codes from set B.

[0520] This orthogonality enables several applications. In spatial mapping contexts, different orthogonal ZIP code sets may be used to encode different classes of molecular information or to map different biological structures simultaneously without cross-interference. In proximity detection applications, orthogonal ZIP code sets enable multiplexed detection of multiple independent co-localization events, for example, antibodies targeting one set of markers may be conjugated to ZIP codes from set A, while antibodies targeting a different set of markers may be conjugated to ZIP codes from set B, such that colocalization of two markers within the same class (e.g., two immune markers or two tumor markers) generates amplification products, while co-localization of markers from different classes (e.g., one immune marker with one tumor marker) does not. In DNA computing applications, orthogonal ZIP code sets may implement independent logic circuits that operate in parallel without crosstalk.

[0521] The k-mers in each orthogonal set may be selected using computational optimization methods that maximize thermodynamic separation between intended within-set hybridization events and unintended between-set hybridization events. In some embodiments, k-mer pools may be designed using distinctAttorney Docket No: 062954-509001 WO nucleotide composition constraints to enhance orthogonality, such as GC-rich k-mers in one pool and AT- rich k-mers in another. In some embodiments, orthogonal ZIP code sets may employ different amplification primer binding sites, enabling selective amplification of junction products from one set or the other by choice of primers. In bridge oligonucleotide-based ZIP code systems, orthogonality may similarly be achieved by designing distinct sets of ZIP code sequences and corresponding bridge oligonucleotides that do not cross-react.Additional Applications

[0522] The junction-selective amplification principle may also be adapted for nucleic acid co-localization detection, such as in situ detection of viral co-infection, fusion transcript detection, or chromosomal translocation events by assigning different ZIP codes to probes targeting different nucleic acid species. In stimulus-responsive biosensing applications, ZIP-coded oligonucleotides may be conjugated to molecular switches or embedded in responsive materials such that conformational changes or environmental stimuli bring different ZIP codes into proximity, enabling conditional amplification. In therapeutic or diagnostic contexts, ZIP-coded oligonucleotides may be conjugated to targeting ligands recognizing disease markers, enabling dual-marker-dependent signal generation or conditional therapeutic activation in cells expressing multiple targets. These and other implementations leverage the same fundamental design principle of conditional amplification based on molecular diversity.General Considerations

[0523] These applications may utilize any of the ZIP code architectures described herein. Bridge oligonucleotide-based systems may be advantageous when high specificity and precise control over junction formation kinetics are desired. Complementary k-mer-based systems may be advantageous when simplified workflows, reduced reagent costs, or high scalability are priorities. The choice of ZIP code architecture for a particular application may depend on factors such as the number of different molecular species to be distinguished, the desired level of multiplexing, the required sensitivity and dynamic range, and the complexity of the sample matrix.

[0524] The amplification products generated in these applications may be detected using gel electrophoresis, quantitative PCR, digital droplet PCR, next-generation sequencing, nanopore sequencing, or other suitable methods. In applications where sequence information is not required, amplification products may be detected in bulk using intercalating dyes or other non-specific nucleic acid detection reagents. In applications where identification of specific junction combinations is required, the amplification products may be sequenced to decode the identities of the participating ZIP codes. The implementations described herein are exemplary and non-limiting, and the junction-selective amplificationAttorney Docket No: 062954-509001 WO principle may be adapted to numerous other contexts where detection of co-localization, proximity, or cooccurrence is desired.Process 700 - Integration

[0525] The protein sequencing methods of Processes 100-400, ULI neighborhood technologies of Process 500, and spatial mapping approaches of Process 600 may be implemented independently or in various combinations to address specific analytical requirements. Integration of these technologies provides enhanced capabilities while maintaining the flexibility to deploy individual processes as needed.

[0526] Process 700 can relate to integration of other aspects or methods provided herein, such as aspects or methods of Processes 100-600. Process 700 may include use of development of: protein sequencing, a ULI neighborhood, a solution-phase, a surface-based approach, spatial biology, multiplexed analysis, multiple attachment points, in situ digestion, a diagnostic methor or application, a microfluidic platform, an alternative molecular architecture provided herein, automated or high-throughput methods or sequencing, platform compatibility or standardization. Any aspect or combination of aspects from Process 700 may be included in a method herein.Integration of Protein Sequencing with ULI Neighborhoods

[0527] In some embodiments, the protein sequencing methods may be enhanced through integration with ULI neighborhood technologies. The ULI neighborhoods provide a framework for assigning location information to individual amino acids cleaved from peptides, enhancing detection efficiency and reducing information loss d...

Claims

1. Attorney Docket No: 062954-509001 WOCLAIMS1. A method for determining amino acid identity and positional information, comprising:(a) providing a peptide, wherein the peptide is coupled to a solid support such that a N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions;(b) providing a chemically -reactive conjugate, the chemically-reactive conjugate comprising: (x) a location nucleic acid, (y) a reactive moiety for binding the N-terminal amino acid residue of the peptide, and (z) an immobilizing moiety for immobilization to the solid support;(c) contacting the peptide with the chemically-reactive conjugate, thereby coupling the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex;(d) immobilizing the conjugate complex to the solid support via the immobilizing moiety;(e) cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby exposing a next amino acid residue as a second N-terminal amino acid residue on the cleaved peptide and providing an immobilized location complex, the immobilized location complex comprising the cleaved and separated N-terminal amino acid residue and the location nucleic acid;(f) providing a second chemically-reactive conjugate, the second chemically-reactive conjugate comprising: (xx) a cycle tag comprising a cycle nucleic acid associated with a cycle number, and (yy) a second reactive moiety for binding and cleaving the second N-terminal amino acid residue of the peptide;(g) contacting the second N-terminal amino acid residue of the peptide with the second chemically-reactive conjugate, thereby coupling the second chemically-reactive conjugate to the second N-terminal amino acid of the peptide to form a coupled-conjugate-complex;(h) forming a transfer complex comprising the immobilized location complex and the coupled- conjugate-complex, thereby bringing the location nucleic acid into proximity with the cycle nucleic acid;(i) within the transfer complex, joining the location nucleic acid or a reverse complement thereof to the cycle nucleic acid to form a recode unit, or otherwise joining information of the location nucleic acid and the cycle nucleic acid, thereby creating an immobilized recode unit comprising location and cycle information of the location nucleic acid and cycle nucleic acid;(j) cleaving and thereby separating the second N-terminal amino acid residue from the peptide, thereby exposing another next amino acid residue as a third N-terminal amino acid residue on the cleaved peptide and liberating a recode unit complex from the solid support, the liberated recode unit complex comprising the cleaved and separated second N-terminal amino acid residue the recode unit, wherein the recode unit is no longer immobilized;(k) contacting the recode unit complex with a binding agent, the binding agent comprising: a binding moiety for preferentially binding to the second N-terminal amino acid residue of the recode unitAttorney Docket No: 062954-509001 WO complex, and a recode tag comprising a recode nucleic acid corresponding with the binding agent, thereby forming an affinity complex, the affinity complex comprising the recode unit complex and the binding agent, and thereby bringing the recode nucleic acid into proximity with the recode unit;(l) transferring information of the recode nucleic acid to the recode unit to generate a recode block;(m) obtaining sequence information of the recode block; and(n) based on the obtained sequence information, determining identity and positional information of the second amino acid residue of the peptide.

2. The method of claim 1 , wherein forming the transfer complex comprises hybridizing a portion of the location nucleic acid with a portion of the cycle nucleic acid.

3. The method of claim 1 , wherein joining the location nucleic acid or a reverse complement thereof to the cycle nucleic acid comprises extending a 3’ end of the cycle tag using the location tag as a template.

4. The method of claim 1 , further comprising repeating steps (f) through (j) for subsequent amino acids of the peptide.

5. The method of claim 1, further comprising washing the immobilized location complex before (f) or before (g).

6. The method of claim 1, further comprising washing the coupled-conjugate-complex before (h).

7. The method of claim 1, further comprising determining a likely three-dimensional structure of the peptide based on the obtained sequence information.

8. The method of claim 1, wherein the location nucleic acid comprises deaza-DNA.

9. The method of claim 1 , wherein the cycle nucleic acid comprises deaza-DNA.

10. The method of claim 1, wherein the recode nucleic acid comprises DNA or RNA.

11. The method of claim 1 , wherein the recode nucleic acid comprises a hybridization-capable code.

12. The method of claim 1, wherein the recode nucleic acid comprises a non-colliding code in an additive vector space.

13. The method of claim 1, wherein any of (b)-(j) are performed in the presence of a Lewis acid, and in the absence of trifluoroacetic acid.

14. The method of claim 1, wherein the binding moiety comprises a peptide, antibody, antibody fragment, antibody derivative, or aptamer.

15. The method of claim 1, wherein the binding moiety binds to a natural amino acid, a post- translationally modified amino acid, a derivatized version of an amino acid, a derivatized or stabilized version of a post-translationally modified amino acid, a synthetic amino acid, an amino acid with a specific side chain, an amino acid with a phosphorylated side chain, an amino acid with a glycosylatedAttorney Docket No: 062954-509001 WO side chain, an amino acid with a methylation modification, or a D-amino acid, phenylthiohydantoin (PTH) derivative or anilinothiazolinone (ATZ) derivative of an amino acid, or binds to a combination thereof.

16. The method of claim 1, wherein the binding moiety binds covalently or non-covalently to the recode unit complex.

17. The method of claim 1, wherein the solid support comprises a bead, a plate, a chip, a glass slide, silica, a resin, a gel, a hydrogel, a membrane, polystyrene, a metal, nitrocellulose, a mineral, plastic, polyacrylamide, latex, or ceramic.

18. The method of claim 1, wherein the peptide comprises a hormone, neurotransmitter, enzyme, antibody, viral protein, bacterial protein, synthetic peptide, bioactive peptide, peptide hormone, oligopeptide, polypeptide, fusion protein, cyclic peptide, branched peptide, recombinant protein, tumor marker, therapeutic peptide, antigenic peptide, or signaling peptide.

19. The method of claim 1, wherein the peptide is derived from a cell lysate, blood sample, plasma sample, serum sample, tissue biopsy, saliva sample, urine sample, cerebrospinal fluid sample, sweat sample, synovial fluid sample, fecal sample, gut microbiome sample, environmental water sample, soil sample, bacterial culture, viral culture, organoid, tumor biopsy, sputum sample, or hair sample.

20. The method of claim 1 , wherein the peptide is associated with a disease.

21. The method of claim 1, wherein said transferring information comprises performing nucleic acid amplification, enzymatic ligation, splint ligation, chemical ligation, template-assisted ligation, use of a ligase enzyme, use of a splint oligonucleotide, use of a catalyst, use of a bridging molecule, use of a condensation agent, use of a coupling reagent, use of a polymerase enzyme, use of a complementary nucleic acid sequence, use of a nicking enzyme, use of a nucleic acid modifying enzyme, use of a recombinase, use of a strand-displacing polymerase, use of a single-strand binding protein, a click chemistry reaction, a phosphodiester bond formation, or a peptide nucleic acid-mediated ligation.

22. The method of claim 1 , wherein the information of the recode nucleic acid comprises a sequence of the recode nucleic acid or a reverse complement of the sequence of the recode nucleic acid.

23. The method of claim 1, wherein said transferring information comprises joining the recode nucleic acid or a reverse complement of the recode nucleic acid with the cycle nucleic acid.

24. The method of claim 1, wherein the second N-terminal amino acid is identified as amino acid position 2, 3, 4, 5, 6, 7, 8, 9, 10, or grearter, within the peptide, based on sequenced information of the cycle nucleic acid.

25. A method for determining identity and positional information of a plurality of amino acid residues of a peptide, the peptide comprising n amino acid residues, the method comprising:Attorney Docket No: 062954-509001 WO(a) coupling the peptide to a solid support such that a N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions;(b) providing a chemically -reactive conjugate, the chemically-reactive conjugate comprising: (x) a location nucleic acid, (y) a reactive moiety for binding and cleaving the N-terminal amino acid residue of the peptide and exposing a next amino acid residue as a N-terminal amino acid residue on the cleaved peptide, and (z) an immobilizing moiety for immobilization to the solid support;(c) contacting the peptide with the chemically-reactive conjugate, thereby coupling the chemically-reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex;(d) immobilizing the conjugate complex to the solid support via the immobilizing moiety;(e) cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby exposing the next amino acid residue as a N-terminal amino acid residue on the cleaved peptide and providing an immobilized location complex, the immobilized location complex comprising the cleaved and separated N-terminal amino acid residue and the location nucleic acid;(f) providing a second chemically-reactive conjugate comprising: (xx) a cycle tag comprising a cycle nucleic acid associated with a cycle number, and (yy) a second reactive moiety for binding and cleaving the N-terminal amino acid residue of the peptide and exposing a next amino acid residue as an additional N-terminal amino acid residue on the cleaved peptide, and optionally (zz) a moiety for joining to an optional second solid support;(g) contacting the peptide with the second chemically-reactive conjugate, thereby coupling the second chemically-reactive conjugate to the additional N-terminal amino acid of the peptide to form a coupled-conjugate-complex;(h) forming one or more transfer complexes, each transfer complex comprising an immobilized location complex and a coupled-conjugate-complex, thereby bringing a location nucleic acid into proximity with a cycle nucleic acid within each formed transfer complex;(i) within each formed transfer complex, joining a location nucleic acid or a reverse complement thereof to a cycle nucleic acid to form an immobilized recode unit, or otherwise joining information of the location nucleic acid and the cycle nucleic acid, thereby creating one or a plurality of immobilized recode units, each recode unit corresponding with a formed transfer complex;(j) cleaving and thereby separating the additional N-terminal amino acid residue from the peptide, thereby exposing the next amino acid residue as a further N-terminal amino acid residue on the cleaved peptide and liberating the immobilized recode unit complex from the solid support, the liberated recode unit complex comprising the cleaved and separated additional N-terminal amino acid residue, cycle and location information;Attorney Docket No: 062954-509001 WO(k) repeating (f) through (j) n-1 times to liberate and collect pools having a plurality of recode unit complexes, each additional plurality of recode unit complexes comprising information associated with cycle 2 to n, accordingly;(l) contacting the collection of recode units complexes with binding agents, the binding agents comprising: a binding moiety for preferentially binding to one, or to a subset, of the recode unit complexes, and a recode tag comprising a recode nucleic acid corresponding with the binding agent, thereby forming one or more affinity complexes, each affinity complex comprising an recode unit complex and the binding agent, thereby bringing a recode nucleic acid into proximity with a recode tag within each formed affinity complex;(m) within each formed affinity complex, joining a recode unit or a reverse complement thereof to a recode tag to form a recode block, or otherwise transferring information of the recode tag to the recode unit complex, thereby creating a plurality of recode blocks, each recode block corresponding with a formed affinity complex;(n) optionally, joining two or more members of the plurality of recode blocks to form a memory oligonucleotide;(o) obtaining sequence information for the recode blocks or memory oligonucleotides; and(p) based on the obtained sequence information, determining identity and positional information of a plurality of amino acid residues of the peptide.

26. The method of claim 25, wherein (f)-(j) are repeated 2, 3, 4, or more times.

27. The method of claim 25, wherein n is an integer greater than or equal to 2.

28. The method of claim 25, wherein each binding agent comprises recode tags with a unique nucleic acid sequence.

29. The method of claim 25, wherein a plurality of binding agents comprises recode tags with the same nucleic acid sequence.

30. The method of claim 25, wherein the binding agents comprises recode tags which have a unique sequence portion and a common sequence portion.

31. The method of claim 25, wherein the binding agents are contacted with combined pool(s) of recode unit complexes.

32. The method of claim 25, further comprising washing the immobilized amino acid complex before said contacting the immobilized amino acid complex with a coupled-conjugate-complex.

33. The method of claim 25, further comprising washing the coupled-conjugate-complex before said contacting the immobilized amino acid complex with a coupled-conjugate-complex.

34. The method of claim 25, further comprising determining a likely three-dimensional structure of the peptide based on the sequence information.Attorney Docket No: 062954-509001 WO35. The method of claim 25, wherein the recode nucleic acid comprises a hybridization-capable code.

36. The method of claim 25, wherein the recode nucleic acid comprises DNA.

37. The method of claim 25, wherein the recode nucleic acid comprises non-colliding codes in an additive vector space.

38. The method of claim 25, wherein the location nucleic acid comprises DNA.

39. The method of claim 25, wherein the cycle nucleic acid comprises DNA.

40. The method of claim 25, wherein the peptide comprises a hormone, neurotransmitter, enzyme, antibody, viral protein, bacterial protein, synthetic peptide, bioactive peptide, peptide hormone, oligopeptide, polypeptide, fusion protein, cyclic peptide, branched peptide, recombinant protein, tumor marker, therapeutic peptide, antigenic peptide, or signaling peptide.

41. The method of claim 25, wherein a plurality of peptides are immobilized to the solid support and analyzed in concert.

42. The method of claim 25, wherein obtaining the sequence information for the memory oligonucleotide comprises performing sequencing.

43. The method of claim 25, wherein obtaining the sequence information for the memory oligonucleotide comprises melt curve analysis for multi-stage encoding.

44. The method of claim 25, wherein the binding moiety comprises an antibody or a fragment thereof, or an aptamer.

45. The method of claim 25, wherein the binding moiety is covalently bound or non-covalently bound to the amino acid.

46. The method of claim 25, wherein the binding moiety binds to a natural amino acid, a derivatized amino acid, a synthetic amino acid, or a D-amino acid.

47. The method of claim 25, wherein the binding moiety binds to a natural amino acid, a post- translationally modified amino acid, a derivatized version of an amino acid, a derivatized or stabilized version of a post-translationally modified amino acid, a synthetic amino acid, an amino acid with a specific side chain, an amino acid with a phosphorylated side chain, an amino acid with a glycosylated side chain, an amino acid with a methylation modification, or a D-amino acid, phenylthiohydantoin (PTH) derivative or anilinothiazolinone (ATZ) derivative of an amino acid, or binds to a combination thereof.

48. The method of claim 25, wherein the solid support comprises a bead, a plate, or a chip.

49. The method of claim 25, wherein the solid support comprises glass slide, silica, a resin, a gel, a hydrogel, a membrane, polystyrene, a metal, nitrocellulose, a mineral, plastic, polyacrylamide, latex, or ceramic.Attorney Docket No: 062954-509001 WO50. The method of claim 25, further comprising protecting or deprotecting the location tag at any appropriate step of the workflow between (b) and (k).

51. The method of claim 25, wherein said transferring information comprises performing nucleic acid amplification, enzymatic ligation, splint ligation, chemical ligation, template-assisted ligation, use of a ligase enzyme, use of a splint oligonucleotide, use of a catalyst, use of a bridging molecule, use of a condensation agent, use of a coupling reagent, use of a polymerase enzyme, use of a complementary nucleic acid sequence, use of a nicking enzyme, use of a nucleic acid modifying enzyme, use of a recombinase, use of a strand-displacing polymerase, use of a single-strand binding protein, a click chemistry reaction, a phosphodiester bond formation, or a peptide nucleic acid-mediated ligation.

52. The method of claim 25, wherein the information of the recode nucleic acid comprises a sequence of the recode nucleic acid or a reverse complement of the sequence of the recode nucleic acid.

53. The method of claim 25, wherein said transferring information comprises joining the recode nucleic acid or a reverse complement of the recode nucleic acid with the cycle nucleic acid.

54. The method of claim 25, wherein determining the identity and positional information of the plurality of amino acid residues of the peptide comprises determining the identity and positional information of all of the amino acid residues of the peptide.

55. The method of claim 25, wherein determining the identity and positional information of the plurality of amino acid residues of the peptide comprises determining the identity and positional information of only a subset of the amino acid residues of the peptide.

56. The method of claim 25, further comprising identifying the peptide by comparing the identity and positional information of the plurality of amino acid residues to a database.

57. A kit for determining identity and positional information of an amino acid residue of a peptide, comprising: a chemically-reactive conjugate comprising (a) a nucleic acid sequence tag and (b) a reactive moiety that couples to a N-terminal amino acid residue of a peptide, and thereby forms a conjugate complex comprising the chemically-reactive conjugate coupled to the N-terminal amino acid of the peptide; a binding agent comprising: a binding moiety for preferentially binding to the conjugate complex and a recode tag comprising a recode nucleic acid corresponding with the binding agent; and a reagent for transferring information of the recode nucleic acid to the cycle nucleic acid of the conjugate complex to generate a recode block.

58. A chemically-reactive conjugate (CRC) comprising or consisting of: (A) a nucleic acid sequence tag; and (B) a reactive moiety for binding and cleaving a N-terminal amino acid residue from a peptide or a cleavable derivative thereofAttorney Docket No: 062954-509001 WO59. A CRC represented by Formula III:LABA B(Formula III), wherein A comprises a cycle tag, and B comprises a reactive moiety, and LAB comprises a linker with option central moiety (C).

60. The CRC of claim 58, comprising a cleavable group between (A) and (B), between (B) and (C), between (A) and (C), between (A) and (B+C), between (B) and (A+C), or between (C) and (A+B), or any combination thereof.

61. The CRC of claim 59, comprising a cleavable group between (A) and (B), between (B) and (C), between (A) and (C), between (A) and (B+C), between (B) and (A+C), or between (C) and (A+B), or any combination thereof.

62. The CRC of claim 58, comprising a cleavable group between (A) and (B), between (B) and (C), or a combination thereof.

63. The CRC of claim 58, wherein the reactive moiety comprises a phenyl isothiocyanate (PITC), an isothiocyanate (ITC), a dansyl chloride, a dinitrofluorobenzene (DNFB), an enzyme or peptide, or a combination or derivative thereof.

64. The CRC of claim 58, wherein the reactive moiety specifically cleaves at a specific amino acid.

65. The CRC of claim 58, wherein the reactive moiety cleaves more than a single amino acid or motif.

66. The CRC of claim 58, wherein the immobilizing moiety comprises a protected thiol group, a protected amine group, or a carboxyl group, an azide, an alkyne, an alkene, an aryl boronic acid, an aryl halide, a haloalkyne, an acryl, a silylalkyne, a Si-H group, a protected or photoprotected reactive group, or a photoactivated reactive group.

67. The CRC of claim 58, wherein the nucleic acid sequence tag is generated upon conjugating the nucleic acid sequence to a group for attaching a nucleic acid sequence comprising a protected oxyamine group, a protected thiol, a protected amine, a protected hydrazine, a tetrazine, an azide, an alkyne, an alkene, a trans-cyclooctene, a DBCO, a bicyclononyne, a norbornene, a strained alkyne, or a strained alkene, or a derivative thereof.

68. The CRC of claim 58, wherein the reactive moiety comprises a group on the CRC for attaching to a cleavable derivatized N-terminal amino acid, comprising a tetrazine, an azide, an alkene, an alkyne, aAttorney Docket No: 062954-509001 WO trans-cyclooctene, a DBCO, a bicyclononyne, a norbomene, a strained alkyne, or a strained alkene, or a derivative thereof.

69. A method of producing a surface that presents spatially discrete molecular zones, the method comprising: a. attaching first and second anchor oligonucleotides to a solid support; b. extending a subset of the first anchor oligonucleotides with a template oligonucleotide that comprises a universal-hybridization site (UHS) sequence, a unique-location identifier (ULI) sequence, and a first primer sequence, thereby defining analyte-binding locations; c. performing bridge amplification using the first primer sequence and a second primer sequence on the second anchor oligonucleotides, whereby multiple copies of the ULI are generated around each analyte-binding location to form a discrete molecular zone; and d. modifying the analyte-binding location to permit covalent attachment of an analyte.

70. The method of claim 69 wherein the template of step (b) comprises iG bases, iC bases, neither iC nor iG, or both iC and iG.

71. The method of claim 69 wherein the anchor oligonucleotides do not comprise adenine or guanine bases.

72. The method of claim 69 wherein the bridge-amplification reaction solution does not comprise iso-dCTP.

73. The method of claim 69 wherein the oligonucleotide of step (b) is supplied at <0.5 mol % of the total extension-reaction oligonucleotides.

74. The method of claim 69 wherein the modification introduced in step (d) comprises incorporation of an alkyne-modified deoxycytidine.

75. A method of producing a surface that presents spatially discrete molecular zones, the method comprising: a. attaching template oligonucleotides that comprise a universal-hybridization site (UHS) sequence, a unique-location identifier (ULI) sequence, and a first primer sequence, and thereby define analyte-binding locations, and template oligonucleotides that comprises a second primer sequence capable to support bridge amplification in combination with the first primer sequence; b. performing bridge amplification of the template oligonucleotides, whereby multiple copies of the ULI are generated around each analyte-binding location to form a discrete molecular zone; c. modifying the analyte-binding location to permit covalent attachment of an analyte.

76. A surface comprising a plurality of discrete molecular zones, each zone containing a. a single analyte-binding site; andAttorney Docket No: 062954-509001 WO b. two or more copies of a zone-specific ULI oligonucleotide, the ULI oligonucleotides of different zones being sequence-distinguishable from one another.

77. The surface of any one of claims 69, 75 or 76, wherein the ULI oligonucleotides in each zone are bounded by an iCiC motif that is unpaired in the absence of iso-dGTP.

78. The surface of any one of claims 69, 75 or 76, wherein each zone further comprises upstream and downstream zone interaction pairing (ZIP) codes that flank the ULI sequence, at least 80 % of adjacent zones differing in their ZIP-code pair.

79. The surface of claim 76, wherein the analyte-binding site comprises an alkyne group suitable for copper-catalyzed azide-alkyne cyclo-addition.

80. A method for determining adjacency relationships between discrete molecular zones of clonal ULIs, the method comprising: a. copying, in each zone, one or more ZIP-flanked oligonucleotides to generate a strand that terminates in a restriction site (RS) and is blocked at a modified-nucleotide motif; b. annealing and ligating a bridge oligonucleotide whose termini are complementary to unlike ZIP codes on neighbouring zones, thereby forming a junction-spanning strand that contains the ULIs of two adjacent zones; c. amplifying the junction-spanning strand with primers that bind to adapter sequences distal to the ZIP codes to generate a junction amplicon; d. removing the amplified junction amplicon from the surface; and e. sequencing the junction amplicon and recording the two ULIs therein as adjacent nodes in a spatial graph.

81. The method of claim 80, wherein the ZIP codes are 4-7 nucleotides in length and comprise at least one LNA base.

82. The method of claim 80, wherein the bridge oligonucleotide carries a 5' phosphate and internal abasic spacers that suppress non-specific hybridization to conserved primer regions.

83. The method of claim 80, wherein the amplification of step (c) is carried out with primers complementary to CT1 and CT2 adapter sequences.

84. The method of claim 80, further comprising, after step (d), digesting the residual surface-bound strand at the RS site with a Type-IIS restriction endonuclease to regenerate the primer site for a subsequent mapping cycle.

85. The method of claim 80, wherein the spatial graph is laid out with a spring-embedding algorithm in which edge weights are proportional to junction-amplicon read counts.

86. The method of claim 80, wherein machine-learning inference is applied to the spatial graph to refine node positions using a training set of reference layouts.Attorney Docket No: 062954-509001 WO87. An oligonucleotide composition comprising N distinct oligonucleotide species, a. wherein each species comprises: i. a first region containing a subset of k-mers selected from a pool of N+l distinct k-mer sequences, wherein exactly one k-mer from said pool is absent from said first region, and ii. a second region comprising the reverse complement of said absent k-mer, b. wherein any two different oligonucleotide species in said composition are capable of forming a heteroduplex via complementary k-mer pairing, and wherein no oligonucleotide species is capable of forming a homoduplex with itself.

88. A method, comprising: removing an N-terminal amino acid from a protein in the presence of tris(pentafluorophenyl)borane (BCF).

89. The method of claim 88, wherein removing the N-terminal amino acid from the protein is performed in the absence of trifluoroacetic acid (TFA).

90. The method of claim 88, wherein removing the N-terminal amino acid from the protein comprises contacting the protein with a chemically reactive conjugate (CRC) comprising an Edman degradation reagent.

91. The method of claim 90, wherein the CRC comprises a nucleic acid tag.

92. The method of claim 91, wherein the nucleic acid tag comprises deaza-nucleotides.

93. The method of claim 91, wherein the nucleic acid tag maintains its integrity.

94. The method of claim 91, further comprising sequencing the nucleic acid tag.

Citation Information

Patent Citations

  • Proximity interaction analysis

    US20210254047A1

  • Determination of protein information by recoding amino acid polymers into DNA polymers

    US20240044909A1

  • Polypeptide capture, in situ fragmentation and identification

    US20240133892A1