Systems and methods for application of artificial intelligence to proteomics data

WO2026170196A1PCT designated stage Publication Date: 2026-08-13ESCALANTE BIOSCIENCES INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-02-10
Publication Date
2026-08-13

Smart Images

  • Figure US2026014745_13082026_PF_FP_ABST
    Figure US2026014745_13082026_PF_FP_ABST
Patent Text Reader

Abstract

Systems, methods, compositions, and kits for identifying protein analytes are provided. Analytes are obtained and contacted with binders to form binder-analyte interactions. Binders include unique barcode sequences and optionally interact nonspecifically with analytes. For a first analyte, a DNA concatemer including barcode sequences is generated, where each barcode sequence corresponds to a binder that forms a binder-analyte interaction with the first analyte and is appended to the first analyte using a DNA polymerase responsive to the interaction. A spatial order of barcode sequences corresponds to a temporal order of formation of binder-analyte interactions between binders and the first analyte. An interaction data structure indicating presence or abundance of binders forming binder-analyte interactions with the first analyte is determined from sequencing the DNA concatemer. Responsive to inputting the interaction data structure to a model, an indication of an identity for the first analyte is received.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.: 139643-5001-WOSYSTEMS AND METHODS FOR APPLICATION OF ARTIFICIAL INTELLIGENCE TO PROTEOMICS DATACROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 756,656, filed February 10, 2025, which is hereby incorporated by reference in its entirety.TECHNICAL FIELD

[0002] This application is directed to identifying analytes, in particular using nucleic acid barcode concatemers that record interactions between binders and analytes.BACKGROUND

[0003] Proteins within biological contexts such as cells work together to perform a wide range of functions through protein-protein interactions (PPIs), regulatory interactions, posttranslational modifications, and subcellular localization. Proteomic technologies allow for the analysis of the roles of proteins within such biological contexts, including biological processes, cellular compartments, and metabolic or signaling pathways. Gaining a complete understanding of this intricate proteomic landscape requires the accurate and specific identification and characterization of proteins in their native physiological environments.

[0004] Traditionally, this has involved obtaining and testing large libraries of compounds to find a small number of compounds that interact with a protein analyte of interest. However, the cost, time, and resources needed to physically assay and characterize each of these compounds is prohibitive to testing them at scale. Moreover, protein datasets are conventionally driven by manual lab processes, which are highly laborious and unsuitable for performing large numbers of complex workflows involving a wide variety of biological contexts.

[0005] Given the above background, what is needed in the art are improved methods for characterizing protein analytes within biological contexts.DB2 / 651687418.2 1Attorney Docket No.: 139643-5001-WOSUMMARY

[0006] As described above, there is a need in the art for improved methods for characterizing protein analytes within biological contexts. In particular, there is a need in the art for scalable, protein-focused models that represent and provide insight into biological states within biological samples (e.g., cells).

[0007] Analysis of nucleic acids alone provides only limited insight into biological phenotypes, particularly for protein analyte expression and production. For instance, genomic or transcriptomic disruption is impractical for numerous protein classes (e.g., developmentally essential genes) and may not faithfully reflect the effects of pharmacological disruption or other perturbations in mature organisms. Genomic and transcriptomic systems, furthermore, do not fully recapitulate the diverse ways in which protein function may be modulated, including partial blockade or activation, tissue-specific targeting of protein analytes, and / or polypharmacology, in which a chemical compound targets multiple protein targets. In contrast, molecular models such as AlphaFold provide highly accurate predictions of protein structure but often lack biological context at the omics level, limiting their biological and therapeutic interpretability.

[0008] The present disclosure addresses the problems identified herein by providing methods, systems, and / or compositions that facilitate characterization and / or identification of one or more proteins in a proteome of a biological sample. In some instances, the methods, systems, and / or compositions disclosed herein provide for the determination of a phenotype of a biological sample under a target condition. In some embodiments, a phenotype includes, but is not limited to, expression or production of protein analytes, conformation of protein analytes, generation of protein-protein complexes, binding affinity, and / or behavior of a protein analyte responsive to exposure to a chemical compound (e.g., “druggability”). In some embodiments, a target condition includes, but is not limited to, a healthy condition, a diseased condition, exposure to chemical compounds, overexpression or overproduction of biological intermediates, and / or repression or silencing of biological intermediates. In some aspects, the methods, systems, and / or compositions disclosed herein provide for the determination of protein analytes with target binding properties, for instance, by identifying protein analytes comprising binding moi eties that bind to a class of binders and / or a class of tags. For instance, protein analytes with specific binding moieties include different proteinDB2 / 651687418.2 2Attorney Docket No.: 139643-5001-WOclasses with specific epitopes that can be preferentially targeted with different small molecules.

[0009] Advantageously, the presently disclosed methods, systems, and / or compositions improve the technical field of proteomics and drug discovery by facilitating high-precision identification and characterization of protein analytes present in biological samples in a manner that preserves biological context. In some implementations, the identification and characterization of protein analytes allow for the elucidation of phenotypic effects in response to target conditions (e.g., small molecule or pharmacological perturbation), as well as the determination of protein analytes that can be targeted by specific binders. This knowledge can then be used to prioritize new targets and / or new drug-target pairs for clinical development.

[0010] One aspect of the present disclosure provides a method for identifying protein analytes, comprising obtaining a plurality of protein analytes from a sample of a subject and contacting the plurality of protein analytes with a plurality of molecular binders and a DNA polymerase, thereby forming a plurality of binder-analyte interactions. In some embodiments, each respective molecular binder in the plurality of molecular binders comprises a corresponding nucleic acid barcode sequence that is unique to the respective molecular binder, and one or more molecular binders in the plurality of molecular binders interacts nonspecifically with each protein analyte in a corresponding subset of the plurality of protein analytes. In some embodiments, the method further includes generating, for a first protein analyte in the plurality of protein analytes, a DNA concatemer comprising a corresponding set of DNA barcode sequences. In some embodiments, each DNA barcode sequence in the corresponding set of DNA barcode sequences i) corresponds to a nucleic acid barcode sequence of a respective molecular binder, in the plurality of molecular binders, that forms a binder-analyte interaction with the first protein analyte, and ii) is appended to the first protein analyte using the DNA polymerase, responsive to formation of the binder-analyte interaction between the respective molecular binder and the first protein analyte, thereby causing a spatial order of DNA barcode sequences in the corresponding set of DNA barcode sequences to correspond to a temporal order of formation of binder-analyte interactions between respective molecular binders, in the plurality of molecular binders, and the first protein analyte. In some embodiments, the method further includes determining an interaction data structure (e.g., a vector, matrix, and / or tensor) for the first protein analyte based on a sequencing of the DNA concatemer, where the interaction data structure comprises an DB2 / 651687418.2 3Attorney Docket No.: 139643-5001-WOindication of a presence or abundance of each respective molecular binder in the plurality of molecular binders that formed a binder-analyte interaction with the first protein analyte. In some embodiments, the method further includes receiving, responsive to inputting the interaction data structure to a model, an indication of an identity for the first protein analyte, as output from the model.

[0011] Another aspect of the present disclosure provides a method for detecting protein analytes, comprising obtaining a plurality of protein analytes from a sample of a subject. In some embodiments, the plurality of protein analytes comprises a first class of protein analytes and a second class of protein analytes, and each protein analyte in the first class of protein analytes is attached to a first initiator tag comprising a first DNA initiator sequence. In some embodiments, the method further includes contacting the plurality of protein analytes with a plurality of molecular binders and a DNA polymerase, thereby forming a plurality of binderanalyte interactions. In some embodiments, each protein analyte in the first class of protein analytes forms a binder-analyte interaction with at least one molecular binder in the plurality of molecular binders, each respective molecular binder in the plurality of molecular binders comprises a corresponding nucleic acid barcode sequence that is unique to the respective molecular binder, and one or more molecular binders in the plurality of molecular binders interacts nonspecifically with each protein analyte in a corresponding subset of the first class of protein analytes. In some embodiments, the method further includes generating, for a first protein analyte in the first class of protein analytes, a DNA concatemer comprising a corresponding set of DNA barcode sequences, where each DNA barcode sequence in the corresponding set of DNA barcode sequences i) corresponds to a nucleic acid barcode sequence of a respective molecular binder, in the plurality of molecular binders, that forms a binder-analyte interaction with the first protein analyte, and ii) is appended to the first DNA initiator sequence of the first protein analyte using the DNA polymerase, responsive to formation of the binder-analyte interaction between the respective molecular binder and the first protein analyte, thereby causing a spatial order of DNA barcode sequences in the corresponding set of DNA barcode sequences to correspond to a temporal order of formation of binder-analyte interactions between respective molecular binders, in the plurality of molecular binders, and the first protein analyte. In some embodiments, the method further includes determining an interaction data structure for the first protein analyte based on a sequencing of the DNA concatemer, where the interaction data structure comprises an indication of a presence or abundance of each respective molecular binder in the plurality ofDB2 / 651687418.2 4Attorney Docket No.: 139643-5001-WOmolecular binders that formed a binder-analyte interaction with the first protein analyte. In some embodiments, the method further includes receiving, responsive to inputting the interaction data structure to a model, an indication of an identity for the first protein analyte, as output from the model.

[0012] Yet another aspect of the present disclosure includes a system, including a memory; one or more processors; and one or more modules stored in the memory and configured for execution by the one or more processors, the one or more modules including instructions for performing any of the methods disclosed above.

[0013] Still another aspect of the present disclosure includes a non-transitory computer readable storage medium, the non-transitory computer readable storage medium storing one or more programs for execution by one or more processors of a computer system, the one or more computer programs including instructions for performing any of the methods disclosed above.

[0014] Another aspect of the present disclosure provides compositions and kits for identifying and / or detecting protein analytes.BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In the drawings, embodiments of the methods, systems, and / or compositions of the present disclosure are illustrated by way of example. It is to be expressly understood that the description and drawings are only for the purpose of illustration and as an aid to understanding, and are not intended as a definition of the limits of the methods, systems, and / or compositions of the present disclosure.

[0016] FIGS. 1A, IB, and 1C collectively illustrate a computer system in accordance with some embodiments of the present disclosure.

[0017] FIGS. 2A, 2B, 2C, 2D, 2E, and 2F collectively illustrate an example workflow for identifying protein analytes, in which optional steps are indicated by dashed lines, in accordance with some embodiments of the present disclosure.

[0018] FIGS. 3 A and 3B collectively illustrate an example workflow for detecting protein analytes, in which optional steps are indicated by dashed lines, in accordance with some embodiments of the present disclosure.DB2 / 651687418.2 5Attorney Docket No.: 139643-5001-WO

[0019] FIGS. 4A, 4B, and 4C collectively illustrate an example schematic for identifying and / or detecting protein analytes, in accordance with some embodiments of the present disclosure.

[0020] FIGS. 5A and 5B collectively illustrate an example workflow for identifying protein analytes, in which optional steps are indicated by dashed lines, in accordance with some embodiments of the present disclosure.

[0021] FIGS. 6A and 6B collectively illustrate an example schematic for identifying and / or detecting protein analytes, in accordance with some embodiments of the present disclosure.

[0022] FIG. 7 illustrates an example schematic for a nucleic acid barcode molecule, in accordance with some embodiments of the present disclosure.

[0023] FIG. 8 illustrates example concatemers generated from appending a plurality of nucleic acid barcodes, in accordance with an embodiment of the present disclosure.

[0024] FIG. 9 A illustrates example plots showing read length, number of barcodes, and distance between constant regions for concatemers generated from a positive control oligonucleotide, in accordance with an embodiment of the present disclosure.

[0025] FIG. 9B illustrates example plots showing read length, number of barcodes, and distance between constant regions for concatemers generated from a pooled library of nucleic acid barcodes, in accordance with an embodiment of the present disclosure.

[0026] FIG. 10 illustrates an example plot showing a correspondence between number of barcode observations in a concatemer and delta-G predictions of biomolecular proximity between barcodes and an initiator tag, in accordance with an embodiment of the present disclosure.

[0027] FIG. 11 illustrates example outputs from a method for obtaining protein analytes comprising initiator tags and molecular binders comprising nucleic acid barcodes, in accordance with an embodiment of the present disclosure.

[0028] FIG. 12 illustrates example concatemers generated from appending a plurality of nucleic acid barcodes to initiator tags using an extension reaction, in accordance with an embodiment of the present disclosure.

[0029] FIG. 13 illustrates example PCR products for concatemers generated from appending a plurality of nucleic acid barcodes to initiator tags, in accordance with an embodiment of the present disclosure.DB2 / 651687418.2 6Attorney Docket No.: 139643-5001-WO

[0030] FIGS. 14A and 14B collectively illustrate example plots showing read length, number of barcodes, and distance between constant regions for sequencing data obtained from the extension reactions illustrated in FIGS. 12 and 13, in accordance with an embodiment of the present disclosure.

[0031] FIGS. 15A and 15B collectively illustrate an example extension reaction performed using initiator-tagged analytes and nucleic acid barcoded molecular binders (circles and pseudo-lariats), in accordance with an embodiment of the present disclosure.

[0032] FIGS. 16A and 16B illustrate example products formed from an extension reaction performed using initiator-tagged analytes and nucleic acid barcoded molecular binders (circles and pseudo-lariats) illustrated in FIGS. 15A-B, in accordance with an embodiment of the present disclosure.

[0033] FIG. 17 illustrates example plots showing read length, number of barcodes, and distance between constant regions for sequencing data obtained from the extension reactions illustrated in FIGS. 15A-B and 16A-B (circles), in accordance with an embodiment of the present disclosure.

[0034] FIGS. 18A and 18B illustrate example plots showing read length, number of barcodes, and distance between constant regions for sequencing data obtained from the extension reactions illustrated in FIGS. 15A-B and 16A-B (pseudo-lariats), in accordance with an embodiment of the present disclosure.

[0035] FIG. 19 illustrates example sequence reads generated from a concatemer reaction demonstrating 0-3 barcode extensions on initiator-tagged protein analytes using the extension reactions illustrated in FIGS. 15A-B and 16A-B (pseudo-lariats), in accordance with an embodiment of the present disclosure.

[0036] Like reference numerals refer to corresponding parts throughout the several views of the drawings.DETAILED DESCRIPTION

[0037] Reference will now be made in detail to embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to one of ordinary skill in the art that the presentDB2 / 651687418.2 7Attorney Docket No.: 139643-5001-WOdisclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0038] Given the above background, there is a need in the art for improved methods for characterizing protein analytes within biological contexts. In particular, there is a need in the art for scalable, protein-focused models that represent and provide insight into biological states within biological samples such as cells.

[0039] Analysis of nucleic acids alone provides only limited insight into biological phenotypes, particularly for protein analyte expression and production. For instance, genomic or transcriptomic disruption is impractical for numerous protein classes (e.g., developmentally essential genes) and may not faithfully reflect the effects of pharmacological disruption or other perturbations in mature organisms. Genomic and transcriptomic systems, furthermore, do not fully recapitulate the diverse ways in which protein function may be modulated, including partial blockade or activation, tissue-specific targeting of protein analytes, and / or polypharmacology, in which a chemical compound targets multiple protein targets. In contrast, molecular models such as AlphaFold provide highly accurate predictions of protein structure but often lack biological context at the omics level, limiting their biological and therapeutic interpretability.

[0040] The present disclosure addresses the problems identified herein by providing methods, systems, and / or compositions that facilitate characterization and / or identification of one or more proteins in a proteome of a biological sample. In some instances, the methods, systems, and / or compositions disclosed herein provide for the determination of a phenotype of a biological sample under a target condition. In some embodiments, a phenotype includes, but is not limited to, expression or production of protein analytes, conformation of protein analytes, generation of protein-protein complexes, binding affinity, and / or behavior of a protein analyte responsive to exposure to a chemical compound (e.g., “druggability”). In some embodiments, a target condition includes, but is not limited to, a healthy condition, a diseased condition, exposure to chemical compounds, overexpression or overproduction of biological intermediates, and / or repression or silencing of biological intermediates. In some aspects, the methods, systems, and / or compositions disclosed herein provide for the determination of protein analytes with target binding properties, for instance, by identifying protein analytes comprising binding moieties that bind to a class of binders. For instance,DB2 / 651687418.2 8Attorney Docket No.: 139643-5001-WOprotein analytes with specific binding moieties include different protein classes with specific epitopes that can be preferentially targeted with different small molecules.

[0041] Advantageously, the presently disclosed methods, systems, and / or compositions improve the technical field of proteomics and drug discovery by facilitating high-precision identification and characterization of protein analytes present in biological samples in a manner that preserves biological context. In some implementations, the identification and characterization of protein analytes allow for the elucidation of phenotypic effects in response to target conditions (e.g., small molecule or pharmacological perturbation), as well as the determination of protein analytes that can be targeted by specific binders. This knowledge can then be used to prioritize new targets and / or new drug-target pairs for clinical development.

[0042] Moreover, the use of automation and machine learning can overcome human limitations in proteomics and drug discovery. For instance, manual analysis is often slow and inefficient. As such, the limitations of manual experimentation can impede the identification and selection of target protein analytes. Conversely, an automated reaction and sequencing platform leverages recent increases in computational power to enable standardized big data that can lead to improved models and identification of new molecules and / or binders for drug discovery. In some embodiments, the use of machine learning models and / or automated reaction and sequencing devices, such as an automated robot, improves the technical field of proteomics and drug discovery.

[0043] The present disclosure addresses the above problems in the art by providing methods, systems, compositions, and kits for identifying protein analytes. In one aspect of the present disclosure, analytes are obtained and contacted with binders to form binder-analyte interactions. Binders include unique barcode sequences and, in some embodiments, interact nonspecifically with analytes. In some embodiments, each molecular binder in the plurality of molecular binders that contacts the analytes interact nonspecifically with the analytes. In some embodiments, each respective molecular binder in a subset of the plurality of molecular binders (e.g., one or more molecular binders) interacts nonspecifically with the analytes. For a first analyte, a DNA concatemer including barcode sequences is generated, where each barcode sequence corresponds to a binder that forms a binder-analyte interaction with the first analyte and is appended to the first analyte using a DNA polymerase responsive to the interaction. In some implementations, the DNA concatemer includes a spatial order of barcode sequences that corresponds to a temporal order of formation of binder-analyte DB2 / 651687418.2 9Attorney Docket No.: 139643-5001-WOinteractions between binders and the first analyte. An interaction data structure indicating presence or abundance of binders forming binder-analyte interactions with the first analyte is determined from sequencing the DNA concatemer. Responsive to inputting the interaction data structure to a model, an indication of an identity for the first analyte is received.

[0044] Definitions.

[0045] It will be understood that, although the terms first, second, etc., may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first subject could be termed a second subject, and, similarly, a second subject could be termed a first subject, without departing from the scope of the present disclosure. The first subject and the second subject are both subjects, but they are not the same subject.

[0046] The terminology used in the present disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used in the description of the invention and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0047] As used herein, the term “if’ may be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” depending on the context. Similarly, the phrase “if it is determined” or “if [a stated condition or event] is detected” may be construed to mean “upon determining” or “in response to determining” or “upon detecting [the stated condition or event]” or “in response to detecting [the stated condition or event],” depending on the context.

[0048] As used interchangeably herein, the terms “protein” and “protein analyte” refer to a polymeric molecule comprising amino acids. In some embodiments, a protein analyte is a biological object that is capable of interacting with a molecular binder. In some embodiments, a protein analyte is a polypeptide. In some embodiments, a protein analyte is a large molecule composed of repeating residues. In some embodiments, a protein analyte is a DB2 / 651687418.2 10Attorney Docket No.: 139643-5001-WOplurality of polymers (e.g., 2 or more, 3, or more, 10 or more, 100 or more, 1000 or more, or 5000 or more polymers), where the respective polymers in the plurality of polymers do not all have the same molecular weight. In some embodiments, the protein analyte includes two or more polypeptides bound to each other. As used herein, the term “polypeptide” means two or more amino acids or residues linked by a peptide bond.

[0049] In some embodiments, the protein analyte includes any number of posttranslational modifications. Thus, in some embodiments, a protein analyte includes those polymers that are modified by acylation, alkylation, amidation, biotinylation, formylation, y-carboxylation, glutamyl ati on, glycosylation, glycylation, hydroxylation, iodination, isoprenylation, lipoylation, cofactor addition (for example, of a heme, flavin, metal, efc.), addition of nucleosides and their derivatives, oxidation, reduction, pegylation, phosphatidylinositol addition, phosphopantetheinylation, phosphorylation, pyroglutamate formation, racemization, addition of amino acids by tRNA (for example, arginylation), sulfation, selenoylation, ISGylation, SUMOylation, ubiquitination, chemical modifications (for example, citrullination and deamidation), and treatment with other enzymes (for example, proteases, phosphatases and kinases). Other types of posttranslational modifications are known in the art and are within the scope of the protein analytes of the present disclosure.

[0050] As used herein, the term “target” refers to an object of interest, such as a protein or protein analyte that is of interest as a primary binding target for a molecular binder. As used herein, the term “off-target” refers to an object that is not the primary binding target, such as a protein or protein analyte that exhibits off-target binding with a molecular binder.

[0051] As used herein, the term “model” refers to a machine learning model or algorithm.

[0052] In some embodiments, a model is an unsupervised learning algorithm. One example of an unsupervised learning algorithm is cluster analysis.

[0053] In some embodiments, a model is a supervised machine learning algorithm.Nonlimiting examples of supervised learning algorithms include, but are not limited to, logistic regression, neural networks, support vector machines, Naive Bayes algorithms, nearest neighbor algorithms, random forest algorithms, decision tree algorithms, boosted trees algorithms, multinomial logistic regression algorithms, linear models, linear regression, GradientBoosting, mixture models, hidden Markov models, Gaussian NB algorithms, linear discriminant analysis, or any combinations thereof. In some embodiments, a model is a multinomial classifier algorithm. In some embodiments, a model is a 2-stage stochastic DB2 / 651687418.2 11Attorney Docket No.: 139643-5001-WOgradient descent (SGD) model. In some embodiments, a model is a deep neural network (e.g., a deep-and-wide sample-level classifier).

[0054] Neural networks. In some embodiments, the model is a neural network (e.g., a convolutional neural network and / or a residual neural network). Neural network algorithms, also known as artificial neural networks (ANNs), include convolutional and / or residual neural network algorithms (deep learning algorithms). Neural networks can be machine learning algorithms that may be trained to map an input data set to an output data set, where the neural network comprises an interconnected group of nodes organized into multiple layers of nodes. For example, the neural network architecture may comprise at least an input layer, one or more hidden layers, and an output layer. The neural network may comprise any total number of layers, and any number of hidden layers, where the hidden layers function as trainable feature extractors that allow mapping of a set of input data to an output value or set of output values. As used herein, a deep learning algorithm can be a neural network comprising a plurality of hidden layers, e.g., two or more hidden layers. Each layer of the neural network can comprise a number of nodes (or “neurons”). A node can receive input that comes either directly from the input data or the output of nodes in previous layers, and perform a specific operation, e.g., a summation operation. In some embodiments, a connection from an input to a node is associated with a parameter (e.g., a weight and / or weighting factor). In some embodiments, the node may sum up the products of all pairs of inputs, xi, and their associated parameters. In some embodiments, the weighted sum is offset with a bias, b. In some embodiments, the output of a node or neuron may be gated using a threshold or activation function, f, which may be a linear or non-linear function. The activation function may be, for example, a rectified linear unit (ReLU) activation function, a Leaky ReLU activation function, or other function such as a saturating hyperbolic tangent, identity, binary step, logistic, arcTan, softsign, parametric rectified linear unit, exponential linear unit, softPlus, bent identity, softExponential, Sinusoid, Sine, Gaussian, or sigmoid function, or any combination thereof.

[0055] The weighting factors, bias values, and threshold values, or other computational parameters of the neural network, may be “taught” or “learned” in a training phase using one or more sets of training data. For example, the parameters may be trained using the input data from a training data set and a gradient descent or backward propagation method so that the output value(s) that the ANN computes are consistent with the examples included in theDB2 / 651687418.2 12Attorney Docket No.: 139643-5001-WOtraining data set. The parameters may be obtained from a back propagation neural network training process.

[0056] Any of a variety of neural networks may be suitable for use in analyzing an image of an eye of a subject. Examples can include, but are not limited to, feedforward neural networks, radial basis function networks, recurrent neural networks, residual neural networks, convolutional neural networks, residual convolutional neural networks, and the like, or any combination thereof. In some embodiments, the machine learning makes use of a pre-trained and / or transfer-learned ANN or deep learning architecture. Convolutional and / or residual neural networks can be used for analyzing an image of a subject in accordance with the present disclosure.

[0057] For instance, a deep neural network model comprises an input layer, a plurality of individually parameterized (e.g., weighted) convolutional layers, and an output scorer. The parameters (e.g., weights) of each of the convolutional layers as well as the input layer contribute to the plurality of parameters (e.g, weights) associated with the deep neural network model. In some embodiments, at least 100 parameters, at least 1000 parameters, at least 2000 parameters or at least 5000 parameters are associated with the deep neural network model. As such, deep neural network models require a computer to be used because they cannot be mentally solved. In other words, given an input to the model, the model output needs to be determined using a computer rather than mentally in such embodiments. See, for example, Krizhevsky etal., 2012, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems 2, Pereira, Burges, Bottou, Weinberger, eds., pp. 1097-1105, Curran Associates, Inc.; Zeiler, 2012 “ADADELTA: an adaptive learning rate method,” CoRR, vol. abs / 1212.5701; and Rumelhart etal., 1988, “Neurocomputing: Foundations of research,” ch. Learning Representations by Back-propagating Errors, pp. 696-699, Cambridge, MA, USA: MIT Press, each of which is hereby incorporated by reference.

[0058] Neural network algorithms, including convolutional neural network algorithms, suitable for use as models are disclosed in, for example, Vincent et al., 2010, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion,” J Mach Learn Res 11, pp. 3371-3408; Larochelle et al., 2009, “Exploring strategies for training deep neural networks,” J Mach Learn Res 10, pp. 1-40; and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of which is hereby incorporated by reference. Additional example neural DB2 / 651687418.2 13Attorney Docket No.: 139643-5001-WOnetworks suitable for use as models are disclosed in Duda et al., 2001, Pattern Classification, Second Edition, John Wiley & Sons, Inc., New York; and Hastie etal., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, each of which is hereby incorporated by reference in its entirety. Additional example neural networks suitable for use as models are also described in Draghici, 2003, Data Analysis Tools for DNA Microarrays, Chapman & Hall / CRC; and Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, each of which is hereby incorporated by reference in its entirety.

[0059] Support vector machines. In some embodiments, the model is a support vector machine (SVM). SVM algorithms suitable for use as models are described in, for example, Cristianini and Shawe-Taylor, 2000, “An Introduction to Support Vector Machines,” Cambridge University Press, Cambridge; Boser et al., 1992, “A training algorithm for optimal margin classifiers,” in Proceedings of the 5th Annual ACM Workshop on Computational Learning Theory, ACM Press, Pittsburgh, Pa., pp. 142-152; Vapnik, 1998, Statistical Learning Theory, Wiley, New York; Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y.; Duda, Pattern Classification, Second Edition, 2001, John Wiley & Sons, Inc., pp. 259, 262-265; and Hastie, 2001, The Elements of Statistical Learning, Springer, New York; and Furey et al., 2000, Bioinformatics 16, 906-914, each of which is hereby incorporated by reference in its entirety. When used for classification, SVMs separate a given set of binary labeled data with a hyper-plane that is maximally distant from the labeled data. For cases in which no linear separation is possible, SVMs can work in combination with the technique of 'kernels', which automatically realizes a non-linear mapping to a feature space. The hyper-plane found by the SVM in feature space can correspond to a non-linear decision boundary in the input space. In some embodiments, the plurality of parameters (e.g., weights) associated with the SVM define the hyper-plane. In some embodiments, the hyper-plane is defined by at least 10, at least 20, at least 50, or at least 100 parameters and the SVM model requires a computer to calculate because it cannot be mentally solved.

[0060] Naive Bayes algorithms. In some embodiments, the model is a Naive Bayes algorithm. Naive Bayes models suitable for use as models are disclosed, for example, in Ng et al., 2002, “On discriminative vs. generative classifiers: A comparison of logistic regression and naive Bayes,” Advances in Neural Information Processing Systems, 14, which is hereby incorporated by reference. A Naive Bayes model is any model in a family of “probabilistic DB2 / 651687418.2 14Attorney Docket No.: 139643-5001-WOmodels” based on applying Bayes’ theorem with strong (naive) independence assumptions between the features. In some embodiments, they are coupled with Kernel density estimation. See, for example, Hastie et al, 2001, The elements of statistical learning: data mining, inference, and prediction, eds. Tibshirani and Friedman, Springer, New York, which is hereby incorporated by reference.

[0061] Nearest neighbor algorithms. In some embodiments, a model is a nearest neighbor algorithm. Nearest neighbor models can be memory-based and include no model to be fit. For nearest neighbors, given a query point xo (a test subject), the k training points X(r), r, ..., k (here the training subjects) closest in distance to xo are identified and then the point xo is classified using the k nearest neighbors. Here, the distance to these neighbors is a function of the abundance values of the discriminating gene set. In some embodiments, Euclidean distance in feature space is used to determine distance asTypically, when the nearest neighbor algorithm is used, the abundance data used to compute the linear discriminant is standardized to have mean zero and variance 1. The nearest neighbor rule can be refined to address issues of unequal class priors, differential misclassification costs, and feature selection. Many of these refinements involve some form of weighted voting for the neighbors. For more information on nearest neighbor analysis, see Duda, Pattern Classification, Second Edition, 2001, John Wiley & Sons, Inc; and Hastie, 2001, The Elements of Statistical Learning, Springer, New York, each of which is hereby incorporated by reference.

[0062] A k-nearest neighbor model is a non-parametric machine learning method in which the input consists of the k closest training examples in feature space. The output is a class membership. An object is classified by a plurality vote of its neighbors, with the object being assigned to the class most common among its k nearest neighbors (k is a positive integer, typically small). If k = 1, then the object is simply assigned to the class of that single nearest neighbor. See, Duda el al., 2001, Pattern Classification, Second Edition, John Wiley & Sons, which is hereby incorporated by reference. In some embodiments, the number of distance calculations needed to solve the k-nearest neighbor model is such that a computer is used to solve the model for a given input because it cannot be mentally performed.

[0063] Random forest, decision tree, and boosted tree algorithms. In some embodiments, the model is a decision tree. Decision trees suitable for use as models are described generally by Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 395-396, which is hereby incorporated by reference. Tree-based methods partition the feature space into a set DB2 / 651687418.2 15Attorney Docket No.: 139643-5001-WOof rectangles, and then fit a model (like a constant) in each one. In some embodiments, the decision tree is random forest regression. One specific algorithm that can be used is a classification and regression tree (CART). Other specific decision tree algorithms include, but are not limited to, ID3, C4.5, MART, and Random Forests. CART, ID3, and C4.5 are described in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 396-408 and pp. 411-412, which is hereby incorporated by reference. CART, MART, and C4.5 are described in Hastie eta!.. 2001, The Elements of Statistical Learning, Springer-Verlag, New York, Chapter 9, which is hereby incorporated by reference in its entirety.Random Forests are described in Breiman, 1999, “Random Forests— Random Features,” Technical Report 567, Statistics Department, U.C. Berkeley, September 1999, which is hereby incorporated by reference in its entirety. In some embodiments, the decision tree model includes at least 10, at least 20, at least 50, or at least 100 parameters (e.g., weights and / or decisions) and requires a computer to calculate because it cannot be mentally solved.

[0064] Regression. In some embodiments, the model uses a regression algorithm. A regression algorithm can be any type of regression. For example, in some embodiments, the regression algorithm is logistic regression. In some embodiments, the regression algorithm is logistic regression with lasso, L2 or elastic net regularization. In some embodiments, those extracted features that have a corresponding regression coefficient that fails to satisfy a threshold value are pruned (removed from) consideration. In some embodiments, a generalization of the logistic regression model that handles multicategory responses is used as the model. Logistic regression algorithms are disclosed in Agresti, An Introduction to Categorical Data Analysis, 1996, Chapter 5, pp. 103-144, John Wiley & Son, New York, which is hereby incorporated by reference. In some embodiments, the model makes use of a regression model disclosed in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York. In some embodiments, the logistic regression model includes at least 10, at least 20, at least 50, at least 100, or at least 1000 parameters (e.g., weights) and requires a computer to calculate because it cannot be mentally solved.

[0065] Linear discriminant analysis algorithms. Linear discriminant analysis (LDA), normal discriminant analysis (ND A), or discriminant function analysis can be a generalization of Fisher’s linear discriminant, a method used in statistics, pattern recognition, and machine learning to find a linear combination of features that characterizes or separates two or more classes of objects or events. The resulting combination can be used as the model (linear model) in some embodiments of the present disclosure.DB2 / 651687418.2 16Attorney Docket No.: 139643-5001-WO

[0066] Mixture model and Hidden Markov model. In some embodiments, the model is a mixture model, such as that described in McLachlan etal., Bioinformatics 18(3):413-422, 2002. In some embodiments, in particular, those embodiments including a temporal component, the model is a hidden Markov model such as described by Schliep etal., 2003, Bioinformatics 19(1): i255-i263.

[0067] Clustering. In some embodiments, the model is an unsupervised clustering model. In some embodiments, the model is a supervised clustering model. Clustering algorithms suitable for use as models are described, for example, at pages 211-256 of Duda and Hart, Pattern Classification and Scene Analysis, 1973, John Wiley & Sons, Inc., New York, (hereinafter "Duda 1973") which is hereby incorporated by reference in its entirety. The clustering problem can be described as one of finding natural groupings in a dataset. To identify natural groupings, two issues can be addressed. First, a way to measure similarity (or dissimilarity) between two samples can be determined. This metric (e.g., similarity measure) can be used to ensure that the samples in one cluster are more like one another than they are to samples in other clusters. Second, a mechanism for partitioning the data into clusters using the similarity measure can be determined. One way to begin a clustering investigation can be to define a distance function and to compute the matrix of distances between all pairs of samples in the training set. If distance is a good measure of similarity, then the distance between reference entities in the same cluster can be significantly less than the distance between the reference entities in different clusters. However, clustering may not use a distance metric. For example, a nonmetric similarity function s(x, x') can be used to compare two vectors x and x'. s(x, x') can be a symmetric function whose value is large when x and x' are somehow “similar.” Once a method for measuring “similarity” or “dissimilarity” between points in a dataset has been selected, clustering can use a criterion function that measures the clustering quality of any partition of the data. Partitions of the data set that extremize the criterion function can be used to cluster the data. Particular exemplary clustering techniques that can be used in the present disclosure can include, but are not limited to, hierarchical clustering (agglomerative clustering using a nearest-neighbor algorithm, farthest-neighbor algorithm, the average linkage algorithm, the centroid algorithm, or the sum-of-squares algorithm), k-means clustering, fuzzy k-means clustering algorithm, and Jarvis-Patrick clustering. In some embodiments, the clustering comprises unsupervised clustering (e.g., with no preconceived number of clusters and / or no predetermination of cluster assignments).DB2 / 651687418.2 17Attorney Docket No.: 139643-5001-WO

[0068] Reinforcement learning. In some embodiments, a model is a reinforcement learning model. In some embodiments, the reinforcement learning system comprises four main elements - an agent, a policy, a reward signal, and a value function, where the behavior of the agent is defined in terms of the policy. In some embodiments, the reinforcement learning system comprises a learning algorithm. In some implementations, the learning algorithm is an on-policy learning algorithm or an off-policy learning algorithm. On-Policy learning algorithms evaluate and improve the same policy which is being used to select the agent’s actions. Off-Policy learning algorithms evaluate and improve policies that are different from the policy being used for action selection. Reinforcement learning models contemplated for use in the present disclosure are further described, for example, in Ibrahim et al., “Comprehensive Overview of Reward Engineering and Shaping in Advancing Reinforcement Learning Applications,” IEEE Access, July 22, 2024, doi: 10.48550 / arXiv.2408.10215, which is hereby incorporated herein by reference in its entirety.

[0069] Ensembles of models and boosting. In some embodiments, an ensemble (two or more) of models is used. In some embodiments, a boosting technique such as AdaBoost is used in conjunction with many other types of learning algorithms to improve the performance of the model. In this approach, the output of any of the models disclosed herein, or their equivalents, is combined into a weighted sum that represents the final output of the boosted model. In some embodiments, the plurality of outputs from the models is combined using any measure of central tendency known in the art, including but not limited to a mean, median, mode, a weighted mean, weighted median, weighted mode, etc. In some embodiments, the plurality of outputs is combined using a voting method. In some embodiments, a respective model in the ensemble of models is weighted or unweighted.

[0070] In some embodiments, the model is a reinforcement learning model. In some embodiments, the reinforcement learning system comprises four main elements - an agent, a policy, a reward signal, and a value function, where the behavior of the agent is defined in terms of the policy. In some embodiments, the reinforcement learning system comprises a learning algorithm. In some implementations, the learning algorithm is an on-policy learning algorithm or an off-policy learning algorithm. On-Policy learning algorithms evaluate and improve the same policy which is being used to select the agent’s actions. Off-Policy learning algorithms evaluate and improve policies that are different from the policy being used for action selection. Reinforcement learning is further described, for example, in Sutton RS, Barto AG, “Reinforcement learning: an introduction,” IEEE Transactions on Neural DB2 / 651687418.2 18Attorney Docket No.: 139643-5001-WONetworks. 1998;9(5): 1054-1054, which is hereby incorporated herein by reference in its entirety. In some embodiments, the reinforcement learning model includes at least 10, at least 100, at least 1000, at least 10,000, at least 100,000, at least 1 x 106, at least 1 x 107, or more parameters. In some embodiments, the reinforcement learning model includes no more than 1 x 108, no more than 1 x 107, no more than 1 x 106, no more than 100,000, no more than 10,000, no more than 1000, or no more than 100 parameters. In some embodiments, the reinforcement learning model consists of from 10 to 1000, from 100 to 100,000, from 10,000 to 1 x 107, or from 1 x 106to 1 x 108parameters. In some embodiments, the plurality of parameters for the reinforcement learning model falls within another range starting no lower than 10 parameters and ending no higher than 1 x 108parameters.

[0071] As used herein, the term “parameter” refers to any coefficient or, similarly, any value of an internal or external element (e.g., a weight and / or a hyperparameter) in an algorithm, model, regressor, and / or classifier that affects (e.g., modify, tailor, and / or adjust) one or more inputs, outputs, and / or functions in the algorithm, model, regressor and / or classifier. For example, in some embodiments, a parameter refers to any coefficient, weight, and / or hyperparameter that is used to control, modify, tailor, and / or adjust the behavior, learning and / or performance of an algorithm, model, regressor, and / or classifier. In some instances, a parameter is used to increase or decrease the influence of an input (e.g., a feature) to an algorithm, model, regressor, and / or classifier. As a nonlimiting example, in some instances, a parameter is used to increase or decrease the influence of a node (e.g., of a neural network), where the node includes one or more activation functions. Assignment of parameters to specific inputs, outputs, and / or functions is not limited to any one paradigm for a given algorithm, model, regressor, and / or classifier but can be used in any suitable an algorithm, model, regressor, and / or classifier architecture for a desired performance. In some embodiments, a parameter has a fixed value. In some embodiments, a value of a parameter is manually and / or automatically adjustable. In some embodiments, a value of a parameter is modified by a validation and / or training process for an algorithm, model, regressor, and / or classifier (e.g., by error minimization and / or backpropagation methods, as described elsewhere herein).

[0072] In some embodiments, an algorithm, model, regressor, and / or classifier of the present disclosure comprises a plurality of parameters. In some embodiments the plurality of parameters is n parameters, where: n > 2; n > 5; n > 10; n > 25; n > 40; n > 50; n > 75; n > 100; n > 125; n > 150; n > 200; n > 225; n > 250; n > 350; n > 500; n > 600; n > 750; n > DB2 / 651687418.2 19Attorney Docket No.: 139643-5001-WO1,000; n > 2,000; n > 4,000; n > 5,000; n > 7,500; n > 10,000; n > 20,000; n > 40,000; n > 75,000; n > 100,000; n > 200,000; n > 500,000, n > 1 x 106, n > 5 x 106, or n > 1 x 107In some embodiments n is between 10,000 and 1 x 107, between 100,000 and 5 x 106, or between 500,000 and 1 x 106.

[0073] As used herein, the term “instruction” refers to an order given to a computer processor by a computer program. On a digital computer, in some embodiments, each instruction is a sequence of 0s and Is that describes a physical operation the computer is to perform. Such instructions can include data transfer instructions and data manipulation instructions. In some embodiments, each instruction is a type of instruction in an instruction set that is recognized by a particular processor type used to carry out the instructions. Examples of instruction sets include, but are not limited to, Reduced Instruction Set Computer (RISC), Complex Instruction Set Computer (CISC), Minimal Instruction Set Computers (MISC), Very Long Instruction Word (VLIW), Explicitly Parallel Instruction Computing (EPIC), and One Instruction Set Computer (OISC).

[0074] As used herein, the terms “nonspecific” or “nonspecifically” (alternatively, “specific,” “specifically,” or “specificity”), in the context of an interaction between two or more entities, refers the degree to which a first entity interacts with (e.g., binds to) a particular second entity and not to other entities. In some embodiments, an entity refers to a biological molecule capable of forming interactions with another entity, including but not limited to polymers, peptides, receptors, small molecules, nucleic acids, proteins, ligands, binders, analytes, protein forms (e.g., proteoforms), and the like. A first entity that forms an interaction with a second entity includes, but is not limited to, a drug and its target receptor, a first protein and a second protein, a protein and a ligand, a molecular binder and an analyte, among others. For instance, in some embodiments, a nonspecific interaction refers to the degree to which a molecular binder interacts with a particular analyte and not to other analytes. In some embodiments, protein analytes participate in specific interactions with just one or a few partners (e.g., molecular binders). In some embodiments, protein analytes participate in nonspecific interactions with numerous different partners. Thus, in some embodiments, interaction specificity refers to a preference of a first entity (e.g., a molecular binder or an analyte) for interactions with a small set of partners over multiple possibilities. See, for example, Eaton BE, Gold L, Zichi DA. Let’s get specific: the relationship between specificity and affinity. Chem Biol. 1995;2(10):633-638, which is hereby incorporated herein by reference in its entirety.DB2 / 651687418.2 20Attorney Docket No.: 139643-5001-WO

[0075] In some implementations, the specificity or nonspecificity of an interaction between a first entity and a second entity is determined based on an interaction affinity between the first entity and the second entity, such as a binding affinity. Binding affinity, as used herein, refers to the strength of interaction between a first entity and a second entity (e.g., a drug and its target receptor, a first protein and a second protein, a protein and a ligand, a molecular binder and an analyte, etc.). In some embodiments, the binding affinity of an entity is an equilibrium dissociation constant, KD, a measure of how tightly a first entity binds to a second entity (e.g., a ligand binding to its target, a molecular binder binding to an analyte, etc.). A smaller KD value indicates a stronger binding affinity. In some embodiments, an interaction between a first entity and a second entity is deemed to be specific or nonspecific when the interaction satisfies an affinity threshold. In some embodiments, an interaction between a first entity and a second entity e.g., between a molecular binder and an analyte) is deemed to be nonspecific when the interaction between the first entity and the second entity is a “low affinity” or “weak affinity” interaction, that is, less than the affinity threshold. In some embodiments, an interaction between a first entity and a second entity (e.g., between a molecular binder and an analyte) is deemed to be specific when the interaction between the first entity and the second entity is a “high affinity” or “strong affinity” interaction, that is, greater than the affinity threshold. In some embodiments, the affinity threshold is a KD of at least 0.01 nM, at least 0.05 nM, at least 0.1 nM, at least 0.5 nM, at least 1 nM, at least 5 nM, at least 10 nM, at least 50 nM, at least 0.1 pM, at least 0.5 pM, at least 1 pM, at least 5 pM, at least 10 pM, at least 50 pM, at least 100 pM, or at least 500 pM. In some embodiments, the affinity threshold is a KD of no more than 1 mM, no more than 500 pM, no more than 200 pM, no more than 100 pM, no more than 50 pM, no more than 10 pM, no more than 1 pM, no more than 0.1 pM, no more than 10 nM, no more than 1 nM, or no more than 0.1 nM. In some embodiments, the affinity threshold is a KD of from 0.01 nM to 0.1 nM, from 0.1 nM to 1 nM, from 1 nM to 0.1 pM, from 0.1 pM to 1 pM, from 1 pM to 50 pM, from 10 pM to 200 pM, from 100 pM to 500 pM, or from 500 pM to 1 mM. In some embodiments, the affinity threshold is a KD falling within another range starting no lower than 0.01 nM and ending no higher than 1 mM. In some cases, however, binding affinity alone does not determine whether a first entity interacts specifically or nonspecifically with a second entity. For instance, and without being limited to any one theory of operation, functionally important binding can occur at a range of affinities from low millimolar to femtomolar. See, for example, Schreiber G, Keating AE. Protein binding specificity versus promiscuity. Current opinion in structural biology.2010;21(l):50, which is hereby incorporated herein by reference in its entirety.DB2 / 651687418.2 21Attorney Docket No.: 139643-5001-WO

[0076] In some implementations, the specificity or nonspecificity of an interaction between a first entity and a second entity is determined according to the types of interaction forces involved. Generally, interactions between entities comprise covalent interactions (e.g., full or partial covalent bonds) and / or noncovalent interactions, including but not limited to hydrogen bonds, electrostatics, van der Waals interactions, and / or hydrophobic effects. Without being limited to any one theory of operation, the presence of covalent interactions allows for greater affinity between entities than the presence of noncovalent interactions alone. For example, studies have shown that small molecule affinity for protein binding sites resulting from noncovalent interactions generally peaks at 10 picomolar (IO11M). See, for example, Smith AJT, Zhang X, Leach AG, Houk KN. Beyond picomolar affinities: quantitative aspects of noncovalent and covalent binding of drugs to proteins. Journal of medicinal chemistry.2009;52(2):225, which is hereby incorporated herein by reference in its entirety. In some implementations, an interaction between a first entity and a second entity is deemed nonspecific when there is an absence of covalent interactions (e.g., an absence of covalent bonds) involved in the entity-entity interaction. The presence of covalent interactions, however, does not necessarily indicate strong affinity, but can also be present in instances of weak affinity. In some embodiments, an interaction between a first entity and a second entity is deemed nonspecific even in the presence of covalent interactions.

[0077] In some implementations, the specificity or nonspecificity of an interaction between a first entity and a second entity is determined according to the concentration and proximity of each entity in a plurality of entities. For instance, in some embodiments, a measure for biologically relevant binding includes that two proteins must be localized near one another at a concentration that promotes interaction. In some such embodiments, specificity is a relative trait that is context dependent. Similarly, in some embodiments, interaction specificity is determined based on a comparison between the strength of an interaction between a first entity and a second entity relative to the strength of an interaction between the first entity and another entity other than the second entity (e.g., designating a protein as “specific” if interaction with a desired partner is tighter than with other proteins). In some such embodiments, interaction specificity for a particular pair of entities is determined relative to other pairs of entities. In some embodiments, a first entity is deemed to interact nonspecifically with a second entity if the strength of the interaction between the first entity and the second entity is lower than or equal to another interaction between the first entity and a third entity other than the second entity. See, for example, Schreiber G, Keating AE. ProteinDB2 / 651687418.2 22Attorney Docket No.: 139643-5001-WObinding specificity versus promiscuity. Current opinion in structural biology. 2010;21(l):50, which is hereby incorporated herein by reference in its entirety.

[0078] In some implementations, the specificity or nonspecificity of an entity is determined according to the number of interactions that the entity forms with other interaction partners. In some embodiments, an entity interacts nonspecifically when it is capable of interacting with (e.g., binding with) a plurality of interaction partners. In some embodiments, an entity interacts nonspecifically when it is capable of interacting with (e.g., binding with) each interaction partner in a subset of the plurality of interaction partners. In some embodiments, each respective entity in a plurality of entities interacts nonspecifically when it is capable of interacting with (e.g., binding with) each respective interaction partner in a corresponding subplurality of interaction partners in the plurality of interaction partners. For instance, in some embodiments, a molecular binder interacts nonspecifically when it is capable of forming a binder-analyte interaction with each protein analyte in a plurality of protein analytes, and / or with each different protein form e.g., proteoforms, variants, structures, and / or modifications thereof) in a plurality of different protein forms. In some embodiments, a molecular binder interacts nonspecifically when it is capable of forming a binder-analyte interaction with each protein analyte in a subset of the plurality of protein analytes, and / or with each different protein form (e.g, proteoforms, variants, structures, and / or modifications thereof) in a subset of the plurality of different protein forms. In some embodiments, each respective molecular binder in a plurality of molecular binders interacts nonspecifically when it is capable of forming a binder-analyte interaction with each respective protein analyte in a corresponding subplurality of protein analytes in the plurality of protein analytes, and / or with each different protein form (e.g., proteoforms, variants, structures, and / or modifications thereof) in a corresponding subplurality of different protein forms in the plurality of different protein forms.

[0079] In some embodiments, an entity is deemed to interact nonspecifically when it is capable of interacting with at least 1, at least 2, at least 5, at least 10, at least 20, at least 30, at least 50, at least 80, at least 100, at least 200, or at least 300 interaction partners. In some embodiments, an entity is deemed to interact nonspecifically when it is capable of interacting with no more than 500, no more than 300, no more than 100, no more than 50, no more than 30, no more than 20, no more than 10, or no more than 5 interaction partners. In some embodiments, an entity is deemed to interact nonspecifically when it is capable of interacting with from 1 to 10, from 5 to 40, from 20 to 100, from 50 to 300, or from 100 to 500DB2 / 651687418.2 23Attorney Docket No.: 139643-5001-WOinteraction partners. In some embodiments, an entity is deemed to interact nonspecifically when it is capable of interacting with another range of interaction partners falling no lower than 1 interaction partner and ending no higher than 500 interaction partners.

[0080] Suitable methods for determining interaction specificity or nonspecificity contemplated for use in the present disclosure are further described, for instance, in Schreiber G, Keating AE. Protein binding specificity versus promiscuity. Current opinion in structural biology. 2010;21(l):50; Smith AJT, Zhang X, Leach AG, Houk KN. Beyond picomolar affinities: quantitative aspects of noncovalent and covalent binding of drugs to proteins. Journal of medicinal chemistry. 2009;52(2):225; Eaton BE, Gold L, Zichi DA. Let’s get specific: the relationship between specificity and affinity. Chem Biol. 1995;2(10):633-638; and Bruns RF, Watson IA. Rules for identifying potentially reactive or promiscuous compounds. J Med Chem. 2012;55(22):9763-9772, each of which is hereby incorporated herein by reference in its entirety. However, the present disclosure is not limited thereto.

[0081] Example Systems for Identifying Protein Analytes

[0082] FIGS. 1 A-C collectively illustrate a computer system 100 (e.g., for identifying protein analytes).

[0083] Referring to FIGS. 1 A-C, in some embodiments, computer system 100 comprises one or more computers. For purposes of illustration in FIGS. 1A-C, the computer system 100 is represented as a single computer that includes all of the functionality of the disclosed computer system 100. However, the present disclosure is not so limited. The functionality of the computer system 100 can be spread across any number of networked computers and / or reside on each of several networked computers and / or virtual machines. One of skill in the art will appreciate that a wide array of different computer topologies is possible for the computer system 100 and all such topologies are within the scope of the present disclosure.

[0084] The computer system 100 comprises one or more processing units (CPUs) 59, a network or other communications interface 84, a user interface 78 (e.g., including an optional display 82 and optional keyboard 80 or other form of input device), a memory 92 (e.g., random access memory, persistent memory, or combination thereof), one or more magnetic disk storage and / or persistent devices 90 optionally accessed by one or more controllers 88, one or more communication busses 12 for interconnecting the aforementioned components, and a power supply 79 for powering the aforementioned components. To the extent that components of memory 92 are not persistent, data in memory 92 can be seamlessly shared DB2 / 651687418.2 24Attorney Docket No.: 139643-5001-WOwith non-volatile memory 90 or portions of memory 92 that are non-volatile / persistent using known computing techniques such as caching. Memory 92 and / or memory 90 can include mass storage that is remotely located with respect to the central processing unit(s) 59. In other words, some data stored in memory 92 and / or memory 90 may in fact be hosted on computers that are external to computer system 100 but that can be electronically accessed by the computer system 100 over an Internet, intranet, or other form of network or electronic cable using network interface 84. In some embodiments, the computer system 100 makes use of models that are run from the memory associated with one or more graphical processing units in order to improve the speed and performance of the system. In some alternative embodiments, the computer system 100 makes use of models that are run from memory 92 rather than memory associated with a graphical processing unit.

[0085] In some embodiments, the memory 92 of the computer system 100 stores:• an optional operating system 34 that includes procedures for handling various basic system services;• an optional network communication module 118 for connecting the system 100 with other devices, or a communication network;• a barcode module 120, optionally comprising, for each respective molecular binder 122 (e.g., 122-1,... 122-L), a corresponding nucleic acid barcode sequence 124 (e.g., 124-1) that is unique to the respective molecular binder 122;• a reaction construct 130, optionally comprising:o instructions for contacting a plurality of protein analytes 142 (e.g., 142- 1,... 142-K) with the plurality of molecular binders 122, where one or more molecular binders in the plurality of molecular binders 122 interacts nonspecifically with each protein analyte 142 in a corresponding subset of the plurality of protein analytes (e.g., forms binder-analyte interactions with each protein analyte in the subset of protein analytes), ando instructions for generating, for at least a first protein analyte 142-1, a DNA concatemer 144 (e.g., 144-1) comprising a corresponding set of DNA barcode sequences 146 (e.g., 146-1,... 146-L);• a sequencing construct 140, optionally comprising instructions for sequencing the DNA concatemer 144;• an interaction module 150, optionally comprising, for at least the first protein analyte 142-1, an interaction data structure 152 (e.g., 152-1) based on a sequencing of theDB2 / 651687418.2 25Attorney Docket No.: 139643-5001-WODNA concatemer 144, comprising an indication of a presence or abundance 154 (e.g., 154-1,... 154-L) of each respective molecular binder 122 in the plurality of molecular binders that formed a binder-analyte interaction with the first protein analyte 142-1; and• a prediction module 160, optionally for receiving, responsive to inputting the interaction data structure 152 to a model, an indication of an identity 162 (e.g., 162-1) for at least the first protein analyte 142-1.

[0086] Optionally, in some embodiments, system 100 comprises one or more of: a barcode module 120, optionally comprising, for each respective protein analyte 142 in a plurality of protein analytes, a corresponding nucleic acid barcode sequence 124 e.g., 124-1) that is unique to the respective protein analyte 142; a reaction construct 130, optionally comprising: instructions for contacting the plurality of protein analytes 142 with a plurality of molecular binders 122, where one or more molecular binders in the plurality of molecular binders 122 interacts nonspecifically with each protein analyte 142 in a corresponding subset of the plurality of protein analytes (e.g., forms binder-analyte interactions with each protein analyte in the subset of protein analytes), and instructions for generating, for at least a first molecular binder 122, a DNA concatemer 144 (e.g., 144-1) comprising a corresponding set of DNA barcode sequences 146 (e.g., 146-1,... 146-L); a sequencing construct 140, optionally comprising instructions for sequencing the DNA concatemer 144; an interaction module 150, optionally comprising, for at least the first molecular binder 122-1, an interaction data structure 152 (e.g., 152-1) based on a sequencing of the DNA concatemer 144, comprising an indication of a presence or abundance 154 (e.g., 154-1,... 154-L) of each respective protein analyte 142 in the plurality of protein analytes that formed a binder-analyte interaction with the first molecular binder 122-1; and a prediction module 160, optionally for receiving, responsive to inputting the interaction data structure 152 to a model, an indication of an identity 162 (e.g., 162-1) for at least a first protein analyte 142-1.

[0087] In some implementations, one or more of the barcode module, the reaction construct, and / or the sequencing construct are performed using an automated reaction device. In some implementations, one or more of the barcode module, the reaction construct, and / or the sequencing construct are performed manually.

[0088] In some implementations, one or more of the above identified data elements or modules of the computer system 100 are stored in one or more of the previously mentioned memory devices, and correspond to a set of instructions for performing a function described DB2 / 651687418.2 26Attorney Docket No.: 139643-5001-WOabove. The above identified data, modules, or programs (e.g., sets of instructions) need not be implemented as separate software programs, procedures or modules, and thus various subsets of these modules may be combined or otherwise re-arranged in various implementations. In some implementations, the memory 92 and / or 90 (and optionally 52) optionally stores a subset of the modules and data structures identified above. Furthermore, in some embodiments the memory 92 and / or 90 (and optionally 52) stores additional modules and data structures not described above.

[0089] Now that a system 100 has been disclosed, methods for performing such methods are detailed with reference to FIGS. 2A-F, FIGS. 3A-B, and FIGS. 4A-C.

[0090] Example Methods for Identifying Protein Analytes

[0091] FIGS. 2A-F collectively illustrate a method 200 for identifying protein analytes 142. Referring to Block 202, in some embodiments, methods include obtaining a plurality of protein analytes 142 from a sample of a subject.

[0092] Samples.

[0093] In some embodiments, the sample of the subject comprises one or more cells.Referring to Block 204, in some embodiments, the obtaining of the plurality of protein analytes comprises generating a cell lysate from the sample of the subject.

[0094] In some embodiments, the sample of the subject is a solid tissue sample or a liquid biopsy sample. In some embodiments, the sample of the subject comprises blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid.

[0095] As used interchangeably herein, the terms “biological sample” and “sample” refer to any sample taken from a subject, which can reflect a biological state associated with the subject. Examples of samples include, but are not limited to, blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid of the subject. In some embodiments, the sample consists of blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid of the subject. In some embodiments, the sample is a stool sample. In some embodiments, a sample includes any tissue or material derived from a living or dead subject. In some embodiments, a sample is a cell-free sample.DB2 / 651687418.2 27Attorney Docket No.: 139643-5001-WO

[0096] In some embodiments, as illustrated in FIG. 4A, a sample is treated to physically disrupt tissue or cell structure (e.g., centrifugation and / or cell lysis), thus releasing intracellular components into a solution, such as a lysate. In some embodiments, a sample further contains enzymes, buffers, salts, detergents, and the like which can be used to prepare the sample for analysis. In some embodiments, obtaining the plurality of protein analytes comprises isolating the protein analytes from the sample. In some embodiments, the plurality of protein analytes is not isolated from the sample.

[0097] In some embodiments, the sample of the subject comprises a target condition. In some embodiments, the target condition is a healthy condition or a disease condition. In some embodiments, the target condition is exposure to a chemical compound. In some embodiments, chemical compounds contemplated for use in the present disclosure include, but are not limited to, drugs, hormones, antibodies, enzymes, nucleic acids, and / or small molecules. For example, in some embodiments, a chemical compound contemplated for use in the present disclosure includes a small molecule inhibitor. In some embodiments, a chemical compound contemplated for use in the present disclosure is a proteolysis targeting chimera (PROTAC). PROTACs suitable for use in the present disclosure are further described, for example, in Liu et al., 2023, “An overview of PROTACs: a promising drug discovery paradigm,” Molecular Biomedicine. 2022; 3:46; and Ribes etal., 2024, “Modeling protac degradation activity with machine learning,” Artificial Intelligence in the Life Sciences 6:100104; each of which is incorporated herein by reference in its entirety.

[0098] In some embodiments, the target condition is a dysregulation of a biological intermediate in the sample of the subject. In some embodiments, the biological intermediate is an enzyme, a kinase, a signaling molecule, a transcription factor, a hormone, a small molecule, a microRNA, a regulatory RNA, or a cytokine, and the dysregulation comprises overexpression, overproduction, repression, or silencing. As used herein, a biological intermediate refers to a biological molecule or substance that directly or indirectly affects expression or production of protein analytes, such as within a biological pathway, including but not limited to enzymes, kinases, signaling molecules, transcription factors (e.g., proteins that bind to nucleic acids to regulate gene transcription), hormones (e.g., signaling molecules that regulate gene expression), small molecules (e.g., drugs or metabolites that bind to specific receptors to influence gene activity), microRNAs (e.g., small RNA molecules that inhibit translation of mRNA), regulatory RNAs (e.g., non-coding RNAs that regulate geneDB2 / 651687418.2 28Attorney Docket No.: 139643-5001-WOexpression at different levels), and / or cytokines (e.g., proteins that signal between cells to modulate gene expression).

[0099] Protein analytes.

[0100] Referring to Block 206, in some embodiments, methods further include contacting the plurality of protein analytes 142 with a plurality of molecular binders 122 and a DNA polymerase, thereby forming a plurality of binder-analyte interactions. In some embodiments, each respective molecular binder 122 in the plurality of molecular binders comprises a corresponding nucleic acid barcode sequence 124 that is unique to the respective molecular binder 122. In some embodiments, one or more molecular binders in the plurality of molecular binders 122 interacts nonspecifically with each protein analyte 142 in a corresponding subset of the plurality of protein analytes (e.g., forms binder-analyte interactions with each protein analyte in the subset of protein analytes).

[0101] Referring to Block 208, in some embodiments, a respective protein analyte in the plurality of protein analytes is an enzyme, a kinase, a protease, an antibody, or a receptor. In some embodiments, the protein analyte is any of the proteins or protein analytes disclosed elsewhere herein (see, for example, the section entitled “Definitions: Protein,” above).

[0102] Referring to Block 210, in some embodiments, a respective protein analyte in the plurality of protein analytes is a protein-protein complex. In some embodiments, the protein-protein complex comprises at least 2, at least 3, at least 4, at least 5, at least 8, at least 10, or at least 15 proteins. In some embodiments, the protein-protein complex comprises no more than 20, no more than 15, no more than 10, no more than 5, or no more than 3 proteins. In some embodiments, the protein-protein complex consists of from 2 to 4, from 2 to 8, from 3 to 7, from 5 to 10, or from 8 to 20 proteins. In some embodiments, the protein-protein complex falls within another range starting no lower than 2 proteins and ending no higher than 20 proteins.

[0103] In some embodiments, the protein-protein complex comprises at least 2, at least 3, at least 4, at least 5, at least 8, at least 10, or at least 15 different proteins. In some embodiments, the protein-protein complex comprises no more than 20, no more than 15, no more than 10, no more than 5, or no more than 3 different proteins. In some embodiments, the protein-protein complex consists of from 2 to 4, from 2 to 8, from 3 to 7, from 5 to 10, or from 8 to 20 different proteins. In some embodiments, the protein-protein complex fallsDB2 / 651687418.2 29Attorney Docket No.: 139643-5001-WOwithin another range starting no lower than 2 different proteins and ending no higher than 20 different proteins.

[0104] Referring to Block 212, in some embodiments, a respective protein analyte in the plurality of protein analytes comprises a post-translational modification. In some embodiments, the post-translational modification includes, but is not limited to, phosphorylation, glycosylation, ubiquitination, nitrosylation, methylation, acetylation, lipidation and / or proteolysis. In some embodiments, the post-translational modification is any of the post-translational modifications disclosed elsewhere herein (see, for example, the section entitled “Definitions: Protein,” above).

[0105] In some embodiments, the plurality of protein analytes comprises at least 100 protein analytes (e.g., instances of protein analytes in the sample). In some embodiments, the plurality of protein analytes comprises at least 20, at least 30, at least 40, at least 50, at least 100, at least 200, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 20,000, at least 100,000, at least 500,000, at least 1 x 106, at least 5 x 106, at least 1 x 107, or at least 1 x 108protein analytes. In some embodiments, the plurality of protein analytes comprises no more than 1 x 109, no more than 1 x 108, no more than 1 x 107, no more than 1 x 106, no more than 100,000, no more than 50,000, no more than 20,000, no more than 10,000, no more than 5000, no more than 1000, no more than 500, no more than 100, no more than 50, or no more than 30 protein analytes (e.g., instances of protein analytes in the sample). In some embodiments, the plurality of protein analytes consists of from 20 to 400, from 100 to 2000, from 1000 to 5000, from 2000 to 10,000, from 10,000 to 50,000, from 50,000 to 1 x 106, from 1 x 106to 1 x 108, or from 1 x 107to 1 x 109protein analytes. In some embodiments, the plurality of protein analytes falls within another range starting no lower than 20 protein analytes and ending no higher than 1 x 109protein analytes. In some embodiments, the plurality of protein analytes comprises a plurality of different protein forms (e.g., proteoforms, variants, structures, and / or modifications thereof). In some embodiments, the plurality of different protein forms comprises at least 20, at least 30, at least 40, at least 50, at least 100, at least 200, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 20,000, at least 100,000, at least 500,000, at least 1 x 106, or at least 5 x 106different protein forms. In some embodiments, the plurality of different protein forms comprises no more than 1 x 107, no more than 1 x 106, no more than 100,000, no more than 50,000, no more than 20,000, no more than 10,000, no more than 5000, no more than 1000, no more than 500, no more than 100, no more than 50, or no more than 30 different protein DB2 / 651687418.2 30Attorney Docket No.: 139643-5001-WOforms. In some embodiments, the plurality of different protein forms consists of from 20 to 400, from 100 to 2000, from 1000 to 5000, from 2000 to 10,000, from 10,000 to 50,000, from 50,000 to 1 x 106, or from 1 x 106to 1 x 107different protein forms. In some embodiments, the plurality of different protein forms falls within another range starting no lower than 20 different protein forms and ending no higher than 1 x 107different protein forms. For instance, in some embodiments, a first protein form in the plurality of different protein forms corresponds to a different set of molecular binders and / or a different DNA concatemer sequence than a second protein form, as described in further detail below. In some embodiments, each respective protein form in the plurality of different protein forms corresponds to a different respective set of molecular binders and / or a different respective DNA concatemer sequence, as described in further detail below (see, for instance, the sections entitled “Molecular binders” and “DNA concatemers and interaction data structures,” below).

[0106] In some embodiments, the plurality of protein analytes comprises human proteins. In some embodiments, the plurality of protein analytes comprises bacterial or yeast proteins. Alternatively, or additionally, in some embodiments, the plurality of protein analytes comprises one or more engineered proteins. Alternatively, or additionally, in some embodiments, the plurality of protein analytes comprises one or more proteins generated using de novo protein design. However, the present disclosure is not limited thereto.

[0107] Referring to Block 214, in some embodiments, a respective protein analyte in the plurality of protein analytes comprises (e.g., “is associated with,” “contains,” “is characterized by”) a corresponding set of protein interaction features in a plurality of protein interaction features. In some embodiments, a protein interaction feature in the plurality of protein interaction features is selected from the group consisting of: hydrogen bond donors, hydrogen bond acceptors, hydroxyl groups, positively charged atoms, negatively charged atoms, aromatic rings, and aliphatic hydrophobic groups.

[0108] In some embodiments, a protein interaction feature in the plurality of protein interaction features is selected from the group consisting of: hydrogen bond donors, hydrogen bond acceptors, hydroxyl groups, positively charged atoms, negatively charged atoms, aromatic rings, and aliphatic hydrophobic groups.

[0109] In some embodiments, a protein interaction feature in the plurality of protein interaction features is selected from the group consisting of: hydrogen bond donors, hydrogenDB2 / 651687418.2 31Attorney Docket No.: 139643-5001-WObond acceptors, hydroxyl groups, positively charged atoms, negatively charged atoms, aromatic rings, aliphatic hydrophobic groups, a metal -binding site (e.g., a thiol, carboxylate, imidazole), a carbonyl group, a carboxylic acid, an amine, an ether, an alkene, a sulfur-containing group, a phosphate, a TI system, and a hydrophobic alkyl chain).

[0110] In some embodiments, the corresponding set of protein interaction features for the first protein analyte is known (e.g., for detection of protein analytes) or unknown (e.g., for identification of protein analytes).

[0111] In some embodiments, the plurality of protein analytes comprises a first class of protein analytes and a second class of protein analytes. In some embodiments, the first class of protein analytes comprises one or more protein analytes with which a plurality of molecular binders interacts. In some embodiments, the second class of protein analytes comprises one or more protein analytes with which the plurality of molecular binders does not interact. For instance, in some embodiments, methods include targeting the first class of protein analytes with the plurality of molecular binders for identification, characterization, and / or detection. In some embodiments, the protein analytes include proteins for which known molecular binders are available and / or proteins for which known molecular binders are unavailable. For example, in some embodiments, methods include targeting known proteins (e.g., for characterization and / or detection). In some embodiments, methods include targeting unknown proteins (e.g., for identification).

[0112] In some embodiments, the protein analytes include proteins of a target class and / or proteins of an off-target class. For instance, in some embodiments, methods include targeting protein analytes of a particular class of interest for identification, characterization, and / or detection. In some embodiments, the target class is any of the classes of protein analytes disclosed herein (e.g., enzymes, kinases, proteases, antibodies, or receptors).

[0113] Initiator tags and binding moieties.

[0114] Referring to Block 216, in some embodiments, each respective protein analyte in the plurality of protein analytes is attached to an initiator tag comprising a DNA initiator sequence. In some embodiments, the DNA concatemer is appended to the DNA initiator sequence of the first protein analyte using the DNA polymerase.

[0115] In some embodiments, methods further include attaching the initiator tag to the first protein analyte. Referring to Block 218, in some embodiments, attaching the initiator tag to a respective protein analyte in the plurality of protein analytes comprisesDB2 / 651687418.2 32Attorney Docket No.: 139643-5001-WOfunctionalizing the respective protein analyte using a linker that attaches the respective protein analyte to the initiator tag.

[0116] In some embodiments, as illustrated in FIG. 4A, the linker is a bifunctional linker, and the respective protein analyte is functionalized with the bifunctional linker using primary amine modification. In some embodiments, the primary amine modification includes modifying a primary amine of the respective protein analyte, thereby obtaining a first attachment moiety comprising a modified primary amine on the respective protein analyte. In some embodiments, the primary amine modification further includes attaching the bifunctional linker to the respective protein analyte, where the bifunctional linker comprises a second attachment moiety and a third attachment moiety, and where the second attachment moiety of the bifunctional linker binds to the first attachment moiety of the respective protein analyte. In some embodiments, the primary amine modification further includes attaching the initiator tag to the bifunctional linker, where the initiator tag comprises a fourth attachment moiety that binds to the third attachment moiety.

[0117] In some embodiments, methods for functionalizing protein analytes include primary amine bioconjugation, site- selective primary amine modification (e.g., site- selective lysine modification, site-selective N-terminal modification, selective N-terminal modification on a-amino groups, selective N-terminal modification on amino acid residues).

[0118] In some embodiments, methods for functionalizing protein analytes include, but are not limited to, lysine conjugation (e.g, via N-hydroxysuccinimide (NHS) esters, isocyanates and isothiocyanates, 4-azidobenzoyl fluoride (ABF), P-lactams, Phospha-Mannich reaction, a,P-unsaturated sulfonamide, and / or sulfonyl acrylate); cysteine single thiol functionalization (e.g., mal eimide, alkynyl carboxylic acid derivatives, 5-methylene pyrrolone (5MP), 5,5'-dithiobis-(2-nitrobenzoic acid) (DTNB), phenyloxadiazole sulfone (PODS), cyclooctyne, phosphonamidate, 3 -arylpropionitrile (APN), perfluoroarene, ethynylbenziodoxolone (EBX), bicyclo[1.1.0]butane (BCB) carboxylic amide, and / or allenamide); cysteine disulfide functionalization (e.g., disulfide rebridging using bis-sulfones, divinylpyrimidine (DVP), 3-bromo-5-methylene pyrrolones (3Br-5MP), arylenedipropiolonitrile (ADPN), 3,3-Bis(bromomethyl)oxetane, dichlorotetrazine, thiol-yne coupling, 2H-azirines-2-carboxamides, and / or dichloroacetophenone); tyrosine functionalization (e.g., Manni ch-type three component reaction, diazonium salt, 4-phenyl-3H-l,2,4-triazole-3,5(4H)-dione (PTAD), and / or phenothiazine); tryptophan functionalization (e.g., 9-azabicyclo [3.3.1]nonane-3-one-N-oxyl (keto-ABNO) and / or N-substitutedDB2 / 651687418.2 33Attorney Docket No.: 139643-5001-WOpyridinium salts); and / or methionine functionalization (e.g., oxaziridine, hypervalent iodine, and / or lumiflavin-catalysis).

[0119] Methods for functionalizing protein analytes are known in the art. Suitable methods for functionalizing protein analytes are further described, for example, in Vught et aL, 2014, “Site-specific functionalization of proteins and their applications to therapeutic antibodies,” Computational and Structural Biotechnology Journal 9: e201402001;Tantipanjaporn and Wong, 2023, “Development and recent advances in lysine and N-terminal bioconjugation for peptides and proteins,” Molecules. 28(3): 1083; and Kang etal., 2021, “Recent developments in chemical conjugation strategies targeting native amino acids in proteins and their applications in antibody-drug conjugates,” Chemical Science 12(41): 13613, each of which is hereby incorporated herein by reference in its entirety.

[0120] In some embodiments, the initiator tag is attached to the respective protein analyte using azide-alkyne cycloaddition. In some embodiments, as illustrated in FIG. 4A, the initiator tag is attached to the bifunctional linker using click chemistry or trans-cyclooctene (TCO) click chemistry.

[0121] Generally, click chemistry employs highly selective and efficient chemical reactions to rapidly link two or more molecules. In some embodiments, click chemistry utilizes copper-catalyzed azide-alkyne cycloaddition (CuAAC), in which an azide group reacts with an alkyne group to form a stable triazole ring. Advantageously, click chemistry reactions are fast, occur under mild conditions, and produce minimal byproducts, making these reactions suitable for various applications. Another click reaction contemplated for use herein is strain-promoted alkyne-azide cycloaddition (SPAAC).

[0122] Other methods for attaching molecules to functionalized protein analytes, including but not limited to click chemistry, are contemplated for use in the present disclosure, for instance as described in Devaraj and Finn, 2021, “Introduction: click chemistry,” Chem Rev. 121(12):6697-6698; Hein et al., 2008, “Click chemistry, a powerful tool for pharmaceutical sciences,” Pharmaceutical research 25(10) :2216; and Fantoni et aL, 2021, “A hitchhiker’s guide to click-chemistry with nucleic acids,” Chem Rev. 121(12):7122-7154, each of which is hereby incorporated herein by reference in its entirety.

[0123] Referring to Block 220, in some embodiments, a respective protein analyte in the plurality of protein analytes comprises one or more binding moieties. Referring to Block 222, in some embodiments, methods further include attaching the initiator tag to a binding DB2 / 651687418.2 34Attorney Docket No.: 139643-5001-WOmoiety of the respective protein analyte in the plurality of protein analytes, where the initiator tag is specific to the binding moiety.

[0124] In some embodiments, protein analytes with specific binding moieties include different protein classes with specific epitopes that can be preferentially targeted with different small molecules. For example, in some embodiments, the methods of the present disclosure further comprise attaching a linker to a binding moiety of a protein analyte, and further attaching an initiator tag to the linker (e.g., via click chemistry).

[0125] Referring to Block 224, in some embodiments, the plurality of protein analytes comprises a first class of protein analytes and a second class of protein analytes, each protein analyte in the first class of protein analytes comprises a set of binding moieties specific to the first class of protein analytes, and each protein analyte in the second class of protein analytes comprises a set of binding moieties specific to the second class of protein analytes. In some embodiments, the set of binding moieties specific to the first class of protein analytes are different from (e.g., not the same as) the set of binding moieties specific to the second class of protein analytes.

[0126] In some embodiments, the first protein is a member of the first class of protein analytes and comprises a binding moiety that is specific to the first class of protein analyte.

[0127] In some embodiments, a binding moiety in the one or more binding moieties is an epitope.

[0128] In some embodiments, the use of class-specific linkers and / or initiator tags allows for preferential targeting of classes of protein analytes. In some embodiments, the method comprises small molecules or chemical probes designed to selectively bind to target proteins to identify protein targets, discover drug candidates, and understand disease mechanisms. In some embodiments, the plurality of protein analytes comprises a first class of protein analytes and a second class of protein analytes, each protein analyte in the first class of protein analytes is attached to a first initiator tag comprising a first DNA initiator sequence, and each protein analyte in the second class of protein analytes is free of an initiator tag comprising the first DNA initiator sequence. In some embodiments, each protein analyte in the second class of protein analytes is attached to a second initiator tag comprising a second DNA initiator sequence, and each protein analyte in the first class of protein analytes is free of an initiator tag comprising the second DNA initiator sequence. In some embodiments, the methods of the present disclosure further include using the first DNA DB2 / 651687418.2 35Attorney Docket No.: 139643-5001-WOinitiator sequence and the second DNA initiator sequence to determine a corresponding class for the first protein analyte and a corresponding class for the second protein analyte. In some embodiments, the first initiator tag preferentially binds to a first epitope on each protein analyte in the first class of protein analytes, and the second initiator tag preferentially binds to a second epitope on each protein analyte in the second class of protein analytes.

[0129] Other methods of attaching initiator tags to protein analytes are possible, as will be apparent to one skilled in the art.

[0130] In some embodiments, the initiator tag is attached directly to the protein analyte without use of a linker.

[0131] In some embodiments, a respective protein analyte in the plurality of analytes comprises one or more initiator tags. In some embodiments, each respective protein analyte comprises at least 1 initiator tag. In some embodiments, each respective protein analyte comprises no more than 1 initiator tag. In some embodiments, each respective protein analyte consists of exactly 1 initiator tag. In some embodiments, a distribution of linkers attached to protein analytes in the plurality of protein analytes follows a Poisson distribution or a near Poisson distribution.

[0132] In some embodiments, a respective protein analyte comprises at least 2, at least 3, at least 4, at least 5, at least 10, at least 50, at least 100, at least 200, or at least 300 initiator tags. In some embodiments, a respective protein analyte comprises no more than 500, no more than 300, no more than 100, no more than 50, no more than 10, no more than 5, or no more than 3 initiator tags. In some embodiments, a respective protein analyte consists of from 2 to 4, from 2 to 8, from 3 to 7, from 5 to 10, from 10 to 80, from 50 to 200, from 100 to 400, or from 300 to 500 initiator tags. In some embodiments, a respective protein analyte comprises another range of tags starting no lower than 2 initiator tags and ending no higher than 500 initiator tags.

[0133] Molecular binders.

[0134] In some embodiments, the plurality of molecular binders comprises at least 10, 50, 100, or 500 different molecular binders.

[0135] In some embodiments, the plurality of molecular binders comprises at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 200, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 20,000, at least 50,000, at least 100,000, at least 500,000, or at least 1 x 106different molecular binders. In some DB2 / 651687418.2 36Attorney Docket No.: 139643-5001-WOembodiments, the plurality of molecular binders consists of no more than 5 x 106, no more than 1 x 106, no more than 500,000, no more than 100,000, no more than 50,000, no more than 20,000, no more than 10,000, no more than 5000, no more than 1000, no more than 500, no more than 100, no more than 50, or no more than 10 different molecular binders. In some embodiments, the plurality of molecular binders consists of from 50 to 400, from 100 to 2000, from 1000 to 5000, from 2000 to 10,000, from 10,000 to 50,000, from 20,000 to 200,000, from 100,000 to 1 x 106, or from 500,000 to 5 x 106different molecular binders. In some embodiments, the plurality of molecular binders falls within another range starting no lower than 5 different molecular binders and ending no higher than 5 x 106different molecular binders.

[0136] Referring to Block 226, in some embodiments, a respective molecular binder in the plurality of molecular binders has a length of at least 20 amino acids.

[0137] In some embodiments, a respective molecular binder comprises at least 5, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 100, at least 300, at least 500, at least 1000, at least 5000 amino acids. In some embodiments, a respective molecular binder consists of no more than 10,000, no more than 5000, no more than 1000, no more than 500, no more than 300, no more than 100, no more than 50, no more than 40, no more than 20, or no more than 10 amino acids. In some embodiments, a respective molecular binder consists of from 5 to 20, from 12 to 30, from 20 to 60, from 40 to 100, from 80 to 200, from 100 to 500, from 200 to 1000, from 500 to 5000, or from 2000 to 10,000 amino acids. In some embodiments, a respective molecular binder consists of another range of amino acids starting no lower than 5 amino acids and ending no higher than 10,000 amino acids.

[0138] In some embodiments, a respective molecular binder in the plurality of molecular binders has a molecular weight of at least 20 kDa. In some embodiments, a respective molecular binder has a molecular weight of at least 5, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 100, or at least 300 kDa. In some embodiments, a respective molecular binder has a molecular weight of no more than 500, no more than 300, no more than 100, no more than 50, no more than 40, no more than 20, or no more than 10 kDa. In some embodiments, a respective molecular binder has a molecular weight of from 5 to 20, from 12 to 30, from 20 to 60, from 40 to 100, from 80 to 200, or from 100 to 500 kDa. In some embodiments, a respective molecular binder has a molecular weight falling within another range starting no lower than 5 kDa and ending no higher than 500 kDa.DB2 / 651687418.2 37Attorney Docket No.: 139643-5001-WO

[0139] Referring to Block 228, in some embodiments, each respective molecular binder in the plurality of molecular binders comprises a peptide. Advantageously, in some instances, synthesizing and sourcing short peptides are easier and cheaper than larger binders, such as antibodies. In some embodiments, each respective molecular binder in the plurality of molecular binders comprises a small molecule. In some embodiments, each respective molecular binder in the plurality of molecular binders comprises an antibody.

[0140] In some embodiments, molecular binders comprise any molecule or chemical compound having the ability to bind to a protein analyte. In some embodiments, the plurality of molecular binders comprises peptides, proteins, polypeptides, DNA, RNA, small molecules, antibodies, and / or ligands. In some embodiments, each respective molecular binder satisfies any two or more rules, any three or more rules, or all four rules of the Lipinski’s rule of Five: comprising (i) not more than five hydrogen bond donors, (ii) not more than ten hydrogen bond acceptors, (iii) a molecular weight under 500 Daltons, and (iv) a LogP under 5.

[0141] In some embodiments, the plurality of molecular binders is obtained from a database or reference library of molecular binders. In some embodiments, the plurality of molecular binders comprises known single-chain variable fragments (scFv) or nanobody molecular binders obtained for common protein analyte targets. In some embodiments, the plurality of molecular binders comprises known binders for common post-translational modifications (e.g., lectins for glycans).

[0142] In some embodiments, the plurality of molecular binders comprises a plurality of random peptide binders obtained from peptide libraries. In some embodiments, the plurality of molecular binders includes completely random short peptides generated by NNK codons (e.g., where N = any nucleotide, and K = G or T). In some embodiments, the plurality of molecular binders includes random short peptides that are directly chemically synthesized. Advantageously, this approach is extremely cheap and generates very large molecular diversity. Without being limited to any one theory of operation, such libraries traditionally have very low hit rates, but may provide higher hit rates for low-affinity, non-specific binders.

[0143] In some embodiments, the plurality of molecular binders comprises a library of semi-random binders. In some embodiments, the semi-random binders are generated computationally. In some embodiments, the library of semi-random binders comprises lessDB2 / 651687418.2 38Attorney Docket No.: 139643-5001-WOthan 107unique molecules. Advantageously, the use of semi-random binders provides higher hit rates by avoiding issues with expression and solubility (e.g., by using random fragments of known protein sequences or samples from generative protein language models such as antibody-specific models).

[0144] In some embodiments, the plurality of molecular binders comprises partially randomized peptide scaffolds. In some embodiments, the plurality of molecular binders is obtained by using a partially-randomized version of an extremely thermostable small peptide scaffold, as described in Blanchard et al., 2023, “Hyperstable synthetic miniproteins as effective ligand scaffolds,” ACS Synthetic Biology 12(12), pp. 3608-3622; and McConnell et al., 2023, “Determinants of developability and evolvability of synthetic mini proteins as ligand scaffolds,” Journal of Molecular Biology 435(24), pp. 168339-168340, each of which is hereby incorporated herein by reference in its entirety. Without being limited to any one theory of operation, hyperstable folds are highly resilient to random mutations, which have, generally, an additive effect on stability. As such, the scaffold is able to accommodate a wide range of sequences to bind diverse targets. In some embodiments, designing the initial scaffold sequence is performed computationally. For instance, in some embodiments, Metropolis-Hastings or other Markov Chain Monte Carlo methods are used to sample sequences that have high predicted confidence using AlphaFold, as well as high predicted thermostability and high likelihood. In some embodiments, high predicted thermostability is optionally determined using a model trained or finetuned on large thermostability datasets as described in Tsuboyama etal., 2023, “Mega-scale experimental analysis of protein folding stability in biology and design,” Nature 620(7973), pp. 434-444, which is hereby incorporated herein by reference in its entirety. In some embodiments, likelihood is optionally determined according to a protein language model like ESM2 as described in Lin et al., 2023, “Evolutionary-scale prediction of atomic-level protein structure with a language model,” Science 379(6637), pp. 1123-1130, which is hereby incorporated herein by reference in its entirety. Sequences jointly satisfying these criteria are likely to be highly expressible, soluble, and stable.

[0145] In some implementations, surface positions are next selected for mutation. In some embodiments, these are chosen based on computational criteria (e.g., which positions are least likely to impact the predicted metrics described above, or predicted impact on structure) or randomized. The scaffolded approach enjoys a few additional advantages to random peptides from the computational point of view. Current computational methods DB2 / 651687418.2 39Attorney Docket No.: 139643-5001-WOstruggle with loop-mediated binding (e.g., in antibodies); by selecting a scaffold and mutation positions that are likely to bind using a different mode (e.g., helix-mediated binding), it is possible to generate molecular binders that are more useful for downstream computational analysis and model training. Similarly, hits from the scaffolded library can be used in exactly the same way as de novo designed scaffolded binders above (e.g., to later generate nonspecific binders).

[0146] In some embodiments, the methods of the present disclosure further include screening an initial selection of molecular binders (e.g., from a database or reference library) prior to contacting the plurality of protein analytes with the plurality of molecular binders.

[0147] In some embodiments, each respective molecular binder in the plurality of molecular binders binds to one or more protein interaction features in a plurality of protein interaction features. In some embodiments, as noted above, protein interaction features include, but are not limited to, hydrogen bond donors, hydrogen bond acceptors, hydroxyl groups, positively charged atoms, negatively charged atoms, aromatic rings, and aliphatic hydrophobic groups. In some embodiments, each respective molecular binder in the plurality of molecular binders binds weakly to the one or more protein interaction features in the plurality of protein interaction features. In some embodiments, the plurality of molecular binders comprises one or more molecular binders generated using de novo design. Example methods for de novo binder design suitable for use in the present disclosure are described, for example, in Pacesa etal., “BindCraft: one-shot design of functional protein binders,” bioRxiv 2024, doi: 10.1101 / 2024.09.30.615802, which is hereby incorporated herein by reference in its entirety. However, the present disclosure is not limited thereto.

[0148] In some embodiments, one or more molecular binders in the plurality of molecular binders is a weak affinity binder. In some embodiments, each molecular binder in the plurality of molecular binders is a weak affinity binder. In some embodiments, the molecular binder has a dissociation constant (KD) of at least 1 micromolar (pM). In some embodiments, the molecular binder has a KD of at least 0.01 nM, at least 0.05 nM, at least 0.1 nM, at least 0.5 nM, at least 1 nM, at least 5 nM, at least 10 nM, at least 50 nM, at least 0.1 pM, at least 0.5 pM, at least 1 pM, at least 5 pM, at least 10 pM, at least 50 pM, at least 100 pM, or at least 500 pM. In some embodiments, the molecular binder has a KD of no more than 1 mM, no more than 500 pM, no more than 200 pM, no more than 100 pM, no more than 50 pM, no more than 10 pM, no more than 1 pM, no more than 0.1 pM, no more than 10 nM, no more than 1 nM, or no more than 0.1 nM. In some embodiments, the molecular DB2 / 651687418.2 40Attorney Docket No.: 139643-5001-WObinder has a KD of from 0.01 nM to 0.1 nM, from 0.1 nM to 1 nM, from 1 nM to 0.1 pM, from 0.1 pM to 1 pM, from 1 pM to 50 pM, from 10 pM to 200 pM, from 100 pM to 500 pM, or from 500 pM to 1 mM. In some embodiments, the molecular binder has a KD falling within another range starting no lower than 0.01 nM and ending no higher than 1 mM.

[0149] In some embodiments, one or more molecular binders in the plurality of molecular binders is a nonspecific binder. In some embodiments, each molecular binder in a subplurality of the plurality of molecular binders is a nonspecific binder. In some embodiments, at least 1%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the plurality of molecular binders are nonspecific binders. In some embodiments, no more than 99%, no more than 90%, no more than 80%, no more than 70%, no more than 60%, no more than 50%, no more than 40%, no more than 30%, or no more than 20% of the plurality of molecular binders are nonspecific binders. In some embodiments, from 1% to 20%, from 10% to 40%, from 30% to 60%, from 40% to 70%, from 60% to 90%, or from 70% to 100% of the plurality of molecular binders are nonspecific binders. In some embodiments, the plurality of molecular binders comprises a different range of nonspecific binders starting no lower than 1% and ending no higher than 99% of the plurality of molecular binders. In some embodiments, each molecular binder in the plurality of molecular binders is a nonspecific binder. In some embodiments, a respective molecular binder in the plurality of molecular binders binds to two or more protein interaction features in the plurality of protein interaction features. In some embodiments, the respective molecular binder binds to two or more protein interaction features of a respective protein analyte.

[0150] Nonspecific interactions contemplated for use in the present disclosure are further disclosed, for instance, in the section entitled “Definitions: Nonspecific,” above. In some embodiments, a respective molecular binder in the plurality of molecular binders binds to two or more protein analytes in the plurality of protein analytes.

[0151] In some embodiments, at least the first protein analyte in the plurality of protein analytes binds to a corresponding set of molecular binders in the plurality of molecular binders. In some embodiments, each respective protein analyte in the plurality of protein analytes binds to a corresponding set of molecular binders in the plurality of molecular binders. In some embodiments, the corresponding set of molecular binders comprises at least 2 molecular binders.DB2 / 651687418.2 41Attorney Docket No.: 139643-5001-WO

[0152] In some embodiments, a corresponding set of molecular binders comprises at least 2, at least 5, at least 10, at least 20, at least 30, at least 50, at least 100, at least 500, at least 1000, at least 2000, or at least 5000 different molecular binders. In some embodiments, the corresponding set of molecular binders consists of no more than 10,000, no more than 5000, no more than 2000, no more than 1000, no more than 500, no more than 200, no more than 100, no more than 50, no more than 30, no more than 10, or no more than 5 different molecular binders. In some embodiments, the corresponding set of molecular binders consists of from 2 to 10, from 8 to 30, from 10 to 60, from 50 to 200, from 100 to 1000, from 500 to 2000, from 1000 to 8000, or from 5000 to 10,000 molecular binders. In some embodiments, the corresponding set of molecular binders falls within another range starting no lower than 2 molecular binders and ending no higher than 10,000 molecular binders.

[0153] As described above, in some embodiments, the plurality of protein analytes 142 in the sample of the subject comprises at least 2, at least 5, at least 10, at least 20, at least 30, at least 50, or at least 100 different protein analytes (e.g., proteoforms, variants, structures, and / or modifications thereof). In some embodiments, the plurality of protein analytes 142 in the sample of the subject comprises at least 25, at least 75, at least 100, at least 200, at least 300, at least 500, or at least 1000 different protein analytes. In some embodiments, the plurality of protein analytes 142 in the sample of the subject comprises at least 250, at least 750, at least 1000, at least 2000, at least 3000, at least 5000, at least 1000, at least 2000, at least 5000, at least 10,000, at least 20,000, at least 100,000, at least 500,000, at least 1 x 106, or at least 5 x 106different protein analytes. In some embodiments, the plurality of protein analytes 142 in the sample of the subject consists of no more than 1 x 107, no more than 1 x 106, no more than 100,000, no more than 50,000, no more than 20,000, no more than 10,000, no more than 5000, no more than 1000, no more than 500, no more than 100, no more than 50, or no more than 30 different protein analytes. In some embodiments, the plurality of protein analytes 142 in the sample of the subject consists of from 200 to 1000, from 80 to 300, from 100 to 600, from 20 to 400, from 100 to 2000, from 1000 to 5000, from 2000 to 10,000, from 10,000 to 50,000, from 50,000 to 1 x 106, or from 1 x 106to 1 x 107different protein analytes. In some embodiments, the plurality of protein analytes 142 in the sample of the subject falls within another range starting no lower than 2 different protein analytes and ending no higher than 1 x 107different protein analytes. In some embodiments, a first protein analyte (e.g., a first proteoform) in the plurality of different protein analytes corresponds to a first set of molecular binders and a second protein analyte (e.g., a secondDB2 / 651687418.2 42Attorney Docket No.: 139643-5001-WOproteoform) in the plurality of different protein analytes corresponds to a second set of molecular binders that is not identical to the first set of molecular binders, for instance, where one or more interaction features available for binding on the first protein analyte are different from one or more interaction features available for binding on the second protein analyte. In some embodiments, and without being limited to any one theory of operation, a first protein analyte in the plurality of different protein analytes is characterized by a different conformation or three-dimensional structure relative to a second protein analyte, such that a first set of interaction features available for binding by the plurality of molecular binders on the first protein is not identical to a second set of interaction features available for binding on the second protein analyte. In some embodiments, each respective different protein analyte (e.g., different proteoforms) in the plurality of different protein analytes corresponds to a different respective set of molecular binders, for instance, where the corresponding set of interaction features available for binding on each respective different protein analyte is not identical to the corresponding set of interaction features available for binding on every other protein analyte in the plurality of different protein analytes.

[0154] In some embodiments, a first protein analyte in the plurality of protein analytes comprises the same or a different number of molecular binders in the corresponding set of molecular binders as a second protein analyte in the plurality of protein analytes. In other words, in some embodiments, two protein analytes may interact with different numbers of binders and / or different sets of binders. In some embodiments, the corresponding set of molecular binders for the first protein analyte collectively represent a corresponding set of protein interaction features for the first protein analyte.

[0155] Advantageously, in some embodiments, methods and systems disclosed herein utilize a set of molecular binders, optionally including one or more nonspecific, weakly interacting molecular binders to characterize all or a portion of the protein interaction features of a respective protein analyte.

[0156] Without being limited to any one theory of operation, folded proteins can have unique biomolecular surfaces, such that two protein analytes of interest will be distinguishable by the surface features available for binding interactions. Accordingly, in some implementations, proteins can be detected and / or identified by their interaction features. The present disclosure advantageously provides molecular binders that interact with various protein interaction features, which are detected and evaluated to provide a unique combinatorial fingerprint that can be used to identify, detect, quantify, and / or characterize a DB2 / 651687418.2 43Attorney Docket No.: 139643-5001-WOprotein analyte of interest within a sample. In some embodiments, the use of one or more nonspecific, weakly interacting binders additionally provides greater flexibility and generalizability by allowing the use of a small number of binders to identify a substantially larger number of protein analytes.

[0157] Consider, for instance, where the plurality of molecular binders comprises 100 molecular binders and each protein analyte in the plurality of analytes interacts with at least 10 molecular binders in the plurality of molecular binders (e.g., a combinatorial fingerprint representing at least 10 surface protein features that characterize each protein analyte). The number of different protein analytes that can be identified using such binder sets is more than C(100,10) or 1013or 10 trillion possible protein analytes. In another instance, the plurality of molecular binders comprises 1000 molecular binders and each protein analyte in the plurality of analytes interacts with at least 2 molecular binders in the plurality of molecular binders; here, the number of different protein analytes that can be identified is, at least, C(1000,2) or -500,000. In yet another example, the plurality of molecular binders comprises 200 molecular binders and each protein analyte in the plurality of analytes interacts with at least 30 molecular binders in the plurality of molecular binders, allowing for the identification of at least C(200,30) or 4 x 1035possible protein analytes. The number of identifiable protein analytes grows considerably larger in light of the fact that, in some implementations, different protein analytes comprise different numbers of interaction features and / or bind to different subsets of binders in the plurality of molecular binders. For instance, in some embodiments, a first protein analyte interacts with a first number of molecular binders, and a second protein analyte interacts with a second number of molecular binders that is different from the first number of molecular binders. In other instances, a first protein analyte interacts with a first subset of molecular binders, and a second protein analyte interacts with a second subset of molecular binders that is different from the first subset of molecular binders.

[0158] Advantageously, in addition, the multiplexing of molecular binders optionally including one or more nonspecific, weakly interacting binders obviates the need to identify highly specific binding partners for each possible protein analyte of interest. Such protein assays are laborious, challenging, and frequently have low rates of success. Accordingly, the presently disclosed systems and methods improve upon the technical field of proteomics and drug discovery by reducing the time and resources needed for protein identification, characterization, and / or detection.DB2 / 651687418.2 44Attorney Docket No.: 139643-5001-WO

[0159] In some embodiments, the plurality of protein analytes comprises a first class of protein analytes and a second class of protein analytes, where each protein analyte in the first class of protein analytes interacts with at least one molecular binder in the plurality of molecular binders, and each protein analyte in the second class of protein analytes does not interact with any molecular binders in the plurality of molecular binders. As described above, for example, in some embodiments, methods include targeting the first class of protein analytes with the plurality of molecular binders (e.g., for identification, characterization, and / or detection).

[0160] In some embodiments, the first class of protein analytes is a target class, and the second class of protein analytes is an off-target class. As described above, for example, in some embodiments, methods include targeting protein analytes of a particular class of interest. In some embodiments, the target class is any of the classes of protein analytes disclosed herein (see, for example, the section entitled “Protein analytes,” above). In some embodiments, the targeting is performed with or without a linker specific to the first class. In some embodiments, the targeting is performed with or without an initiator tag specific to the first class. See, for example, the sections above entitled “Protein analytes” and “Initiator tags and binding moieties,” above.

[0161] In some embodiments, the protein analyte is a protein-protein complex and the plurality of binder-analyte interactions comprises interactions between (i) a first set of molecular binders that form binder-analyte interactions with a first protein in the proteinprotein complex, and (ii) a second set of molecular binders that form binder-analyte interactions with a second protein in the protein-protein complex. In some embodiments, the protein analyte is a protein-protein complex, and the plurality of binder-analyte interactions comprises interactions between a set of molecular binders that form binder-analyte interactions with both (i) a first protein in the protein-protein complex, and (ii) a second protein in the protein-protein complex.

[0162] In some embodiments, the plurality of binder-analyte interactions comprises two, three, four, five, six, seven, eight, nine, or ten binder-analyte interactions between one or more protein analytes and one or more molecular binders. In some embodiments, the plurality of binder-analyte interactions comprises at least 2, at least 5, at least 10, at least 20, at least 30, at least 50, at least 100, at least 500, at least 1000, at least 2000, or at least 5000 binderanalyte interactions between one or more protein analytes and one or more molecular binders. In some embodiments, the plurality of binder-analyte interactions comprises no more than DB2 / 651687418.2 45Attorney Docket No.: 139643-5001-WO10,000, no more than 5000, no more than 2000, no more than 1000, no more than 500, no more than 200, no more than 100, no more than 50, no more than 30, no more than 10, or no more than 5 binder-analyte interactions between one or more protein analytes and one or more molecular binders. In some embodiments, the plurality of binder-analyte interactions consists of from 2 to 10, from 8 to 30, from 10 to 60, from 50 to 200, from 100 to 1000, from 500 to 2000, from 1000 to 8000, or from 5000 to 10,000 binder-analyte interactions between one or more protein analytes and one or more molecular binders. In some embodiments, the plurality of binder-analyte interactions falls within another range of molecular binders starting no lower than 2 binder-analyte interactions and ending no higher than 10,000 binderanalyte interactions.

[0163] In some embodiments, the protein analyte comprises one or more post-translational modifications and the plurality of binder-analyte interactions comprises interactions between the protein analyte and one or more molecular binders that interact preferentially with the one or more post-translational modifications.

[0164] Barcodes.

[0165] Referring again to Block 206, in some embodiments, each respective molecular binder in the plurality of molecular binders comprises a corresponding nucleic acid barcode, where the nucleic acid barcode comprises a nucleic acid barcode sequence that is unique to the respective molecular binder. In some embodiments, the corresponding nucleic acid barcode sequence for each respective molecular binder is selected from a plurality of nucleic acid barcode sequences (e.g., unique nucleic acid barcode sequences).

[0166] In some embodiments, the plurality of nucleic acid barcode sequences comprises at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 200, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 20,000, at least 50,000, at least 100,000, at least 500,000, or at least 1 x 106nucleic acid barcode sequences. In some embodiments, the plurality of nucleic acid barcode sequences comprises no more than 5 x 106, no more than 1 x 106, no more than 500,000, no more than 100,000 no more than 50,000, no more than 20,000, no more than 10,000, no more than 5000, no more than 1000, no more than 500, no more than 100, no more than 50, or no more than 10 nucleic acid barcode sequences. In some embodiments, the plurality of nucleic acid barcode sequences consists of from 50 to 400, from 100 to 2000, from 1000 to 5000, from 2000 to 10,000, from 10,000 to 50,000, from 20,000 to 200,000, from 100,000 to 1 x 106, orDB2 / 651687418.2 46Attorney Docket No.: 139643-5001-WOfrom 500,000 to 5 x 106nucleic acid barcode sequences. In some embodiments, the plurality of nucleic acid barcode sequences falls within another range starting no lower than 5 nucleic acid barcode sequences and ending no higher than 5 x 106nucleic acid barcode sequences.

[0167] Referring to Block 230, in some embodiments, the nucleic acid barcode sequence comprises DNA. In some embodiments, the nucleic acid barcode sequence is ssDNA or dsDNA. In some embodiments, the nucleic acid barcode comprises RNA, DNA, or a combination thereof.

[0168] As used herein, the term “nucleic acid” refers to deoxyribonucleic acid (DNA), ribonucleic acid (RNA) or any hybrid or fragment thereof. In some embodiments, the nucleic acid barcode comprises any nucleic acid structure. In some embodiments, the nucleic acid barcode comprises any secondary or tertiary structure. In some embodiments, the nucleic acid barcode comprises a linear structure, a loop, a helix, and / or a hairpin.

[0169] In some embodiments, the nucleic acid barcode sequence comprises from 3 to 30 nucleotides. In some embodiments, the nucleic acid barcode sequence comprises at least 3, at least 5, at least 10, at least 20, at least 30, at least 50, or at least 100 nucleotides. In some embodiments, the nucleic acid barcode sequence comprises no more than 200, no more than 100, no more than 50, no more than 30, no more than 10, or no more than 5 nucleotides. In some embodiments, the nucleic acid barcode sequence consists of from 3 to 10, from 8 to 30, from 10 to 60, or from 50 to 200 nucleotides. In some embodiments, the nucleic acid barcode sequence comprises another range of nucleotides starting no lower than 3 nucleotides and ending no higher than 200 nucleotides.

[0170] In some embodiments, a corresponding nucleic acid barcode sequence for a first molecular binder in the plurality of molecular binders comprises a first nucleotide sequence that is reverse complementary to a second nucleotide sequence of at least one other molecular binder in the plurality of molecular binders. Referring to Block 232, in some embodiments, a first nucleic acid barcode sequence for a first molecular binder in the plurality of molecular binders comprises a first portion positioned at a 5’ end of the nucleic acid barcode sequence, a second nucleic acid barcode sequence for a second molecular binder in the plurality of molecular binders comprises a second portion positioned at a 3’ end of the nucleic acid barcode sequence, and the first portion of the first nucleic acid barcode sequence hybridizes to a reverse complement of the second portion of the second nucleic acid barcode sequence.DB2 / 651687418.2 47Attorney Docket No.: 139643-5001-WO

[0171] In some embodiments, as illustrated in FIG. 4A, a first barcode sequence in the plurality of nucleic acid barcode sequences comprises the same first constant portion and the same second constant portion as a second barcode sequence in the plurality of nucleic acid barcode sequences, such that the second constant portion of the second barcode sequence is hybridizable to a reverse complement of the first constant portion of the first barcode sequence. Advantageously, in some embodiments, the use of reverse complementary sequences (e.g., constant portions) between different barcodes provides “sticky ends” after DNA extension to which subsequent barcodes may hybridize, allowing for concatenation of consecutive barcode sequences.

[0172] In some embodiments, the nucleic acid barcode sequence comprises a first constant portion and a second constant portion, wherein the first constant portion is positioned at a 5’ end of the nucleic acid barcode sequence, the second constant portion is positioned at a 3’ end of the nucleic acid barcode sequence, and the first constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the second constant portion.

[0173] In some embodiments, each respective barcode sequence in the plurality of nucleic acid barcode sequences comprises the same first portion and the same second portion, such that each respective barcode sequence in the plurality of barcode sequences is hybridizable to at least a portion of a reverse complement of every other barcode sequence in the plurality of barcode sequences. In some embodiments, different subsets of barcode sequences comprise different constant portion sequences, such that only barcode sequences within a respective subset of barcode sequences can be consecutively concatenated. Referring to Block 234, in some embodiments, the plurality of molecular binders comprises a first subset of molecular binders and a second subset of molecular binders; for each molecular binder in the first subset of molecular binders, the corresponding nucleic acid barcode sequence comprises a first constant portion and a second constant portion, where the first constant portion is positioned at a 5’ end of the nucleic acid barcode sequence and the second constant portion is positioned at a 3’ end of the nucleic acid barcode sequence; for each molecular binder in the second subset of molecular binders, the corresponding nucleic acid barcode sequence comprises a third constant portion and a fourth constant portion, where the third constant portion is positioned at a 5’ end of the nucleic acid barcode sequence and the fourth constant portion is positioned at a 3’ end of the nucleic acid barcode sequence; the first constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the DB2 / 651687418.2 48Attorney Docket No.: 139643-5001-WOfourth constant portion; and the second constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the third constant portion.

[0174] Advantageously, and without being limited to any one theory of operation, the use of different subsets of barcode sequences reduces the likelihood that two identical barcodes corresponding to the same molecular binder will be consecutively concatenated. In some embodiments, the plurality of nucleic acid barcode sequences comprises a plurality of subsets of barcode sequences, each respective subset of barcode sequences corresponding to a respective pair of constant sequences. In some embodiments, the plurality of subsets of barcode sequences comprises at least 2, at least 5, at least 10, at least 20, at least 30, at least 50, or at least 100 subsets of barcode sequences. In some embodiments, the plurality of subsets of barcode sequences comprises no more than 200, no more than 100, no more than 50, no more than 30, no more than 10, or no more than 5 subsets. In some embodiments, the plurality of subsets of barcode sequences consists of from 2 to 10, from 8 to 30, from 10 to 60, or from 50 to 200 subsets. In some embodiments, the plurality of subsets of barcode sequences falls within another range starting no lower than 2 subsets and ending no higher than 200 subsets.

[0175] Various chemical methods have been developed to connect peptides to DNAs either covalently or non-covalently, catering for different chemical and biological purposes. Covalent ligations produce a stable DNA-peptide hybrid macromolecule and, in most cases, are realized via aminoacylations, orthogonal reactions, Michael additions, and thiol oxidations. Non-covalent conjugations, however, take advantage of high binding affinity between two cognate entities (one attached to DNAs and the other fused to peptides) to confine the DNA and peptide modalities in proximity via protein-protein, protein-ligand, and ligand-ligand interactions. Methods for generating molecular binders comprising nucleic acid barcodes are further described, for example, in Danielson etal., 2023, “Peptide-DNA conjugates as building blocks for de novo design of hybrid nanostructures,” Cell Reports Physical Science 4(10): 101620, which is hereby incorporated herein by reference in its entirety.

[0176] Referring to Block 236, in some embodiments, the methods of the present disclosure further include generating the plurality of molecular binders using mRNA display, cDNA display, in vivo mRNA display, LABEL-seq, Molecular Indexing of Proteins by SelfAssembly (MIPSA), or a combination thereof. However, the present disclosure is not limited thereto.DB2 / 651687418.2 49Attorney Docket No.: 139643-5001-WO

[0177] In some embodiments, as illustrated in FIG. 4A, the plurality of molecular binders comprising the nucleic acid barcode sequence is generated using mRNA display or cDNA display. Methods for mRNA display and cDNA display include modifying mRNA with the translation inhibitor puromycin. When this molecule is translated, the peptide product is covalently linked to the RNA that generated it. As these methods rely on using chemically modified template RNA molecules, they can only be used for in vitro translation, and as such they do not scale well to large preparations, because of the cost of reagents to do cell-free protein synthesis. Methods for mRNA display and cDNA display contemplated for use in the present disclosure are further described, for example, in Yamaguchi etal., 2009, “Cdna display: a novel screening method for functional disulfide-rich peptides by solid-phase synthesis and stabilization of mma-protein fusions,” Nucleic Acids Research 37(16):el08, which is hereby incorporated herein by reference in its entirety.

[0178] Methods for Molecular Indexing of Proteins by Self-Assembly (MIPS A) include modifications of mRNA display, in which a truncated mRNA is reverse transcribed to cDNA and covalently attached to the protein. In some embodiments, the nucleic acid that is displayed is an arbitrary barcode, made out of DNA, and covalently attached to the polypeptide. In some embodiments, the protein coding sequence is modified to have an N-term halo tag, and the 5’ untranslated region of the mRNA is barcoded and then reverse transcribed with a chloroalkane primer. During in vitro translation, the nascent protein product contains a halo domain which reacts in cis with the chloroalkane primer forming a covalent link. Methods for MIPSA contemplated for use in the present disclosure are further described, for example, in Credle eta!.. 2022, “Unbiased discovery of autoantibodies associated with severe COVID-19 via genome-scale self-assembled DNA-barcoded protein libraries,” Nat Biomed Eng. 6(8), pp. 992-1003, which is hereby incorporated herein by reference in its entirety.

[0179] Methods for in vivo mRNA display and LABEL-seq include engineering bacterial cells with a library of plasmids that encode both a nucleic acid barcode, and a molecular binder protein that is tagged with a specific nucleic-acid binding domain. In some embodiments, delivery of the plasmid material to the bacterial cells is bottlenecked such that exactly one plasmid molecule is transformed into each bacterial cell. When a bacterial cell produces the two components encoded on this plasmid, the cell itself acts as a microcompartment in which the two parts can associate with each other to form a specific binder-barcode complex. As bacteria do not exchange cellular contents with each other, DB2 / 651687418.2 50Attorney Docket No.: 139643-5001-WOdifferent bacterial cells in the same culture flask can encode different barcoded binder complexes without affecting the binder-barcode mapping, allowing pools of barcoded binders to be prepared at once. Furthermore, as bacteria reproduce asexually, the binder-barcode association is maintained while the bacteria propagate, allowing for high scalability. In some embodiments, there is no upper bound to how much barcoded material can be made.

[0180] Generally, in vivo mRNA display and LABEL-seq comprise methods similar to conventional mRNA display, but use an MS2 binder fused to a protein of interest to bind to an mRNA engineered to contain MS2 loop in an untranslated region. Biosynthesis of the parts and subsequent in-cell association forms a tight but noncovalent RNA-protein complex that can then be worked with in a pool. In some embodiments, different variants of this workflow produce proteins that bind to their entire mRNA (e.g., in vivo mRNA display) or to a short barcode that specifies protein identity (e.g., LABEL-seq). Methods for in vivo mRNA display and LABEL-seq contemplated for use in the present disclosure are further described, for example, in Oikonomou et al., 2020, “Zw vivo mRNA display enables large-scale proteomics by next generation sequencing” Proc Natl Acad Sci U S A. 117(43), pp. 26710-26718; and Simon etal., 2024, “Multiplexed profiling of intracellular protein abundance, activity, interactions and druggability with LABEL-seq,” Nat Methods. 21(11), pp. 2094-2106, each of which is hereby incorporated herein by reference in its entirety.

[0181] Reactions.

[0182] Referring to Block 237, in some embodiments, the methods of the present disclosure further include generating, for a first protein analyte 142 in the plurality of protein analytes, a DNA concatemer 144 comprising a corresponding set of DNA barcode sequences 146. In some embodiments, each DNA barcode sequence 146 in the corresponding set of DNA barcode sequences: i) corresponds to a nucleic acid barcode sequence 124 of a respective molecular binder 122, in the plurality of molecular binders, that forms a binderanalyte interaction with the first protein analyte 142, and ii) is appended to the first protein analyte 142 using the DNA polymerase, responsive to formation of the binder-analyte interaction between the respective molecular binder 122 and the first protein analyte 142, thereby causing a spatial order of DNA barcode sequences 146 in the corresponding set of DNA barcode sequences to correspond to a temporal order of formation of binder-analyte interactions between respective molecular binders 122, in the plurality of molecular binders, and the first protein analyte 142.DB2 / 651687418.2 51Attorney Docket No.: 139643-5001-WO

[0183] In some embodiments, the DNA polymerase comprises DNA polymerase I or T4 polymerase. DNA polymerases are known in the art. DNA polymerases contemplated for use in the present disclosure include, but are not limited to, DNA polymerase I, DNA polymerase II, DNA polymerase III, DNA polymerase IV, DNA polymerase V, DNA polymerase a, DNA polymerase P, DNA polymerase y, DNA polymerase 8, DNA polymerase a, T4 polymerase, Taq polymerase, Pfu polymerase, and / or reverse transcriptase. Other DNA polymerases may be possible for use in the present disclosure, as will be apparent to one skilled in the art.

[0184] In some embodiments, contacting the plurality of protein analytes with the plurality of molecular binders and the DNA polymerase further comprises contacting the plurality of protein analytes with a plurality of deoxynucleotide triphosphates (dNTPs). In some embodiments, the plurality of dNTPs comprises at least 3 DNA base identities. In some embodiments, the plurality of dNTPs consists of 3 DNA base identities.

[0185] In some embodiments, the plurality of dNTPs consists of 3 DNA base identities selected from adenine (dATP), cytosine (dCTP), guanine (dGTP), and thymine (dTTP). For instance, in some embodiments, the plurality of dNTPs does not include guanine (e.g., the plurality of dNTPs consists of dATP, dCTP, and dTTP). In some embodiments, the plurality of dNTPs does not include cytosine. In some embodiments, the plurality of dNTPs does not include adenosine. In some embodiments, the plurality of dNTPs does not include thymine.

[0186] In some embodiments, the nucleic acid barcode comprises: (i) a nucleic acid sequence consisting of three nucleic acid base identities complementary to the three DNA base identities in the plurality of dNTPs, and (ii) a stop position consisting of a fourth nucleic acid base identity, and where an extension reaction catalyzed by the DNA polymerase ceases upon reaching the stop position.

[0187] Referring to FIGS. 4A-B, for example, in some embodiments, the DNA polymerase is used to extend a DNA concatemer using the nucleic acid barcode sequence of a respective molecular binder upon binding of the molecular binder to a protein analyte (e.g., the first protein analyte). In some embodiments, the method includes using a first base identity as a stop base, for instance, by placing the stop base at an end (e.g, a 5’ end) of the nucleic acid barcode and using only the remaining three base identities for the remainder of the nucleic acid barcode sequence. Without being limited to any one theory of operation, inDB2 / 651687418.2 52Attorney Docket No.: 139643-5001-WOsome embodiments, a DNA extension reaction catalyzed by the DNA polymerase is halted upon reaching the stop base when the reaction mixture is supplied with only the remaining dNTP base identities. Thus, consider the example in which the stop base is a guanine (G), the remainder of the nucleic acid barcode sequence consists of adenine (A), thymine (T), and cytosine (C), and the reaction mixture for the DNA extension reaction is supplied with dATP, dGTP, and dTTP. In some such implementations, the DNA extension reaction ceases upon reaching the stop base G due to a lack of dCTP in the reaction mixture.

[0188] In some embodiments, the plurality of dNTPs consists of all 4 DNA base identities, including dATP, dCTP, dGTP, and dTTP.

[0189] In some embodiments, the DNA concatemer is generated manually. Referring to Block 238, in some embodiments, the DNA concatemer is generated using an automated reaction device.

[0190] In some embodiments, the automated reaction device is an autonomous laboratory or automated chemical synthesis robot. Generally, performing chemistry on automation can differ from manually performed chemistry. Automated chemistry reduces the need for individual labor and training, with the added advantage of standardizing experiments and data read outs. Variables such as human error, time of day, order of addition of reagents, and laboratory temperature can lead to varying data outputs even when using common workflows. Conversely, due to the high number of reactions performed during automated chemistry, automated approaches are sensitive to the conditions and reagents in the reaction in order to achieve successful outcomes. In some embodiments, the automated device interacts with system 100 to integrate one or more instruments. In some embodiments, the automated device comprises one or more integration software tools for scheduling, control, and / or automation of the one or more instruments. Various automated devices and integration modules are contemplated for use in the present disclosure, as will be apparent to one skilled in the art.

[0191] DNA concatemers and interaction data structures.

[0192] Referring to Block 240, in some embodiments, the methods of the present disclosure further include determining an interaction data structure 152 for the first protein analyte 142 based on sequencing of the DNA concatemer 144, where the interaction data structure 152 comprises an indication of a presence or abundance 154 of each respectiveDB2 / 651687418.2 53Attorney Docket No.: 139643-5001-WOmolecular binder 122 in the plurality of molecular binders that formed a binder-analyte interaction with the first protein analyte 142.

[0193] In particular, referring again to Block 237, in some embodiments, the DNA concatemer comprises DNA barcode sequences appended to a respective protein analyte (e.g., the first protein analyte) using the DNA polymerase, responsive to formation of a binder-analyte interaction between the respective molecular binder and the first protein analyte. In some embodiments, the DNA barcode sequence appended to a respective protein analyte represents a spatial order of DNA barcode sequences in the corresponding set of DNA barcode sequences that correspond to a temporal order of formation of binder-analyte interactions between respective molecular binders, in the plurality of molecular binders, and the respective protein analyte.

[0194] In some embodiments, the DNA concatemer comprises DNA barcode sequences for two, three, four, five, six, seven, eight, nine, or ten different molecular binders. Without being limited to any one theory of operation, in some embodiments, the number of DNA barcode sequences in the DNA concatemer indicates that a corresponding number of molecular binders in the plurality of molecular binders interacted with the respective protein analyte. In some embodiments, the DNA concatemer comprises DNA barcode sequences for at least 2, at least 5, at least 10, at least 20, at least 30, at least 50, at least 100, at least 500, at least 1000, at least 2000, or at least 5000 different molecular binders. In some embodiments, the DNA concatemer comprises DNA barcode sequences for no more than 10,000, no more than 5000, no more than 2000, no more than 1000, no more than 500, no more than 200, no more than 100, no more than 50, no more than 30, no more than 10, or no more than 5 different molecular binders. In some embodiments, the DNA concatemer consists of from 2 to 10, from 8 to 30, from 10 to 60, from 50 to 200, from 100 to 1000, from 500 to 2000, from 1000 to 8000, or from 5000 to 10,000 DNA barcode sequences for different molecular binders. In some embodiments, the DNA concatemer comprises DNA barcode sequences for another range of molecular binders starting no lower than 2 molecular binders and ending no higher than 10,000 molecular binders.

[0195] In some embodiments, the DNA concatemer comprises the DNA barcode sequence for the same molecular binder more than two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, sixteen, seventeen, eighteen, nineteen, or twenty times. This indicates that a particular molecular binder type interacted with the respective protein analyte multiple times.DB2 / 651687418.2 54Attorney Docket No.: 139643-5001-WO

[0196] In some embodiments, the DNA concatemer comprises the DNA barcode sequence for molecular binders in the plurality of molecular binders more than two, three, four, five, six, seven, eight, nine, or ten times. This indicates that molecular binders in the plurality of molecular binders, collectively interacted with the respective protein analyte multiple times.

[0197] In some embodiments, each DNA barcode sequence in the DNA concatemer represents a different molecular binder. This also indicates that molecular binders in the plurality of molecular binders, collectively interacted with the respective protein analyte multiple times.

[0198] Accordingly, in some embodiments, the systems and methods provided herein allow transient interactions between molecular binder and analyte to be captured as persistent DNA records through a polymerase-mediated concatemer process.

[0199] In some embodiments, each DNA barcode sequence in the corresponding set of DNA barcode sequences for a respective DNA concatemer corresponds to a respective molecular binder in the plurality of molecular binders that forms a binder-analyte interaction with a respective protein analyte (e.g., the first protein analyte). In some embodiments, the corresponding set of DNA barcode sequences for the respective DNA concatemer represents the set of molecular binders that interact with the respective protein analyte and the set of protein interaction features associated therewith that characterizes the respective protein analyte. Accordingly, in some such embodiments, the DNA concatemer provides a DNA fingerprint that represents the set of protein interaction features unique to the protein analyte. For instance, as noted above, in some embodiments, a first protein analyte (e.g., a first proteoform) in a plurality of different protein analytes corresponds to a first DNA concatemer and a second protein analyte (e.g. , a second proteoform) in the plurality of different protein analytes corresponds to a second DNA concatemer that is not identical to the first DNA concatemer, for instance, where one or more interaction features available for binding on the first protein analyte are different from one or more interaction features available for binding on the second protein analyte. In some embodiments, and without being limited to any one theory of operation, a first protein analyte in the plurality of different protein analytes is characterized by a different amino acid sequence, secondary structure, tertiary structure, and / or conformation relative to a second protein analyte, such that a first set of interaction features available for binding by the plurality of molecular binders on the first protein is not identical to a second set of interaction features available for binding on the second protein DB2 / 651687418.2 55Attorney Docket No.: 139643-5001-WOanalyte. In some such embodiment, the transient interactions between each molecular binder in a first set of molecular binders for the first protein analyte is captured resulting in a first DNA concatemer that is different compared to a second DNA concatemer capturing transient interactions between a second set of molecular binders and the second protein analyte. In some embodiments, each respective different protein analyte (e.g., different proteoforms) in the plurality of different protein analytes corresponds to a different respective DNA concatemer, for instance, where the corresponding set of interaction features available for binding on each respective different protein analyte is not identical to the corresponding set of interaction features available for binding on every other protein analyte in the plurality of different protein analytes.

[0200] In some embodiments, various implementations are contemplated for appending DNA barcode sequences to a respective protein analyte. In some embodiments, protein analytes are contacted with each molecular binder in the plurality of molecular binders one at a time. In some embodiments, protein analytes are contacted with individual subsets of molecular binders, in a plurality of subsets of the plurality of molecular binders. In some embodiments, molecular binders are pooled and added to the plurality of protein analytes in a single pool. Without being limited to any one theory of operation, in some embodiments, two or more molecular binders form binder-analyte interactions with a single respective protein analyte (e.g, the first protein analyte) simultaneously. In some embodiments, and without being limited to any one theory of operation, dissociation of molecular binders from the protein analyte allows for additional molecular binders to form binder-analyte interactions with the respective protein analyte. Depending on the protein interaction features, the size of the protein analyte, and the size of the molecular binders, in some embodiments, one or more molecular binders form binder-analyte interactions with a single respective protein analyte (e.g., the first protein analyte) at any time, as will be apparent to one skilled in the art.

[0201] Referring to Block 242, in some embodiments, generating the DNA concatemer comprises, for each respective molecular binder in the plurality of molecular binders that forms a binder-analyte interaction with the first protein analyte: (i) hybridizing a 5’ portion of the corresponding nucleic acid barcode sequence for the respective molecular binder to a 3’ portion of the DNA concatemer, or a DNA initiator sequence appended to the first protein analyte; (ii) extending the DNA concatemer, or the DNA initiator sequence, using the DNA polymerase by generating a reverse complement of the nucleic acid barcode DB2 / 651687418.2 56Attorney Docket No.: 139643-5001-WOsequence; and (iii) releasing the corresponding nucleic acid barcode. FIG. 4B illustrates an example process of hybridizing, extending, and releasing that is repeated for each subsequent molecular binder in the plurality of molecular binders that forms a binder-analyte interaction with the first protein analyte. Referring again to Block 206, in some embodiments, nucleic acid barcodes comprise one or more constant portions and / or sequences that allow for hybridization to one or more consecutive nucleic acid barcode sequences or reverse complements thereof. In some embodiments, the nucleic acid barcode sequences corresponding to the plurality of molecular binders comprise the same or different constant portions or sequences thereof.

[0202] Referring to Block 244, in some embodiments, generating the DNA concatemer comprises adding each nucleic acid barcode sequence to the DNA concatemer using auto-cycling proximity recording (APR). Generally, autocycling proximity recording (APR) includes repeatedly producing proximity records of nearby DNA barcodes. For instance, in some implementations, APR generates proximity data autonomously and repeatedly, at tunable distances, by non-destructively copying DNA-barcoded probes. The result is a complete set of proximity information and complete reconstruction. The APR mechanism provides for high signal levels, as well as allows resampling of the same molecular targets in different proximity arrangements. APR is further described, for example, in Schaus et al., 2017, “A DNA nanoscope via auto-cycling proximity recording,” Nat Commun. 8(1), p. 696, which is hereby incorporated herein by reference in its entirety.

[0203] In some embodiments, as illustrated in FIG. 4C, the methods of the present disclosure further include sequencing the DNA concatemer, thereby generating a plurality of nucleic acid sequence reads. In some embodiments, methods further include, prior to the sequencing, isolating, purifying, enriching, and / or fragmenting the DNA concatemer.Methods for nucleic acid sequencing, including processing nucleic acid samples for sequencing, are known in the art.

[0204] In some embodiments, nucleic acid sequencing is performed by various commercial systems or sequencing platforms. In some embodiments, nucleic acid sequencing is performed using, e.g., nucleic acid amplification, polymerase chain reaction (PCR) (e.g., digital PCR and droplet digital PCR (ddPCR), quantitative PCR, real time PCR, multiplex PCR, PCR-based singleplex methods, emulsion PCR), and / or isothermal amplification. In some embodiments, nucleic acid sequencing comprises DNA hybridization methods (e.g., Southern blotting), restriction enzyme digestion methods, Sanger sequencing methods, next-82 / 651687418.2 57Attorney Docket No.: 139643-5001-WOgeneration sequencing methods (e.g., single-molecule real-time sequencing, nanopore sequencing, and Polony sequencing), second-generation sequencing (SGS) methods, ligation methods, and microarray methods. In some embodiments, nucleic acid sequencing comprises targeted sequencing, single molecule real-time sequencing, exon sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy termination sequencing, wholegenome sequencing, sequencing by hybridization, pyrosequencing, capillary electrophoresis, gel electrophoresis, duplex sequencing, cycle sequencing, single-base extension sequencing, solid-phase sequencing, high-throughput sequencing, massively parallel signature sequencing, co-amplification at lower denaturation temperature-PCR (COLD-PCR), sequencing by reversible dye terminator, paired-end sequencing, near-term sequencing, exonuclease sequencing, sequencing by ligation, short-read sequencing, single-molecule sequencing, sequencing-by-synthesis, real-time sequencing, reverse-terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, and any combinations thereof. In some embodiments, nucleic acid sequencing comprises short-read sequencing and / or long-read sequencing.

[0205] In some embodiments, the methods of the present disclosure further include preprocessing the nucleic acid sequence reads prior to determining an interaction data structure from the DNA concatemer. In some embodiments, the interaction data structure comprises an interaction vector, an interaction matrix, and / or an interaction tensor. In some embodiments, the interaction data structure comprises any format suitable for comprising a plurality of indications of a presence or abundance of abundance of each respective molecular binder in a plurality of molecular binders that forms a binder-analyte interaction with one or more protein analytes in a plurality of protein analytes.

[0206] In some embodiments, raw DNA sequence reads are preprocessed by standard filtering for quality. In some embodiments, the preprocessing further includes rejecting sequence reads that are shorter than a predetermined length; in some embodiments, the predetermined length is determined based on the sequencing technology used.

[0207] Referring to Block 246, in some embodiments, the indication of the presence or abundance of each respective molecular binder in the plurality of molecular binders is a count of DNA barcode sequences, in the DNA concatemer, that correspond to the respective molecular binder.DB2 / 651687418.2 58Attorney Docket No.: 139643-5001-WO

[0208] For instance, in some embodiments, the methods of the present disclosure further include generating, for each respective molecular binder in the plurality of molecular binders, a set (e.g. , a vector) of counts, where each respective count in the set of counts corresponds to a number of instances that a DNA barcode associated with the respective molecular binder is observed in the plurality of sequence reads. In some implementations, for a single sequence read, the set of counts comprises a vector c E NB, where B is the number of molecular binders in the plurality of molecular binders (e.g., hundreds to thousands) and N is the set of natural numbers. This is done by splitting the read on known separator sequences between barcodes and assigning each barcode to a molecular binder in the plurality of molecular binders according to edit distance to a known set of binder barcodes. In some embodiments, barcode sequences too far from any known binder barcode are discarded. In some embodiments, barcode sequences that are more than 4, 5, 6, 7, 8, 9, 10, 11, or 12 edits from any binder barcode in the plurality of barcodes used for the plurality of molecular binders are discarded. In some embodiments, the methods of the present disclosure further comprise performing error correction after splitting the sequence read to improve for barcode detection.

[0209] Alternatively or additionally, in some embodiments, the interaction data structure is determined by a procedure comprising: mapping a plurality of sequence reads obtained from the sequencing of the DNA concatemer to a plurality of reference sequences comprising DNA barcode sequences associated with each molecular binder in the plurality of molecular binders, and determining, for each respective molecular binder in the plurality of molecular binders, a corresponding count of the corresponding DNA barcode sequence, wherein each entry in the interaction data structure comprises a count of DNA barcode sequences corresponding to the respective molecular binder.

[0210] Referring to Block 248, in some embodiments, the methods of the present disclosure further include determining a respective interaction data structure (e.g., vector) for each protein analyte in the plurality of protein analytes, thereby generating a plurality of interaction data structures (e.g., vectors), and using the plurality of interaction data structures (e.g., vectors) to obtain an interaction matrix dimensioned by the plurality of protein analytes and the plurality of molecular binders, where each entry in the interaction matrix comprises a count of DNA barcode sequences corresponding to a respective molecular binder for a respective protein analyte. FIG. 4C further illustrates an example interaction matrix comprising counts C of DNA barcode sequences (where C is either zero or a positive DB2 / 651687418.2 59Attorney Docket No.: 139643-5001-WOinteger), for each protein analyte in the plurality of protein analytes and each molecular binder in the plurality of molecular binders. In some embodiments, as noted above, an interaction data structure (e.g., one or more interaction vectors and / or an interaction matrix thereof) represents a “fingerprint,” specific to a respective protein analyte, of the molecular binders that form binder-analyte interactions with the protein analyte.

[0211] Protein analyte indications.

[0212] Referring to Block 249, in some embodiments, the methods of the present disclosure further include receiving, responsive to inputting the interaction data structure (e.g., vector 152) to a model, an indication of an identity 162 for the first protein analyte 142, as output from the model.

[0213] Referring to Block 250, in some embodiments, the indication of the identity for the first protein analyte is a protein identity of the first protein analyte. Referring to Block 252, in some embodiments, the indication of the identity for the first protein analyte is a probability of each candidate identity, in a set of candidate identities. In some embodiments, the model generates, responsive to inputting an interaction data structure, a probability that a respective protein analyte has a candidate identity, for each respective candidate identity in a set of candidate identities. Referring to Block 254, in some embodiments, the indication of the identity for the first protein analyte is a probability of a candidate identity, and methods further include applying an identification threshold to the probability to assign the candidate identity to the first protein analyte. In some embodiments, the identification threshold is a probability of at least 40%.

[0214] In some embodiments, the identification threshold is at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 95, at least 98, or at least 99 percent. In some embodiments, the identification threshold is no more than 100, no more than 99, no more than 98, no more than 95, no more than 90, no more than 80, no more than 70, no more than 60, no more than 50, or no more than 40 percent. In some embodiments, the identification threshold is from 30 to 70, from 50 to 80, from 40 to 90, from 70 to 99, or from 80 to 100 percent. In some embodiments, the identification threshold falls within another range starting no lower than 30 percent and ending no higher than 100 percent.

[0215] In some embodiments, methods further include receiving, as output from the model, an abundance of the first protein analyte.DB2 / 651687418.2 60Attorney Docket No.: 139643-5001-WO

[0216] Referring to Block 256, in some embodiments, the indication of the identity for the first protein analyte is obtained based on sampling from a reference distribution, where the reference distribution comprises, for each respective protein analyte in the plurality of protein analytes, for each respective molecular binder in the plurality of molecular binders, a corresponding rate of production for the DNA barcode sequence corresponding to the nucleic acid barcode sequence that is unique to the respective molecular binder.

[0217] In some embodiments, protein identities are determined by a process comprising: (i) for each respective protein analyte in a plurality of known protein analytes, determining a corresponding set of protein interaction features and a set of molecular binders associated therewith (e.g., determining protein surface features for the protein analyte and the set of molecular binders that form binder-analyte interactions with the protein surface features). In some embodiments, the process further includes (ii) generating, for each respective protein analyte in the plurality of known protein analytes, a reference barcode fingerprint comprising, for each respective molecular binder that forms a binder-analyte interaction with the respective protein analyte, the corresponding DNA barcode sequence that is uniquely associated with the respective molecular binder. Accordingly, the reference barcode fingerprints indicate the unique sets of barcodes that can be used to identify and differentiate various protein analytes. In some embodiments, the process further includes (iii) obtaining, for each respective protein analyte in a plurality of unknown protein analytes, a corresponding interaction data structure using any of the methods disclosed above, and (iv) using the interaction data structure to determine an identity of an unknown protein analyte based on the reference barcode fingerprints. In some implementations, determining the identity of the unknown protein comprises performing a comparison of the interaction data structure against the reference barcode fingerprints (e.g., using a lookup table). In some implementations, determining the identity of the unknown protein comprises obtaining a probability of an identity based on a sampling from a reference distribution of reference barcode fingerprints (e.g., using Bayesian inference).

[0218] In some embodiments, as illustrated in FIG. 4C, methods further include training a model comprising a plurality of at least 1000 parameters to generate the indication of the identity responsive to inputting the interaction data structure. In some embodiments, training the model comprises applying a respective difference to a loss function to obtain a respective output of the loss function, where the respective difference is between, for each respective training analyte in a plurality of training analytes, (a) a predicted indication of DB2 / 651687418.2 61Attorney Docket No.: 139643-5001-WOidentity and (b) a measured indication of identity. In some embodiments, the training further comprises using the respective output of the loss function to adjust the one or more parameters in the plurality of parameters.

[0219] In some embodiments, the plurality of parameters includes at least 10, at least 100, at least 1000, at least 10,000, at least 100,000, at least 1 x 106, at least 1 x 107, or more parameters. In some embodiments, the plurality of parameters includes no more than 1 x 108, no more than 1 x 107, no more than 1 x 106, no more than 100,000, no more than 10,000, no more than 1000, or no more than 100 parameters. In some embodiments, the plurality of parameters consists of from 10 to 1000, from 100 to 100,000, from 10,000 to 1 x 107, or from 1 x 106to 1 x 108parameters. In some embodiments, the plurality of parameters falls within another range starting no lower than 10 parameters and ending no higher than 1 x 108parameters.

[0220] FIG. 4C illustrates an example implementation in which outputted probabilities are sampled from multinomial distribution of rates of production. In some implementations, rate distributions used for sampling are determined experimentally (e.g., using measured binder-protein kinetics) or predicted using a model fitted to known binderprotein kinetic data. The output is a collection of DNA sequence reads, each of which is generated by a single, unknown protein target from a known set of potential targets. In some embodiments, a protein target refers to a single protein, a known complex of proteins, or known post-translational modifications of single proteins or complexes of proteins. Each read is composed of a set of DNA barcode sequences, each of which identifies a single molecular binder that interacted with the protein target during the reaction. The indication outputted from the model is then used to identify which protein target generated each read.

[0221] Referring to Block 258, in some embodiments, methods further include using the indication of the identity to determine a presence or abundance of the first protein analyte responsive to a target condition. Referring to Block 260, in some embodiments, methods further include using the indication of the identity to determine a change in a presence or abundance of the first protein analyte responsive to a target condition. In some embodiments, methods and systems disclosed herein provide for the determination of a phenotype of a biological sample under a target condition.

[0222] In some embodiments, a phenotype includes, but is not limited to, expression or production of protein analytes, conformation of protein analytes, generation of protein-DB2 / 651687418.2 62Attorney Docket No.: 139643-5001-WOprotein complexes, binding affinity, and / or behavior of a protein analyte responsive to exposure to a chemical compound (e.g., “druggability”). In some embodiments, the target condition comprises a healthy condition, a normal condition, an unhealthy condition, a disease condition, a pre-treatment condition, a post-treatment condition, and / or a dysregulation of a physiological state. In some embodiments, the target condition comprises any of the target conditions disclosed elsewhere herein, including but not limited to healthy conditions, disease conditions, exposure to a chemical compound, and / or dysregulation of a biological intermediate in the sample of the subject.

[0223] In some embodiments, the methods of the present disclosure further include performing any of the methods and / or embodiments disclosed herein for each respective protein analyte in a plurality of protein analytes (e.g., for each respective protein analyte, other than the first protein analyte, in the plurality of protein analytes).

[0224] Example classification.

[0225] In some embodiments, each interaction data structure (e.g., binder count vector) c is classified as coming from one of a set of known protein targets {1, ..., P} using statistical or machine learning models. In an example implementation, the model uses a rate nmatrix M G R+' that specifies the rate of production of each barcode for each protein target. The interaction data structure (e.g., count vector) c is modeled given a protein identity i E { 1, ..., P} as sampled from a multinomial distribution, where P is a positive integer representing the number of unique protein analytes that are tracked (e.g, 10 or more, 100 or more, 1000 or more, or any of the ranges of protein analytes disclosed herein), and where:

[0226] Here, n = I c, and is the total number observed barcodes in the read. In some implementations, n is not used in the probabilistic model as it is roughly constant for each sequence read in such implementations. Given a prior probability over protein analyte species, A G AP, the posterior probability of each protein analyte species is computed given the interaction data structure (e.g, count vector) c:p(i | c) ocfp(c | i).

[0227] In some implementations, is the uniform distribution. The read is then classified as the protein species with the highest posterior probability while, optionally,DB2 / 651687418.2 63Attorney Docket No.: 139643-5001-WOrejecting or marking as anomalous reads with diffuse or low posterior probability.Advantageously, given the matrix M, this method is extremely computationally efficient.

[0228] Many variations on this method are possible; for instance, in some embodiments, ignoring the counts of each barcode and simply using the presence or absence of each barcode will be sufficient and, in some cases, more robust. In some such implementations, independent Bernoulli distributions are used instead of a multinomial distribution. In some embodiments, an end-to-end machine learning model is used to map the interaction data structure (e.g., count vector) c to a distribution over protein species directly. In some embodiments, such a model is either trained on experimental data or simulated data and, optionally, is fed additional information about the experimental conditions (e.g., concentrations of each binder species).

[0229] Example interaction modelling.

[0230] As noted above, in some embodiments, the model uses a rate matrix M for classification. In some implementations, a parameter for this analysis is the rate matrix M (or for simplified models a binary version of M). In some implementations, a pair of kinetics parameters for each binder-protein pair is known or predicted: the on-rate konG R+and the off-rate koy G / ?+. In some embodiments, the kinetics parameters depend on additional experimental conditions, such as buffer makeup, temperature, etc., that are incorporated in the model. In some embodiments, these dependencies are suppressed. The on-rate (times the binder concentration) is the rate at which an unbound molecule of the target protein becomes bound, and the off-rate is that rate at which a binder-target complex disassociates. A good approximation for the rate matrix is then:

[0231] Here Ct is the concentration of binder i in the experiment. This is a simple model of the binding kinetics of a single binder to a single protein. In some embodiments, more complex models of binding kinetics are used. However, this approximation is quite accurate (e.g., koiron the order of 100 / s, high binder concentration, etc.). In some embodiments, the requirement that the kinetics parameters be known is bypassed by directly estimating M (up to a multiplicative constant) from large amounts of experimental data collected in similar conditions to the classification experiment. Simpler likelihood models (e.g., general independent Bernoulli models or even Bernoulli models with only two values DB2 / 651687418.2 64Attorney Docket No.: 139643-5001-WOfor p) are also possible and do not require as detailed modelling; any of the models described herein can be easily translated to this scenario.

[0232] While the kinetics parameters for each binder-protein pair can be measured experimentally using standard techniques (e.g., Biolayer Interferometry (BLI) and / or Surface Plasmon Resonance (SPR)), this would be prohibitively expensive. Instead, in some implementations, computational models are used to fit these parameters to experimental data. These models fall into two classes: those that estimate kinetics parameters from experimental data on known binders and targets, and those that predict kinetics parameters for novel binders and targets directly from sequences or structures. These two classes overlap (for instance, in some embodiments, a machine learning model that operates on sequence data to predict kinetics parameters for known binders and targets is used) but are a useful distinction.

[0233] Each of these models takes a binder and target as input and output a pair of kinetics parameters konand kOff. In some embodiments, these models are fitted to a large collection of labelled experimental data (where the target proteins are known) using maximum likelihood under the probabilistic model described above (e.g., the posterior probability equation described above). Under this model konand fc are only identifiable up to a multiplicative constant. Additionally, in some embodiments, it can be reasonably assumed konis identical for all binders as it is diffusion-limited and most binders share a common fold. For the first class of models, k (and possibly kon) are simply estimated directly from a wide range of experimental data using the above probabilistic model equation, possibly with regularization.

[0234] The second class of models is much more involved, both computationally and in terms of complexity. In some embodiments, these models use a variety of machine learning techniques, including deep learning, to predict kinetics parameters from sequence or structure. In some embodiments, these models also use a variety of additional data sources, including experimental data or pretrained models.

[0235] In some implementations, the model comprises one or more black-box sequence-based models. These models map a binder sequence b and a target sequence t directly to either fc (and possibly kon . In some embodiments, they use a variety of pretrained models, for instance embeddings from protein language models such as ESM2 or structure prediction models such as AlphaFold and can be pretrained on existing proteinprotein interaction data. For instance, one model computes the single and pair embeddingsDB2 / 651687418.2 65Attorney Docket No.: 139643-5001-WOfrom AlphaFol d2 (AF2) when given as input the two chains b and t, and does a simple attention-based reduction (to remove sequence dimensions) before passing the input though a multilayer perceptron to predict koff.

[0236] In some embodiments, the methods of the present disclosure include adding additional outputs to a pretrained network (e.g., AlphaFold) and finetuning the model. While black-box models are extremely flexible they may require prohibitively large amounts of data to train and be prohibitively computationally expensive. Structure-based models are an alternative that may be more efficient but require good models of the structure of the bindertarget complex. One option here is to sample many models from AlphaFold (using, for instance, AFSample), and take the top models as ranked by AF2’s confidence metrics as input to our model. Given complex models, in some embodiments, a graph neural network (similar to ProteinMPNN) applied to each model is trained to predict koff. In some embodiments, custom models are trained on in-house data to predict poses as input to interaction models. For computational tractability, in some embodiments, these models are trained or pre-trained with only a subset of negative examples or to predict a binary indicator of binding.

[0237] In some embodiments, the model is any of the models and / or classifiers disclosed herein (see, for example, the section entitled “Definitions: Models,” above).

[0238] Example Methods for Detecting Protein Analytes

[0239] As illustrated in FIGS. 3A-B, another aspect of the present disclosure provides a method 300 for detecting protein analytes 142. Referring to Block 302, in some embodiments, methods include obtaining a plurality of protein analytes 142 from a sample of a subject, where the plurality of protein analytes 142 comprises a first class of protein analytes and a second class of protein analytes, and each protein analyte 142 in the first class of protein analytes is attached to a first initiator tag comprising a first DNA initiator sequence.

[0240] Referring to Block 304, in some embodiments, methods further include contacting the plurality of protein analytes 142 with a plurality of molecular binders 122 and a DNA polymerase, thereby forming a plurality of binder-analyte interactions. In some embodiments, each protein analyte 142 in the first class of protein analytes forms a binderanalyte interaction with at least one molecular binder 122 in the plurality of molecular binders, each respective molecular binder 122 in the plurality of molecular binders comprises DB2 / 651687418.2 66Attorney Docket No.: 139643-5001-WOa corresponding nucleic acid barcode sequence 124 that is unique to the respective molecular binder, and one or more molecular binders in the plurality of molecular binders 122 interacts nonspecifically with each protein analyte 142 in a corresponding subset of the first class of protein analytes (e.g., forms binder-analyte interactions with each protein analyte in the subset of protein analytes).

[0241] Referring to Block 306, in some embodiments, methods further include generating, for a first protein analyte 142 in the first class of protein analytes, a DNA concatemer 144 comprising a corresponding set of DNA barcode sequences 146, where each DNA barcode sequence 146 in the corresponding set of DNA barcode sequences: i) corresponds to a nucleic acid barcode sequence 124 of a respective molecular binder 122, in the plurality of molecular binders, that forms a binder-analyte interaction with the first protein analyte 142, and ii) is appended to the first DNA initiator sequence of the first protein analyte 142 using the DNA polymerase, responsive to formation of the binder-analyte interaction between the respective molecular binder 122 and the first protein analyte, thereby causing a spatial order of DNA barcode sequences 146 in the corresponding set of DNA barcode sequences to correspond to a temporal order of formation of binder-analyte interactions between respective molecular binders 122, in the plurality of molecular binders, and the first protein analyte 142.

[0242] Referring to Block 307, in some embodiments, the methods of the present disclosure further include determining a first interaction data structure 152 for the first protein analyte 142 based on a sequencing of the DNA concatemer 144, where the first interaction data structure 152 comprises an indication of a presence or abundance 154 of each respective molecular binder 122 in the plurality of molecular binders that formed a binderanalyte interaction with the first protein analyte 142. Referring to Block 308, in some embodiments, the methods of the present disclosure further include receiving, responsive to inputting the first interaction data structure 152 to a first model, an indication of an identity 162 for the first protein analyte 142, as output from the first model.

[0243] In some embodiments, each protein analyte in the second class of protein analytes is free of an initiator tag comprising the first DNA initiator sequence.

[0244] Referring to Block 310, in some embodiments, each protein analyte in the second class of protein analytes is attached to a second initiator tag comprising a second DNA initiator sequence, further comprising: generating, for a second protein analyte in theDB2 / 651687418.2 67Attorney Docket No.: 139643-5001-WOsecond class of protein analytes, a DNA concatemer comprising a corresponding set of DNA barcode sequences, where each DNA barcode sequence in the corresponding set of DNA barcode sequences: i) corresponds to a nucleic acid barcode sequence of a respective molecular binder, in the plurality of molecular binders, that forms a binder-analyte interaction with the second protein analyte, and ii) is appended to the second DNA initiator sequence of the second protein analyte using the DNA polymerase, responsive to formation of the binderanalyte interaction between the respective molecular binder and the second protein analyte, thereby causing a spatial order of DNA barcode sequences in the corresponding set of DNA barcode sequences to correspond to a temporal order of formation of binder-analyte interactions between respective molecular binders, in the plurality of molecular binders, and the second protein analyte.

[0245] Referring to Block 312, in some embodiments, the methods of the present disclosure further include determining a second interaction data structure for the second protein analyte based on a sequencing of the DNA concatemer, where the second interaction data structure comprises an indication of a presence or abundance of each respective molecular binder in the plurality of molecular binders that formed a binder-analyte interaction with the second protein analyte.

[0246] Referring to Block 314, in some embodiments, the methods further include receiving, responsive to inputting the second interaction data structure to a second model, an indication of an identity for the second protein analyte, as output from the second model.

[0247] In some embodiments the first model and the second model are the same.

[0248] In some embodiments the first model and the second model are different.

[0249] In some embodiments the first model and the second model are any of the models described in the present disclosure.

[0250] In some embodiments, each protein analyte in the first class of protein analytes is free of an initiator tag comprising the second DNA initiator sequence. In some embodiments, the methods of the present disclosure further include using the first DNA initiator sequence and the second DNA initiator sequence to determine a corresponding class for the first protein analyte and a corresponding class for the second protein analyte. In some embodiments, the first initiator tag preferentially binds to a first epitope on each protein analyte in the first class of protein analytes, and the second initiator tag preferentially binds to a second epitope on each protein analyte in the second class of protein analytes.DB2 / 651687418.2 68Attorney Docket No.: 139643-5001-WO

[0251] Additional Example Methods for Identifying Protein Analytes

[0252] In some embodiments, as illustrated in FIGS. 5A-B and 6A-B, the systems, methods, compositions, and kits disclosed herein further comprise identifying and / or detecting protein analytes using DNA concatemers formed by concatenating DNA barcodes uniquely associated with protein analytes onto molecular binders that form binder-analyte interactions with the protein analytes (e.g., “write to binder”). In some embodiments, each respective protein analyte in the plurality of protein analytes comprises a corresponding nucleic acid barcode, where the nucleic acid barcode comprises a nucleic acid barcode sequence that is unique to the respective protein analyte. In some embodiments, each respective molecular binder in the plurality of molecular binders is attached to an initiator tag comprising a DNA initiator sequence, and the DNA concatemer is appended to the DNA initiator sequence of a first molecular binder using the DNA polymerase. In some embodiments, the systems and methods of the present disclosure further include determining a respective interaction data structure (e.g., vector) for each molecular binder in the plurality of molecular binders, thereby generating a plurality of first interaction data structures (e.g., vectors) that carry a record of the protein analytes that formed binder-analyte interactions with each molecular binder. In some embodiments, systems and methods disclosed herein further include using the plurality of first interaction data structures (e.g., vectors) to obtain a second interaction data structure (e.g., an interaction matrix) dimensioned by the plurality of protein analytes and the plurality of molecular binders, where each entry in the interaction matrix comprises a count of DNA barcode sequences corresponding to a respective protein analyte for a respective molecular binder. As illustrated in FIGS. 4C, 5A-B, and 6A-B, and as described above with reference to FIG. 4C, in some implementations, an interaction matrix comprises counts C of DNA barcode sequences (where C is either zero or a positive integer), for each protein analyte in the plurality of protein analytes and each molecular binder in the plurality of molecular binders. In some embodiments, as noted above, a second interaction data structure (e.g., one or more interaction vectors and / or an interaction matrix thereof) represents a “fingerprint,” specific to a respective protein analyte, of the molecular binders that form binder-analyte interactions with the respective protein analyte. In some embodiments, systems and methods disclosed herein further include inputting an interaction data structure (e.g., one or more vectors 152, or an interaction matrix thereof) to a model, thus obtaining an indication of an identity 162 for at least a first protein analyte 142, as output from the model.DB2 / 651687418.2 69Attorney Docket No.: 139643-5001-WO

[0253] As illustrated in FIGS. 5A-B and 6A-B, another aspect of the present disclosure provides a method 500 for detecting protein analytes 142. Referring to Block 502, in some embodiments, method 500 includes obtaining a plurality of protein analytes 142 from a sample of a subject.

[0254] Referring to Block 504, in some embodiments, method 500 further includes contacting the plurality of protein analytes 142 with a plurality of molecular binders 122 and a DNA polymerase, thereby forming a plurality of binder-analyte interactions. In some embodiments, each respective protein analyte 142 in the plurality of protein analytes comprises a corresponding nucleic acid barcode sequence 124 that is unique to the respective protein analyte 142, and one or more molecular binders 122 in the plurality of molecular binders interacts nonspecifically with each protein analyte 142 in a corresponding subset of the plurality of protein analytes (e.g., forms binder-analyte interactions with each protein analyte in the subset of protein analytes).

[0255] Referring to Block 506, in some embodiments, method 500 further includes generating, for a first molecular binder 122-1 in the plurality of molecular binders, a DNA concatemer 144 comprising a corresponding set of DNA barcode sequences 146, where each DNA barcode sequence 146 in the corresponding set of DNA barcode sequences: i) corresponds to a nucleic acid barcode sequence 124 of a respective protein analyte 142, in the plurality of protein analytes, that forms a binder-analyte interaction with the first molecular binder 122-1, and ii) is appended to the first molecular binder 122-1 using the DNA polymerase, responsive to formation of the binder-analyte interaction between the respective protein analyte 142 and the first molecular binder 122-1, thereby causing a spatial order of DNA barcode sequences 146 in the corresponding set of DNA barcode sequences to correspond to a temporal order of formation of binder-analyte interactions between respective protein analytes 142, in the plurality of protein analytes, and the first molecular binder 122-1.

[0256] Referring to Block 507, in some embodiments, method 500 further includes determining an interaction data structure 152 for the first molecular binder 122-1 based on a sequencing of the DNA concatemer 144, wherein the interaction data structure comprises an indication of a presence or abundance of each respective protein analyte 142 in the plurality of protein analytes that formed a binder-analyte interaction with the first molecular binder.

[0257] Referring to Block 508, in some embodiments, method 500 further includes receiving, responsive to inputting at least the interaction data structure 152 to a model, anDB2 / 651687418.2 70Attorney Docket No.: 139643-5001-WOindication of an identity 162 for a first protein analyte 142 in the plurality of protein analytes, as output from the model.

[0258] Any of the systems, methods, and / or embodiments disclosed elsewhere herein for identifying a protein analyte, including samples, protein analytes, initiator tags, binding moieties, molecular binders, barcodes, reactions, DNA concatemers, interaction data structures, protein analyte indications, classification, and interaction models, are similarly contemplated for use in method 500, as will be apparent to one skilled in the art.

[0259] In some embodiments, each respective molecular binder in the plurality of molecular binders is attached to an initiator tag comprising a DNA initiator sequence. In some embodiments, the DNA concatemer is appended to the DNA initiator sequence of the first molecular binder using the DNA polymerase. In some embodiments, method 500 further includes attaching the initiator tag to the first molecular binder. In some embodiments, the attaching the initiator tag to a respective molecular binder in the plurality of molecular binders comprises functionalizing the respective molecular binder using a linker that attaches the respective molecular binder to the initiator tag. In some embodiments, the linker is a bifunctional linker, and the respective molecular binder is functionalized with the bifunctional linker using primary amine modification by a process comprising: modifying a primary amine of the respective molecular binder, thereby obtaining a first attachment moiety comprising a modified primary amine on the respective molecular binder, attaching the bifunctional linker to the respective molecular binder, where the bifunctional linker comprises a second attachment moiety and a third attachment moiety, and where the second attachment moiety of the bifunctional linker binds to the first attachment moiety of the respective molecular binder, and attaching the initiator tag to the bifunctional linker, where the initiator tag comprises a fourth attachment moiety that binds to the third attachment moiety. In some embodiments, the initiator tag is attached to the respective molecular binder using azide-alkyne cycloaddition. Alternatively, or in addition, in some embodiments, the initiator tag is attached to the bifunctional linker using click chemistry or trans-cyclooctene (TCO) click chemistry.

[0260] In some embodiments, a corresponding nucleic acid barcode sequence for a first protein analyte in the plurality of protein analytes comprises a first nucleotide sequence that is reverse complementary to a second nucleotide sequence of at least one other protein analyte in the plurality of protein analytes.DB2 / 651687418.2 71Attorney Docket No.: 139643-5001-WO

[0261] In some embodiments, a first nucleic acid barcode sequence for a first protein analyte in the plurality of protein analytes comprises a first portion positioned at a 5’ end of the nucleic acid barcode sequence, a second nucleic acid barcode sequence for a second protein analyte in the plurality of protein analytes comprises a second portion positioned at a 3’ end of the nucleic acid barcode sequence, and the first portion of the first nucleic acid barcode sequence hybridizes to a reverse complement of the second portion of the second nucleic acid barcode sequence.

[0262] In some embodiments, the nucleic acid barcode sequence comprises a first constant portion and a second constant portion, where the first constant portion is positioned at a 5’ end of the nucleic acid barcode sequence, the second constant portion is positioned at a 3’ end of the nucleic acid barcode sequence, and the first constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the second constant portion.

[0263] In some embodiments, the plurality of protein analytes comprises a first subset of protein analytes and a second subset of protein analytes. In some embodiments, for each protein analyte in the first subset of protein analytes, the corresponding nucleic acid barcode sequence comprises a first constant portion and a second constant portion, where the first constant portion is positioned at a 5’ end of the nucleic acid barcode sequence, and the second constant portion is positioned at a 3’ end of the nucleic acid barcode sequence. In some embodiments, for each protein analyte in the second subset of protein analytes, the corresponding nucleic acid barcode sequence comprises a third constant portion and a fourth constant portion, where the third constant portion is positioned at a 5’ end of the nucleic acid barcode sequence, and the fourth constant portion is positioned at a 3’ end of the nucleic acid barcode sequence. In some embodiments, the first constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the fourth constant portion, and the second constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the third constant portion.

[0264] In some embodiments, the generating comprises, for each respective protein analyte in the plurality of protein analytes that forms a binder-analyte interaction with the first molecular binder: hybridizing a 5’ portion of the corresponding nucleic acid barcode sequence for the respective protein analyte to a 3’ portion of the DNA concatemer, or a DNA initiator sequence appended to the first molecular binder, extending the DNA concatemer, or the DNA initiator sequence, using the DNA polymerase by generating a reverse complement of the nucleic acid barcode sequence, and releasing the corresponding nucleic acid barcode. DB2 / 651687418.2 72Attorney Docket No.: 139643-5001-WO

[0265] In some embodiments, the generating comprises adding each nucleic acid barcode sequence to the DNA concatemer using auto-cycling proximity recording (APR).

[0266] In some embodiments, the indication of the presence or abundance of each respective protein analyte in the plurality of protein analytes is a count of DNA barcode sequences, in the DNA concatemer, that correspond to the respective protein analyte.

[0267] In some embodiments, the interaction data structure is determined by a procedure comprising: mapping a plurality of sequence reads obtained from the sequencing of the DNA concatemer to a plurality of reference sequences comprising DNA barcode sequences associated with each protein analyte in the plurality of protein analytes, and determining, for each respective protein analyte in the plurality of protein analytes, a corresponding count of the corresponding DNA barcode sequence, where each entry in the interaction data structure comprises a count of DNA barcode sequences corresponding to the respective protein analyte.

[0268] Referring to Block 310, in some embodiments, method 500 further includes determining a respective interaction data structure for each molecular binder in the plurality of molecular binders, thereby generating a plurality of interaction data structures, and using the plurality of interaction data structures to obtain an interaction matrix dimensioned by the plurality of protein analytes and the plurality of molecular binders, where each entry in the interaction matrix comprises a count of DNA barcode sequences corresponding to a respective protein analyte for a respective molecular binder.

[0269] Referring to Block 312, in some embodiments, the indication of the identity for the first protein analyte is obtained based on a sampling from a reference distribution, where the reference distribution comprises, for each respective protein analyte in the plurality of protein analytes, for each respective molecular binder in the plurality of molecular binders, a corresponding rate of production for the DNA barcode sequence corresponding to the nucleic acid barcode sequence that is unique to the respective protein analyte.

[0270] Further Example Methods for Identifying Protein Analytes

[0271] Yet another aspect of the present disclosure provides a method for identifying protein analytes, comprising: A) obtaining a plurality of protein analytes from a sample of a subject; and B) contacting the plurality of protein analytes with a plurality of first molecular binders and a DNA polymerase, thereby forming a plurality of binder-analyte interactions. In some embodiments, the method further includes C) generating a DNA concatemerDB2 / 651687418.2 73Attorney Docket No.: 139643-5001-WOcomprising a corresponding set of DNA barcode sequences, where a spatial order of DNA barcode sequences in the corresponding set of DNA barcode sequences corresponds to a temporal order of formation of binder-analyte interactions between respective first molecular binders, in the plurality of first molecular binders, and a respective protein analyte in the plurality of protein analytes. In some embodiments, each DNA barcode sequence in the corresponding set of DNA barcode sequences: i) corresponds to a nucleic acid barcode sequence of a respective first molecular binder, in the plurality of first molecular binders, that forms a binder-analyte interaction with the respective protein analyte, and ii) is appended to the respective protein analyte using the DNA polymerase, responsive to formation of the binder-analyte interaction between the respective first molecular binder and the respective protein analyte.

[0272] In some embodiments, the method further includes D) determining an interaction data structure for the respective protein analyte based on a sequencing of the DNA concatemer, where the interaction data structure comprises an indication of a presence or abundance of each respective first molecular binder in the plurality of first molecular binders that formed a binder-analyte interaction with the respective protein analyte. In some embodiments, the method further includes E) receiving, responsive to inputting the interaction data structure to a model, an indication of an identity for the respective protein analyte, as output from the model.

[0273] In some embodiments, the contacting B) further comprises contacting the plurality of protein analytes with one or more additional molecular binders, other than the plurality of first molecular binders. In some embodiments, the one or more additional molecular binders are different from the plurality of first molecular binders. In some embodiments, a respective additional molecular binder in the one or more additional molecular binders is a different molecular binder from the plurality of first molecular binders. In some embodiments, a respective additional molecular binder in the one or more additional molecular binders is the same as one or more first molecular binders in the plurality of first molecular binders.

[0274] In some embodiments, where the contacting B) further comprises contacting the plurality of protein analytes with a plurality of second molecular binders. In some embodiments, each respective protein analyte in a first subset of the plurality of protein analytes forms a binder-analyte interaction with a second molecular binder in the plurality of second molecular binders. In some embodiments, each respective protein analyte in a second DB2 / 651687418.2 74Attorney Docket No.: 139643-5001-WOsubset of the plurality of protein analytes does not form binder-analyte interactions with the plurality of second molecular binders.

[0275] In some embodiments, each respective first molecular binder in the plurality of first molecular binders comprises a peptide.

[0276] In some embodiments, each respective second molecular binder in the plurality of second molecular binders comprises a small molecule. In some embodiments, a small molecule is a chemical compound of low molecular weight. In some embodiments, a small molecule has a molecular weight of less than 2000 daltons (Da). In some embodiments, the small molecule has a molecular weight of less than 2000, less than 1500, less than 1000, less than 800, less than 500, less than 300, or less than 100 Da. In some embodiments, the small molecule has a molecular weight of at least 10, at least 50, at least 100, at least 500, at least 800, at least 1000, or at least 1500 Da. In some embodiments, the small molecule has a molecular weight of from 10 to 100, from 50 to 300, from 200 to 500, from 300 to 1000, from 500 to 1200, or from 800 to 2000 Da. In some embodiments, the small molecule has a molecular weight falling within another range starting no lower than 10 Da and ending no higher than 2000 Da.

[0277] In some embodiments, each respective second molecular binder satisfies any two or more rules, any three or more rules, or all four rules of the Lipinski's rule of Five: (i) not more than five hydrogen bond donors, (ii) not more than ten hydrogen bond acceptors, (iii) a molecular weight under 500 Daltons, and (iv) a LogP under 5.

[0278] In some embodiments, the method further includes attaching, to each respective second molecular binder in the plurality of second molecular binders, an initiator tag comprising a DNA initiator sequence, where the DNA concatemer is appended to the DNA initiator sequence of the second molecular binder using the DNA polymerase.

[0279] In some embodiments, the method further includes, prior to contacting the plurality of protein analytes with the plurality of second molecular binders, functionalizing each respective second molecular binder in the plurality of second molecular binders with a corresponding linker comprising an attachment moiety for attachment with the initiator tag.

[0280] In some embodiments, the initiator tag is attached, via the corresponding linker, to each respective second molecular binder in the plurality of second molecular binders prior to contacting the plurality of protein analytes with the plurality of second molecular binders.DB2 / 651687418.2 75Attorney Docket No.: 139643-5001-WO

[0281] In some embodiments, the initiator tag is attached, via the corresponding linker, to each respective second molecular binder in the plurality of second molecular binders after contacting the plurality of protein analytes with the plurality of second molecular binders.

[0282] In some embodiments, the linker is a bifunctional linker, and each respective second molecular binder in the plurality of second molecular binders is functionalized with the bifunctional linker using primary amine modification by a process comprising: modifying a primary amine of the respective second molecular binder, thereby obtaining a first attachment moiety comprising a modified primary amine on the respective second molecular binder, attaching the bifunctional linker to the respective second molecular binder, where the bifunctional linker comprises a second attachment moiety and a third attachment moiety, and where the second attachment moiety of the bifunctional linker binds to the first attachment moiety of the respective second molecular binder, and attaching the initiator tag to the bifunctional linker, where the initiator tag comprises a fourth attachment moiety that binds to the third attachment moiety.

[0283] In some embodiments, the initiator tag is attached to the respective second molecular binder using azide-alkyne cycloaddition. In some embodiments, the initiator tag is attached to the bifunctional linker using click chemistry or trans-cyclooctene (TCO) click chemistry. In some embodiments, a respective second molecular binder in the plurality of second molecular binder comprises one or more binding moieties. In some embodiments, the method further includes attaching the initiator tag to a binding moiety of the respective second molecular binder, where the initiator tag is specific to the binding moiety. In some embodiments, a binding moiety in the one or more binding moieties is an epitope.

[0284] In some embodiments, a respective second molecular binder in the plurality of second molecular binders comprises one or more initiator tags.

[0285] In some embodiments, each respective second molecular binder in the plurality of second molecular binders is the same. In some embodiments, each respective second molecular binder in the plurality of second molecular binders is not the same. For instance, in some embodiments, the plurality of second molecular binders comprises distinct types or classes of second molecular binders (e.g., different small molecules). In some embodiments, instances of the same molecular binder (e.g., the same small molecule) are tagged with the same DNA initiator tag, and / or the method does not differentiate betweenDB2 / 651687418.2 76Attorney Docket No.: 139643-5001-WOmolecular binders (e.g., small molecules) that bind to protein analytes. In some embodiments, all molecular binders that bind to protein analytes (whether they are the same or different) are tagged with the same DNA initiator tag.

[0286] Alternatively or additionally, in some embodiments, the plurality of second molecular binders comprises a plurality of different second molecular binders (e.g., a plurality of different small molecules), where each respective second molecular binder in the plurality of different second molecular binders comprises a corresponding unique initiator tag comprising a DNA initiator sequence that is unique to the respective second molecular binder.

[0287] In some embodiments, the unique initiator tag is attached to each respective second molecular binder in the plurality of different second molecular binders prior to contacting the plurality of protein analytes with the plurality of second molecular binders. For instance, in some embodiments, instances of the same molecular binder e.g., the same small molecule) are tagged with the same DNA initiator tag, whereas different molecular binders (e.g., different small molecules) are uniquely tagged with DNA initiator tags having unique DNA sequences. In some such embodiments, the method differentiates between molecular binders (e.g., small molecules) that differentially bind to different subsets of protein analytes. For example, in some implementations, particular molecular binders, and / or protein analytes that bind to such molecular binders, are identifiable as tagged with the unique corresponding DNA initiator tag.

[0288] In some embodiments, the method further includes attaching, to each respective protein analyte in the plurality of protein analytes, an initiator tag comprising a DNA initiator sequence, where the DNA concatemer is appended to the DNA initiator sequence of the respective protein analyte using the DNA polymerase.

[0289] In some embodiments, the initiator tag is attached to the respective protein analyte using azide-alkyne cycloaddition. In some embodiments, the attaching the initiator tag to a respective protein analyte in the plurality of protein analytes comprises functionalizing the respective protein analyte using a linker comprising an attachment moiety for attachment with the initiator tag. In some embodiments, the initiator tag is attached to the linker using click chemistry or trans-cyclooctene (TCO) click chemistry. In some embodiments, a respective protein analyte in the plurality of protein analytes comprises one or more binding moieties, further comprising attaching the initiator tag to a binding moiety ofDB2 / 651687418.2 77Attorney Docket No.: 139643-5001-WOthe respective protein analyte in the plurality of protein analytes, where the initiator tag is specific to the binding moiety. In some embodiments, a binding moiety in the one or more binding moieties is an epitope. In some embodiments, a respective protein analyte in the plurality of analytes comprises one or more initiator tags.

[0290] In some embodiments, each respective second molecular binder in the plurality of second molecular binders comprises a corresponding nucleic acid barcode sequence that is unique to the respective second molecular binder.

[0291] In some embodiments, each respective first molecular binder in the plurality of first molecular binders comprises a corresponding nucleic acid barcode sequence that is unique to the respective first molecular binder.

[0292] In some embodiments, a respective first molecular binder in the plurality of first molecular binders has a dissociation constant (KD) of at least 1 micromolar (pM). In some embodiments, a respective first molecular binder in the plurality of first molecular binders binds to one or more protein interaction features in a plurality of protein interaction features. In some embodiments, a respective first molecular binder in the plurality of first molecular binders binds to two or more protein interaction features in the plurality of protein interaction features. In some embodiments, a respective molecular binder in the plurality of molecular binders binds to two or more protein analytes in the plurality of protein analytes.

[0293] Compositions

[0294] Another aspect of the present disclosure provides a composition, comprising: a DNA polymerase; a plurality of dNTPs; one or more protein analytes, where each respective protein analyte in the one or more protein analytes is attached to an initiator tag comprising a DNA initiator sequence; and a plurality of molecular binders, where: each protein analyte in the plurality of protein analytes forms a binder-analyte interaction with at least one molecular binder in the plurality of molecular binders, each respective molecular binder in the plurality of molecular binders comprises a corresponding nucleic acid barcode sequence that is unique to the respective molecular binder, and one or more molecular binders in the plurality of molecular binders interacts nonspecifically with each protein analyte in a corresponding subset of the plurality of protein analytes (e.g., forms binderanalyte interactions with each protein analyte in the subset of protein analytes).

[0295] In some embodiments, a corresponding nucleic acid barcode sequence for a first molecular binder in the plurality of molecular binders comprises a first nucleotide DB2 / 651687418.2 78Attorney Docket No.: 139643-5001-WOsequence that is reverse complementary to a second nucleotide sequence of at least one other molecular binder in the plurality of molecular binders.

[0296] In some embodiments, a first nucleic acid barcode sequence for a first molecular binder in the plurality of molecular binders comprises a first portion positioned at a 5’ end of the nucleic acid barcode sequence, a second nucleic acid barcode sequence for a second molecular binder in the plurality of molecular binders comprises a second portion positioned at a 3’ end of the nucleic acid barcode sequence, and the first portion of the first nucleic acid barcode sequence hybridizes to a reverse complement of the second portion of the second nucleic acid barcode sequence.

[0297] In some embodiments, for each respective molecular binder in the plurality of molecular binders, the corresponding nucleic acid barcode sequence further comprises a first constant portion and a second constant portion, where the first constant portion is positioned at a 5’ end of the nucleic acid barcode sequence, the second constant portion is positioned at a 3’ end of the nucleic acid barcode sequence, and the first constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the second constant portion.

[0298] In some embodiments, the plurality of molecular binders comprises a first subset of molecular binders and a second subset of molecular binders; for each molecular binder in the first subset of molecular binders, the corresponding nucleic acid barcode sequence comprises a first constant portion and a second constant portion, where the first constant portion is positioned at a 5’ end of the nucleic acid barcode sequence, and the second constant portion is positioned at a 3’ end of the nucleic acid barcode sequence; for each molecular binder in the second subset of molecular binders, the corresponding nucleic acid barcode sequence comprises a third constant portion and a fourth constant portion, where the third constant portion is positioned at a 5’ end of the nucleic acid barcode sequence, and the fourth constant portion is positioned at a 3’ end of the nucleic acid barcode sequence; the first constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the fourth constant portion; and the second constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the third constant portion.

[0299] Another aspect of the present disclosure provides a composition, comprising: a DNA polymerase; a plurality of dNTPs; one or more protein analytes, where each respective protein analyte in the one or more protein analytes comprises a corresponding nucleic acid barcode sequence that is unique to the respective protein analyte; and a plurality of molecularDB2 / 651687418.2 79Attorney Docket No.: 139643-5001-WObinders, where: each protein analyte in the plurality of protein analytes forms a binder-analyte interaction with at least one molecular binder in the plurality of molecular binders, each respective molecular binder in the one or more molecular binders is attached to an initiator tag comprising a DNA initiator sequence, and one or more molecular binders in the plurality of molecular binders interacts nonspecifically with each protein analyte in a corresponding subset of the plurality of protein analytes (e.g., forms binder-analyte interactions with each protein analyte in the subset of protein analytes).

[0300] Any of the embodiments disclosed elsewhere herein for identifying and / or detecting a protein analyte, including samples, protein analytes, initiator tags, binding moieties, molecular binders, barcodes, reactions, DNA concatemers, interaction data structures, protein analyte indications, classification, and interaction models, are similarly contemplated for use in the presently disclosed compositions, including any modifications or substitutions thereof, as will be apparent to one skilled in the art.

[0301] Kits

[0302] Another aspect of the present disclosure provides a kit, comprising: a DNA polymerase; a plurality of dNTPs; and a plurality of molecular binders, where: each protein analyte in a plurality of protein analytes forms a binder-analyte interaction with at least one molecular binder in the plurality of molecular binders, each respective molecular binder in the plurality of molecular binders comprises a corresponding nucleic acid barcode sequence that is unique to the respective molecular binder, and one or more molecular binders in the plurality of molecular binders interacts nonspecifically with each protein analyte in a corresponding subset of the plurality of protein analytes (e.g., forms binder-analyte interactions with each protein analyte in the subset of protein analytes).

[0303] In some embodiments, a corresponding nucleic acid barcode sequence for a first molecular binder in the plurality of molecular binders comprises a first nucleotide sequence that is reverse complementary to a second nucleotide sequence of at least one other molecular binder in the plurality of molecular binders.

[0304] In some embodiments, a first nucleic acid barcode sequence for a first molecular binder in the plurality of molecular binders comprises a first portion positioned at a 5’ end of the nucleic acid barcode sequence, a second nucleic acid barcode sequence for a second molecular binder in the plurality of molecular binders comprises a second portion positioned at a 3’ end of the nucleic acid barcode sequence, and the first portion of the first DB2 / 651687418.2 80Attorney Docket No.: 139643-5001-WOnucleic acid barcode sequence hybridizes to a reverse complement of the second portion of the second nucleic acid barcode sequence.

[0305] In some embodiments, for each respective molecular binder in the plurality of molecular binders, the corresponding nucleic acid barcode sequence further comprises a first constant portion and a second constant portion, where the first constant portion is positioned at a 5’ end of the nucleic acid barcode sequence, the second constant portion is positioned at a 3’ end of the nucleic acid barcode sequence, and the first constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the second constant portion.

[0306] In some embodiments, the plurality of molecular binders comprises a first subset of molecular binders and a second subset of molecular binders; for each molecular binder in the first subset of molecular binders, the corresponding nucleic acid barcode sequence comprises a first constant portion and a second constant portion, wherein the first constant portion is positioned at a 5’ end of the nucleic acid barcode sequence, and the second constant portion is positioned at a 3’ end of the nucleic acid barcode sequence; for each molecular binder in the second subset of molecular binders, the corresponding nucleic acid barcode sequence comprises a third constant portion and a fourth constant portion, wherein the third constant portion is positioned at a 5’ end of the nucleic acid barcode sequence, and the fourth constant portion is positioned at a 3’ end of the nucleic acid barcode sequence; the first constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the fourth constant portion; and the second constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the third constant portion.

[0307] Another aspect of the present disclosure provides a kit, comprising: a DNA polymerase; a plurality of dNTPs; a plurality of nucleic acid barcode sequences; and a plurality of molecular binders, where: each protein analyte in a plurality of protein analytes forms a binder-analyte interaction with at least one molecular binder in the plurality of molecular binders, each respective molecular binder in the plurality of molecular binders is attached to an initiator tag comprising a DNA initiator sequence, and one or more molecular binders in the plurality of molecular binders interacts nonspecifically with each protein analyte in a corresponding subset of the plurality of protein analytes (e.g., forms binderanalyte interactions with each protein analyte in the subset of protein analytes).

[0308] Any of the embodiments disclosed elsewhere herein for identifying and / or detecting a protein analyte, including samples, protein analytes, initiator tags, bindingDB2 / 651687418.2 81Attorney Docket No.: 139643-5001-WOmoieties, molecular binders, barcodes, reactions, DNA concatemers, interaction data structures, protein analyte indications, classification, and interaction models, are similarly contemplated for use in the presently disclosed kits, including any modifications or substitutions thereof, as will be apparent to one skilled in the art.

[0309] Additional Example Embodiments.

[0310] Another aspect of the present disclosure includes a system, including a memory; one or more processors; and one or more modules stored in the memory and configured for execution by the one or more processors, the one or more modules including instructions for performing any of the methods disclosed above.

[0311] Another aspect of the present disclosure includes a non-transitory computer readable storage medium, the non-transitory computer readable storage medium storing one or more programs for execution by one or more processors of a computer system, the one or more computer programs including instructions for performing any of the methods disclosed above.

[0312] EXAMPLES.

[0313] Example 1. Generation of nucleic acid barcodes for use in concatemer extension.

[0314] In example assays, biomolecular proximity detection was performed in a model DNA-DNA system. Here, the goal was to evaluate whether biomolecular proximity could be detected in a model system in the absence of protein analytes. In this DNA-DNA model system, extension reactions including nucleic acid barcode molecules (e.g., comprising nucleic acid barcode sequences) were performed to evaluate the ability to generate DNA concatemers. The biomolecular proximity detection model can then be applied to proteinprotein models (e.g., protein analytes comprising initiator tags and molecular binders comprising nucleic acid barcode sequences) to perform protein-protein proximity detection.

[0315] The biomolecular proximity detection assay was set up by bringing DNA initiators (e.g., initiator tags comprising DNA initiator sequences) into proximity with nucleic acid barcode molecules by DNA-DNA base pairing. In this example, nucleic acid barcode molecules comprised pseudo-lariat structures (e.g., “tennis racket”), as illustrated in FIG. 7. For instance, as depicted in FIG. 7, in some implementations, a nucleic acid barcode molecule included a double-stranded “handle” region, an optional variable length region, and a single-stranded region including a first constant region (e.g., “Constant Region A”), a DNA DB2 / 651687418.2 82Attorney Docket No.: 139643-5001-WObarcode sequence (e.g., “Barcode Region”), and a second constant region (e.g., “Constant Region B”). All or a portion of the first constant region has a nucleic acid sequence that is the same as at least a portion of the second constant region such that extension of one of the first constant region and the second constant region by nucleic acid synthesis generates a reverse complementary sequence that can be hybridized to by the other of the first constant region and the second constant region. Referring again to the schematic in FIGS. 4A-C, in some implementations, extension of a DNA concatemer (e.g., by appending to an initiator tag) synthesizes a reverse complementary sequence comprising, at least, the first constant region, the barcode region, and the second constant region, thereby forming a new single-stranded hybridization region that can be hybridized to by a subsequent nucleic acid barcode molecule after each round of extension.

[0316] The length of the “variable length” region and / or the double-stranded region was tunable, such that the distance of the initiator tag from the extendable portion of the barcode molecule (e.g., including the constant regions and the barcode region) could be modified. In this way, it was possible to manipulate the interaction strength and / or dwell time of each nucleic acid barcode molecule during each round of extension and evaluate the effect of barcode length and distance on concatemer formation. In particular, a library of different nucleic acid barcode molecule sequences was used, where the variable length of each barcode molecule stem (e.g., “handle” of a tennis racket) was systematically increased in size. Each of these oligonucleotides had its own barcode sequence (shown in FIG. 7 as “Barcode Region”) so that it was possible to read out the different effects of a pool of these barcode molecules in a single biochemistry and sequencing experiment.

[0317] First, two samples of pseudo-lariat (e.g., “tennis racket”) barcode molecules were made. A first sample of barcode molecules included a positive control oligonucleotide (“oAT_751”) and a second sample of barcode molecules included pooled nucleic acid barcode molecules generated using the library of barcode molecule sequences (“oligo pool”). In both samples, each barcode molecule included a first constant region including 6-8 DNA bases and a second constant region including 6-8 DNA bases, where at least a portion of the first constant region was the same as at least a portion of the second constant region. The first and second constant regions flanked a barcode region including a variable span of 10 DNA bases, as illustrated in FIG. 7. Furthermore, the unpaired region (e.g., single-stranded or loop region) including the first constant region, the second constant region, and the barcode region was flanked by two pairing sequences (e.g., hairpin sequences) that were reverseDB2 / 651687418.2 83Attorney Docket No.: 139643-5001-WOcomplementary such that they could be hybridized to form a double-stranded “handle” region.

[0318] For the positive control oligonucleotide oAT_751 and each barcode molecule in the oligo pool, nucleic acid barcode sequences were inserted into the 10-base variable barcode region from a plurality of randomly generated barcode sequences. Optionally, barcode molecules were further filtered by matching each barcode sequence with a next closest barcode sequence in the plurality of randomly generated barcode sequences, thus obtaining a plurality of barcode sequence pairs, and removing one of the barcode sequences from each pair in the plurality of barcode sequence pairs. In this way, the uniqueness of barcode sequences in the resulting pool of barcode molecules could be enhanced.

[0319] The resulting barcode molecule designs were then synthesized, annealed to form the pseudo-lariat structure, and optionally crosslinked using psoralen. Here, a first aliquot of the synthesized barcode molecules were crosslinked with psoralen, and a second aliquot of the synthesized barcode molecules were not treated with psoralen. Samples were then cleaned and 5 pL of the clean-up material was run at 5 ng / pL (25 ng) on a 15% urea-PAGE gel. The sample was visualized by staining the gel with a nucleic acid gel stain (SYBR Gold) for 10 minutes and then imaging with a Cy2 channel. The gel showed robust formation of barcode molecules in both the positive control oligonucleotide oAT_751 and the oligo pool, both with and without psoralen crosslinking (data not shown). Additionally, no intensity difference was observed between the positive control oligonucleotide oAT_751 and the oligo pool, either with or without psoralen treatment.

[0320] Example 2. Concatemer extension using nucleic acid barcodes.

[0321] Next, the barcode molecules generated in Example 1 were used to assess concatemer extension using pseudo-lariat barcode molecules (e.g., tennis rackets) and a biotinylated initiator (e.g., an initiator tag comprising a DNA initiator sequence and a biotin affinity moiety).

[0322] A concatemer extension reaction was performed by incubating the barcode molecules generated in Example 1 with a biotinylated initiator (“oAT_707”) and a reaction mix including the following enzymes and reagents: rCutSmart buffer at a reaction concentration of 1 pM, a deoxynucleotide (dNTP) solution mix at a reaction concentration of 1 pM, and Bst 3.0 DNA polymerase at a reaction concentration of 400 pM. Concatemer generation was assayed for each of the positive control oligonucleotide oAT_751 and the DB2 / 651687418.2 84Attorney Docket No.: 139643-5001-WOpooled oligos, with and without psoralen crosslinking, at two final reaction concentrations of 0.2 pM and 2 pM (e.g., 8 experimental conditions). Concatemer generation was performed at an ambient temperature of 37 °C for 16 hours and inactivated at 80 °C for 5 hours. A urea-PAGE gel was then used to assess the generation of nucleic acid concatemers under each of the experimental conditions. 2 pL of each reaction was run on a 15% urea-PAGE gel without staining and imaged with a Cy3 channel.

[0323] FIG. 8 illustrates the concatemers generated using the DNA-DNA model system, in which the nucleic acid barcode molecules were appended to biotinylated initiators during the extension reaction. The positive control barcode molecule oAT_751 resulted in greater than 4 concatemer extensions when incubated at a concentration of 0.2 pM, although not at a concentration of 2 pM. Additionally, the pooled oligo sample showed greater than 2 concatemer extensions when incubated at a concentration of 2 pM. Similar results were observed with and without psoralen crosslinking of the barcode molecules. In general, concatemer formation was observed across all experimental conditions (e.g., positive control oligonucleotide and pooled oligos, with and without psoralen crosslinking, and at both barcode molecule concentrations).

[0324] Example 3. Purification and sequencing of nucleic acid concatemers.

[0325] Next, library preparation for sequencing was performed using streptavidin affinity purification of the nucleic acid concatemers generated in Example 2.

[0326] Streptavidin affinity purification was used to capture the biotinylated initiators that were extended to form the nucleic acid concatemers onto beads. Briefly, 50 pL of the streptavidin bead slurry was pre-washed with 200 pL Tris-EDTA (TE) buffer. The supernatant was removed and discarded, and the beads were resuspended into 90 pL TE. 10 pL of each sample was aliquoted into tubes and rotated at 37 °C for 1 hour. Next, samples were incubated with DNA ligase for on-bead ligation, and further incubated with a PCR mix for on-bead PCR. Samples were cleaned up and the libraries were sequenced using nanopore sequencing.

[0327] FIG. 9 A illustrates example plots showing read length, number of barcodes, and distance between constant regions for concatemers generated for the positive control oligonucleotide oAT_751 after sequencing. For instance, FIG. 9A shows that read length for concatemers generated using the positive control oligonucleotide peaks at 300-400 base pairs (“bp”) (top panel), number of barcodes for concatemers peaks at approximately 15 barcodes DB2 / 651687418.2 85Attorney Docket No.: 139643-5001-WO(middle panel), and a clear spacing of 9 base pairs between constant regions (bottom panel), corresponding to the variable barcode region flanked by the first and second constant regions of each barcode molecule.

[0328] FIG. 9B further illustrates example plots showing read length, number of barcodes, and distance between constant regions for concatemers generated from the pooled library of nucleic acid barcodes (“pooled oligos”). For instance, FIG. 9B shows that read length for concatemers generated using the pooled oligos peaks at approximately 200 base pairs (top panel), the number of barcodes for concatemers peaks at approximately 5 barcodes (middle panel), and illustrates a clear spacing of 9 base pairs between constant regions (bottom panel). The spacing between constant regions is characteristic of concatemeric reads. FIGS. 9A-B illustrate clear evidence of concatemer formation in both sample types (positive control oligonucleotide and pooled oligonucleotide library).

[0329] As noted above in Example 1, the tunable length of the “variable length” region and / or the double-stranded region of each barcode molecule allowed for modulation of the distance between the initiator tag and the extension region of the barcode molecule (e.g., including the constant regions and the barcode region). In this way, it was possible to manipulate the interaction strength and / or dwell time of each nucleic acid barcode molecule during each round of extension (e.g., by increasing the length to reduce interaction strength and / or dwell time, and / or by decreasing the length to increase interaction strength and / or dwell time).

[0330] The sequencing data allowed for quantification of binding and extension events using the following process: first, barcode frequency within concatemers was counted. Next, the barcode molecule corresponding to each barcode sequence was determined (e.g., where each barcode molecule comprises a corresponding unique barcode sequence that allows for the barcode molecule to be identified). Since different barcode molecules had proximity stems of varying length (e.g., variable regions and / or hairpin regions), the barcode sequences contained within each concatemer could be used to determine identity of each barcode molecule and the AG of the biomolecular proximity between the barcode molecule and the initiator.

[0331] Generally, the AG (interchangeably, “delta G”) of the binding and extension event between a barcode molecule and the initiator (or an extended concatemer thereof) is a measure of the change in Gibbs free energy that provides a quantitative measure of howDB2 / 651687418.2 86Attorney Docket No.: 139643-5001-WOenergetically favorable the reaction is under certain conditions. In this case, the AG is a measure of the effect of the variable length of hairpin regions of each barcode molecule, or the biomolecular proximity of each barcode molecule and the initiator or concatemer, on the favorability of the extension reaction, as illustrated in FIG. 10.

[0332] The plot in FIG. 10 shows the correspondence between the number of barcode observations in a concatemer and the AG predictions of biomolecular proximity between barcodes and an initiator tag. Thus, more energetically favorable interactions (e.g., more negative AG) result in greater observations of barcode molecule binding and extension, while less energetically favorable interactions result in fewer observations.

[0333] The DNA-DNA model described in Examples 1-3 provides evidence of a clear signal for biomolecular proximity detection using concatemer formation from barcode molecules. Advantageously, the signal response in these experiments spanned a range of approximately 1 pM to 10 pM, which is useful for in vitro wet-lab experimentation.

[0334] Example 4. Generation of initiator-tagged protein analytes and DNA-barcoded molecular binders.

[0335] Next, assays were performed to evaluate concatemer formation in a proteinprotein interaction model. As an initial step, DNA constructs encoding protein libraries were generated and used to express initiator-tagged protein analytes and barcoded molecular binder pools.

[0336] First, an empty plasmid backbone was constructed and cloned, which included a protein coding region and an RNA expressing region. The protein coding region included an empty cloning site into which a variable molecular binder sequence could be inserted, while the RNA expressing region included an empty cloning site into which a variable nucleic acid sequence could be inserted.

[0337] The protein coding region optionally included additional domains, including one or more affinity tags, solubility tags, RNA binding domains, protease sites, and / or covalent tags. The RNA expressing region also optionally included additional domains, including one or more promoters, one or more RNA binding domains, one or more detectable tags, and / or one or more terminator domains.

[0338] Next, for barcoded molecular binders, a library of molecular binders (e.g., proteins) was cloned into the empty cloning site of the protein coding region and a library of nucleic acid barcode sequences was cloned into the empty cloning site of the RNA expressing DB2 / 651687418.2 87Attorney Docket No.: 139643-5001-WOregion using a two-stage cloning approach for multi-fragment DNA constructs (e.g., two-step Golden Gate assembly). The library of molecular binders included approximately 7,000 molecular binders (“lib023”), and the library of nucleic acid barcode sequences included either circular barcode molecules or pseudo-lariat barcode molecules (e.g., “tennis rackets”). During this step, an additional library of positive control spike-in peptides (“lib024”) with ground-truth affinity constants (e.g., dissociation constant (Kd), association rate constant (Kon), and dissociation rate constant (Koir)) was mixed in to the cloning mixture at 5%.

[0339] For initiator-tagged analyte proteins, a library of 700 analytes (“lib016”) was cloned into the empty cloning site of the protein coding region, and a DNA initiator sequence was cloned into the empty cloning site of the RNA expressing region using the two-stage cloning approach. During this step, the library of positive control spike-in peptides (“lib024”) was also mixed into the cloning mixture at 5%.

[0340] The assembled plasmids were transformed and harvested, and PCR was performed to amplify relevant regions from the plasmids. The success of assembly was evaluated by nanopore sequencing. Binders comprising the circular DNA barcode molecules had a successful assembly rate of 75%, while binders comprising the pseudo-lariat DNA barcode molecules had a successful assembly rate of 59%. Protein analytes comprising initiator tags had a successful assembly rate of 66%. These results indicated that both the initiator-tagged protein analytes and barcoded molecular binders could be constructed and assembled successfully.

[0341] Additionally, sequencing of the assembled plasmids allowed the different protein molecular binders to be associated with their corresponding DNA barcode sequence for use in subsequent protein writing and identification. For instance, sequencing of the assembled plasmids revealed the following counts of total and unique variants for two aliquots each of the circular barcode molecules, pseudo-lariat barcode molecules, and protein analytes: circles (aliquot A): total variants (filtered): 19782, total variants (unique): 2588; circle (aliquot B): total variants (filtered): 4887, total variants (unique): 1424; pseudo-lariat (aliquot A): total variants (filtered): 40604, total variants (unique): 3184; pseudo-lariat (aliquot B): total variants (filtered): 9014, total variants (unique): 1858; protein analytes (aliquot A): total variants (filtered): 50121, total variants (unique): 714; and protein analytes (aliquot B): total variants (filtered): 2117, total variants (unique): 462.DB2 / 651687418.2 88Attorney Docket No.: 139643-5001-WO

[0342] Protein and RNA from each plasmid was expressed in BL21DE3 bacteria and purified with nickel affinity (e.g., for His8 tag on protein). Each protein was expected to copurify with a corresponding RNA, which could be visualized with a fluorescent detectable tag located in the RNA expressing region of the plasmid construct.

[0343] Next, reverse transcription was performed to convert the barcode region from the RNA expressing region into cDNA. By priming the expressed protein product with a DNA oligo chemically modified to react with a covalent tag domain in the protein, a covalent link was formed between the cDNA and the protein product. The DNA oligo primer was also fluorescently tagged enabling easy visualization of the conjugate products.

[0344] The protein-cDNA conjugate product was then treated with RNase and Tev protease, isolated by performing affinity capture with an affinity tag expressed on the protein (e.g., a FLAG tag), and washed to remove N-terminal protein fragments. Optionally, for barcoded molecular binders comprising circular barcode molecules, the circular DNA was ligated. The output of these processes included covalent DNA-protein products with minimal protein context.

[0345] For each of the barcoded molecular binders and initiator-tagged protein analytes, 5 pL of each sample type (e.g., input, flow-through, washes 1 and 6, and elution) was mixed with 5 pL of a fluorescent-compatible sample buffer. The mixture was heated at 95 °C for 3 minutes and run on a 10-20% Tricine-PAGE gel. The gel was imaged first with Cy2 channel without staining (data not shown), then stained with a protein stain (e.g. , Coomassie) and imaged with a visible channel, as depicted in FIG. 11. The gel showcases the ability to generate protein analytes comprising initiator tags (e.g, comprising a DNA initiator sequence for appending concatemers), as well as the ability to generate molecular binders comprising nucleic acid barcode sequences, which can be used to form concatemers by appending to the initiator tag via extension.

[0346] Example 5. Extension reactions using protein analyte libraries and molecular binders comprising nucleic acid barcode sequences.

[0347] Assays were next performed to evaluate extension reactions using initiator-tagged protein analytes and DNA barcoded molecular binders.

[0348] The initiator-tagged protein analytes and molecular binders comprising pseudo-lariat barcode molecules were obtained as described in Example 4. Concatemer extension reactions were performed for combinations of: (i) an initiator tag sequence control DB2 / 651687418.2 89Attorney Docket No.: 139643-5001-WO(“oAT_977”), (ii) pooled protein analytes comprising an initiator tag sequence (“analytes” or “e282-bs”), (iii) pseudo-lariat barcode controls (“oAT_1058”), and (iv) molecular binders comprising pseudo-lariat barcodes (“binders”).

[0349] The concatemer extension reactions were performed by incubating the initiator controls, barcode controls, protein analytes, and molecular binders, alone and in combination, with a reaction mix including the following enzymes and reagents: b44 buffer at a reaction concentration of 1 pM, a deoxynucleotide (dNTP) solution mix at a reaction concentration of 1 pM, and Bst 3.0 DNA polymerase at a reaction concentration of 400 pM. Concatemer generation was performed at an ambient temperature of 37 °C for 16 hours and inactivated at 80 °C for 5 hours. 1 pL of each reaction was mixed with 1 pL of a 2X RNA loading dye and run on a 15% urea-PAGE gel without staining and imaged with a Cy2 channel to assess the generation of nucleic acid concatemers under each of the experimental conditions.

[0350] FIG. 12 illustrates the concatemers generated from the reactions including the initiator tag control alone (oAT_977), barcode control alone (oAT_1058), initiator tag and barcode (oAT_977 + oAT_1058), protein analyte with initiator tag (analytes), molecular binders with nucleic acid barcodes (binder), and various combinations thereof. These results show formation of concatemers at least for extension reactions including initiator tags and barcode controls (oAT_977 + oAT_1058) and for extension reactions including initiator tags and binders comprising nucleic acid barcodes (oAT_977 + binders).

[0351] Next, the concatemers were processed for sequencing using streptavidin purification, on-bead ligation, and on-bead PCR, as described above in Example 3. The readout for the amplified sequencing product was visualized on a 4% E-Gel with no stain, as shown in FIG. 13. These results illustrate that concatemers were additionally formed when pooled analytes comprising initiator tags (e282_bs) were incubated with nucleic acid barcodes (“oAT_1058 racket”). In particular, concatemers of +2 and +3 (e.g., having 2 to 3 barcode extensions) could be observed when the control barcodes were incubated with the pooled, initiator-tagged protein analytes.

[0352] FIGS. 14A and 14B illustrate the read length, number of barcodes, and distance between constant regions of the samples shown in FIGS. 12 and 13 after sequencing. FIG. 14A indicates that most of the nucleic acids were not extended during the concatemer extension reaction. However, as illustrated in FIG. 14B, filtering the sequencing reads to those that were extended to +2 and +3 concatemers (e.g., as shown in FIG. 13) showsDB2 / 651687418.2 90Attorney Docket No.: 139643-5001-WOexpected patterns in read length, barcode number, and distance between constant regions that reflect correct formation of concatemers, similar to those shown in FIGS. 9A-B.

[0353] Example 6. Extension reactions using initiator-tagged protein analytes and molecular binders comprising nucleic acid barcode sequences generate concatemers .

[0354] Extension reactions were performed using analytes conjugated to initiator DNA and binders conjugated to nucleic acid barcodes generated as described in Examples 4 and 5. Extension reactions were performed using the initiator-tagged analytes and binders conjugated to either circular barcodes or pseudo-lariat barcodes, as illustrated in FIG. 15 A. Additional control reactions were performed using control initiator oligonucleotides, control circular barcode oligonucleotides, and control pseudo-lariat barcode oligonucleotides, illustrated in FIG. 15 A. FIG. 15B depicts a gel readout of products generated from the extension reactions, in which the presence of product is visible for lanes corresponding to analytes and binders (e.g., initiator and Circle, initiator and Racket) but not in lanes in which either the analytes or the nucleic acid barcodes are not present.

[0355] An in-solution PCR was performed to amplify the products generated from each extension reaction detailed in FIGS. 15A-B. PCR products were cleaned and concentrations were measured using a microvolume UV-Vis spectrophotometer (Nanodrop). Results are shown below in Table 1, including the concentrations for replicates of the analyte conjugated to DNA initiator (Initiator) combined with either binders conjugated to circular barcodes (Circle) or binders conjugated to pseudo-lariat barcodes (Racket).

[0356] Table 1. Concentration of PCR- Amplified Extension Reaction ProductsDB2 / 651687418.2 91Attorney Docket No.: 139643-5001-WO

[0357] FIGS. 16A-B further provide gel readouts of products from the second PCR reaction before and after clean-up, illustrating the presence of material for sequencing.

[0358] The amplified material was sequenced and analyzed for read length, number of barcodes, and distance between constant regions. FIG. 17 illustrates the plots obtained for extension reactions generated when using binders conjugated to circular barcodes. The plots for read length and number of barcodes demonstrate that barcode extension occurred.Additionally, the distance between constant regions exhibited in the sequencing data was consistent with the concatemers observed in FIGS. 9A-B.

[0359] FIG. 18A illustrates the plots obtained for extension reactions generated when using binders conjugated to pseudo-lariat barcodes. As in FIG. 17, the plots for read length and number of barcodes demonstrate that barcode extension occurred, and the distance between constant regions exhibited in the sequencing data was consistent with concatemer formation. Filtering the pseudo-lariat barcode sequencing data for sequence reads that were extended in FIG. 18B, the expansion in read length and barcode integration could be visualized at greater resolution, in that extension by up to 7 barcodes could be observed. FIG.19 provides example schematics of the actual sequences obtained for extension reaction products after sequencing the amplified products from FIG. 18B (e282_hy and e828_hx). FIG. 19 illustrates, in particular, evidence of formation of concatemers having extensions of +1, +2, and +3.

[0360] The data described in Examples 1-6 illustrate the ability to successfully clone a plasmid backbone, successfully assemble plasmid libraries for initiator-tagged analytes and barcoded molecular binders, express and purify protein conjugates, and generate concatemers using both DNA-DNA models and protein-protein interaction models that leverage biomolecular proximity-based extension reactions. The methods and compositions disclosed herein can be successfully used, among other applications, to generate concatemers that can be sequenced for protein barcode writing, readout, and identification.CONCLUSION

[0361] The foregoing description, for purposes of explanation, has been described with reference to specific implementations. However, the illustrative discussions above are not intended to be exhaustive or to limit the implementations to the precise forms disclosed. Many modifications and variations are possible in view of the above teachings. TheDB2 / 651687418.2 92Attorney Docket No.: 139643-5001-WOimplementations were chosen and described in order to best explain the principles and their practical applications, to thereby enable others skilled in the art to best utilize the implementations and various implementations with various modifications as are suited to the particular use contemplated.DB2 / 651687418.2 93

Claims

Attorney Docket No.: 139643-5001-WOWhat is claimed is:

1. A method for identifying protein analytes, comprising:A) obtaining a plurality of protein analytes from a sample of a subject;B) contacting the plurality of protein analytes with a plurality of molecular binders and a DNA polymerase, thereby forming a plurality of binder-analyte interactions, wherein:each respective molecular binder in the plurality of molecular binders comprises a corresponding nucleic acid barcode sequence that is unique to the respective molecular binder, andone or more molecular binders in the plurality of molecular binders interacts nonspecifically with each protein analyte in a corresponding subset of the plurality of protein analytes;C) generating, for a first protein analyte in the plurality of protein analytes, a DNA concatemer comprising a corresponding set of DNA barcode sequences, wherein each DNA barcode sequence in the corresponding set of DNA barcode sequences:i) corresponds to a nucleic acid barcode sequence of a respective molecular binder, in the plurality of molecular binders, that forms a binder-analyte interaction with the first protein analyte, andii) is appended to the first protein analyte using the DNA polymerase, responsive to formation of the binder-analyte interaction between the respective molecular binder and the first protein analyte, thereby causing a spatial order of DNA barcode sequences in the corresponding set of DNA barcode sequences to correspond to a temporal order of formation of binder-analyte interactions between respective molecular binders, in the plurality of molecular binders, and the first protein analyte; D) determining an interaction data structure for the first protein analyte based on a sequencing of the DNA concatemer, wherein the interaction data structure comprises an indication of a presence or abundance of each respective molecular binder in the plurality of molecular binders that formed a binder-analyte interaction with the first protein analyte; and E) receiving, responsive to inputting the interaction data structure to a model, an indication of an identity for the first protein analyte, as output from the model.

2. The method of claim 1, wherein the sample of the subject comprises one or more cells.DB2 / 651687418.2 94Attorney Docket No.: 139643-5001-WO3. The method of claim 1 or 2, wherein obtaining the plurality of protein analytes comprises generating a cell lysate from the sample of the subject.

4. The method of any one of claims 1-3, wherein the sample of the subject is a solid tissue sample or a liquid biopsy sample.

5. The method of any one of claims 1-4, wherein the sample of the subject comprises blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid.

6. The method of any one of claims 1-5, wherein the sample of the subject comprises a target condition.

7. The method of claim 6, wherein the target condition is a healthy condition or a disease condition.

8. The method of claim 6, wherein the target condition is exposure to a chemical compound.

9. The method of claim 6, wherein the target condition is a dysregulation of a biological intermediate in the sample of the subject.

10. The method of claim 9, wherein the biological intermediate is an enzyme, a kinase, a signaling molecule, a transcription factor, a hormone, a small molecule, a microRNA, a regulatory RNA, or a cytokine, and the dysregulation comprises overexpression, overproduction, repression, or silencing.

11. The method of any one of claims 1-10, wherein a respective protein analyte in the plurality of protein analytes is an enzyme, a kinase, a protease, an antibody, or a receptor.DB2 / 651687418.2 95Attorney Docket No.: 139643-5001-WO12. The method of any one of claims 1-11, wherein a respective protein analyte in the plurality of protein analytes is a protein-protein complex.

13. The method of any one of claims 1-12, wherein a respective protein analyte in the plurality of protein analytes comprises a post-translational modification.

14. The method of any one of claims 1-13, wherein the plurality of protein analytes comprises at least 100 protein analytes.

15. The method of any one of claims 1-14, wherein a respective protein analyte in the plurality of protein analytes comprises a corresponding set of protein interaction features in a plurality of protein interaction features.

16. The method of claim 15, wherein a protein interaction feature in the plurality of protein interaction features is selected from the group consisting of: hydrogen bond donors, hydrogen bond acceptors, hydroxyl groups, positively charged atoms, negatively charged atoms, aromatic rings, and aliphatic hydrophobic groups.

17. The method of any one of claims 1-6, wherein each respective protein analyte in the plurality of protein analytes is attached to an initiator tag comprising a DNA initiator sequence, and wherein the DNA concatemer is appended to the DNA initiator sequence of the first protein analyte using the DNA polymerase.

18. The method of claim 17, further comprising attaching the initiator tag to the first protein analyte.

19. The method of claim 17 or 18, wherein the attaching the initiator tag to a respective protein analyte in the plurality of protein analytes comprises functionalizing the respective protein analyte using a linker that attaches the respective protein analyte to the initiator tag.

20. The method of claim 19, wherein:DB2 / 651687418.2 96Attorney Docket No.: 139643-5001-WOthe linker is a bifunctional linker, andthe respective protein analyte is functionalized with the bifunctional linker using primary amine modification by a process comprising:modifying a primary amine of the respective protein analyte, thereby obtaining a first attachment moiety comprising a modified primary amine on the respective protein analyte,attaching the bifunctional linker to the respective protein analyte, wherein the bifunctional linker comprises a second attachment moiety and a third attachment moiety, and wherein the second attachment moiety of the bifunctional linker binds to the first attachment moiety of the respective protein analyte, andattaching the initiator tag to the bifunctional linker, wherein the initiator tag comprises a fourth attachment moiety that binds to the third attachment moiety.

21. The method of claim 20, wherein the initiator tag is attached to the respective protein analyte using azide-alkyne cycloaddition.

22. The method of claim 20, wherein the initiator tag is attached to the bifunctional linker using click chemistry or trans-cyclooctene (TCO) click chemistry.

23. The method of any one of claims 17-22, wherein a respective protein analyte in the plurality of protein analytes comprises one or more binding moieties.

24. The method of claim 23, further comprising attaching the initiator tag to a binding moiety of the respective protein analyte in the plurality of protein analytes, wherein the initiator tag is specific to the binding moiety.

25. The method of claim 23 or 24, wherein:the plurality of protein analytes comprises a first class of protein analytes and a second class of protein analytes,each protein analyte in the first class of protein analytes comprises a set of binding moieties specific to the first class of protein analytes, andDB2 / 651687418.2 97Attorney Docket No.: 139643-5001-WOeach protein analyte in the second class of protein analytes comprises a set of binding moi eties specific to the second class of protein analytes, wherein the set of binding moi eties specific to the first class of protein analytes are different from the set of binding moi eties specific to the second class of protein analytes.

26. The method of claim 25, wherein the first protein is a member of the first class of protein analytes and comprises a binding moiety that is specific to the first class of protein analyte.

27. The method of any one of claims 23-26, wherein a binding moiety in the one or more binding moieties is an epitope.

28. The method of any one of claims 1-27, wherein a respective protein analyte in the plurality of analytes comprises one or more initiator tags.

29. The method of any one of claims 1-28, wherein the plurality of molecular binders comprises at least 10, 50, 100, or 500 molecular binders.

30. The method of any one of claims 1-29, wherein a respective molecular binder in the plurality of molecular binders has a length of at least 20 amino acids.

31. The method of any one of claims 1-30, wherein a respective molecular binder in the plurality of molecular binders has a molecular weight of at least 20 kDa.

32. The method of any one of claims 1-31, wherein each respective molecular binder in the plurality of molecular binders comprises a peptide.

33. The method of any one of claims 1-31, wherein each respective molecular binder in the plurality of molecular binders comprises a small molecule.

34. The method of any one of claims 1-31, wherein each respective molecular binder in the plurality of molecular binders comprises an antibody.DB2 / 651687418.2 98Attorney Docket No.: 139643-5001-WO35. The method of any one of claims 1-34, wherein each respective molecular binder satisfies any two or more rules, any three or more rules, or all four rules of the Lipinski's rule of Five: (i) not more than five hydrogen bond donors, (ii) not more than ten hydrogen bond acceptors, (iii) a molecular weight under 500 Daltons, and (iv) a LogP under 5.

36. The method of any one of claims 1-35, wherein each respective molecular binder in the plurality of molecular binders binds to one or more protein interaction features in a plurality of protein interaction features.

37. The method of any one of claims 1-36, wherein each molecular binder in the plurality of molecular binders is a weak affinity binder.

38. The method of any one of claims 1-37, wherein the molecular binder has a dissociation constant (KD) of at least 1 micromolar (pM).

39. The method of any one of claims 1-38, wherein each molecular binder in the plurality of molecular binders is a nonspecific binder.

40. The method of any one of claims 1-39, wherein a respective molecular binder in the plurality of molecular binders binds to two or more protein interaction features in the plurality of protein interaction features.

41. The method of any one of claims 1-40, wherein a respective molecular binder in the plurality of molecular binders binds to two or more protein analytes in the plurality of protein analytes.

42. The method of any one of claims 1-41, wherein at least the first protein analyte in the plurality of protein analytes binds to a corresponding set of molecular binders in the plurality of molecular binders.DB2 / 651687418.2 99Attorney Docket No.: 139643-5001-WO43. The method of claim 42, wherein the corresponding set of molecular binders comprises at least 2 molecular binders.

44. The method of claim 42 or 43, wherein the corresponding set of molecular binders for the first protein analyte collectively represent a corresponding set of protein interaction features for the first protein analyte.

45. The method of any one of claims 1-44, wherein the plurality of protein analytes comprises a first class of protein analytes and a second class of protein analytes, wherein: each protein analyte in the first class of protein analytes interacts with at least one molecular binder in the plurality of molecular binders, and each protein analyte in the second class of protein analytes does not interact with any molecular binders in the plurality of molecular binders.

46. The method of claim 45, wherein the first class of protein analytes is a target class and the second class of protein analytes is an off-target class.

47. The method of any one of claims 1-46, wherein the nucleic acid barcode sequence comprises DNA.

48. The method of claim 47, wherein the nucleic acid barcode sequence is ssDNA or dsDNA.

49. The method of any one of claims 1-48, wherein the nucleic acid barcode sequence comprises from 3 to 30 nucleotides.

50. The method of any one of claims 1-48, wherein a corresponding nucleic acid barcode sequence for a first molecular binder in the plurality of molecular binders comprises a first nucleotide sequence that is reverse complementary to a second nucleotide sequence of at least one other molecular binder in the plurality of molecular binders.

51. The method of any one of claims 1-50, wherein:DB2 / 651687418.2 100Attorney Docket No.: 139643-5001-WOa first nucleic acid barcode sequence for a first molecular binder in the plurality of molecular binders comprises a first portion positioned at a 5’ end of the nucleic acid barcode sequence,a second nucleic acid barcode sequence for a second molecular binder in the plurality of molecular binders comprises a second portion positioned at a 3’ end of the nucleic acid barcode sequence, andthe first portion of the first nucleic acid barcode sequence hybridizes to a reverse complement of the second portion of the second nucleic acid barcode sequence.

52. The method of any one of claims 1-51, wherein the nucleic acid barcode sequence comprises a first constant portion and a second constant portion, wherein the first constant portion is positioned at a 5’ end of the nucleic acid barcode sequence, the second constant portion is positioned at a 3’ end of the nucleic acid barcode sequence, and the first constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the second constant portion.

53. The method of any one of claims 1-51, wherein the plurality of molecular binders comprises a first subset of molecular binders and a second subset of molecular binders, and wherein:for each molecular binder in the first subset of molecular binders, the corresponding nucleic acid barcode sequence comprises a first constant portion and a second constant portion, wherein the first constant portion is positioned at a 5’ end of the nucleic acid barcode sequence, and the second constant portion is positioned at a 3’ end of the nucleic acid barcode sequence,for each molecular binder in the second subset of molecular binders, the corresponding nucleic acid barcode sequence comprises a third constant portion and a fourth constant portion, wherein the third constant portion is positioned at a 5’ end of the nucleic acid barcode sequence, and the fourth constant portion is positioned at a 3’ end of the nucleic acid barcode sequence,the first constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the fourth constant portion, andDB2 / 651687418.2 101Attorney Docket No.: 139643-5001-WOthe second constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the third constant portion.

54. The method of any one of claims 1-53, further comprising generating the plurality of molecular binders using mRNA display, cDNA display, in vivo mRNA display, LABEL-seq, Molecular Indexing of Proteins by Self-Assembly (MIPSA), or a combination thereof.

55. The method of any one of claims 1-54, wherein the DNA polymerase comprises DNA polymerase I or T4 polymerase.

56. The method of any one of claims 1-55, wherein the contacting B) further comprises contacting the plurality of protein analytes with a plurality of deoxynucleotide triphosphates (dNTPs).

57. The method of claim 56, wherein the plurality of dNTPs comprises at least 3 DNA base identities.

58. The method of claim 56, wherein the plurality of dNTPs consists of 3 DNA base identities.

59. The method of claim 58, wherein the nucleic acid barcode comprises: (i) a nucleic acid sequence consisting of 3 nucleic acid base identities complementary to the 3 DNA base identities in the plurality of dNTPs, and (ii) a stop position consisting of a 4th nucleic acid base identity, and wherein: an extension reaction catalyzed by the DNA polymerase ceases upon reaching the stop position.

60. The method of any one of claims 1-59, wherein the DNA concatemer is generated using an automated reaction device.DB2 / 651687418.2 102Attorney Docket No.: 139643-5001-WO61. The method of any one of claims 1-60, wherein generating C) comprises, for each respective molecular binder in the plurality of molecular binders that forms a binder-analyte interaction with the first protein analyte:hybridizing a 5’ portion of the corresponding nucleic acid barcode sequence for the respective molecular binder to a 3’ portion of the DNA concatemer, or a DNA initiator sequence appended to the first protein analyte,extending the DNA concatemer, or the DNA initiator sequence, using the DNA polymerase by generating a reverse complement of the nucleic acid barcode sequence, andreleasing the corresponding nucleic acid barcode.

62. The method of any one of claims 1-61, wherein the generating C) comprises adding each nucleic acid barcode sequence to the DNA concatemer using auto-cycling proximity recording (APR).

63. The method of any one of claims 1-62, wherein the indication of the presence or abundance of each respective molecular binder in the plurality of molecular binders is a count of DNA barcode sequences, in the DNA concatemer, that correspond to the respective molecular binder.

64. The method of any one of claims 1-63, wherein the interaction data structure is determined by a procedure comprising:mapping a plurality of sequence reads obtained from the sequencing of the DNA concatemer to a plurality of reference sequences comprising DNA barcode sequences associated with each molecular binder in the plurality of molecular binders, and determining, for each respective molecular binder in the plurality of molecular binders, a corresponding count of the corresponding DNA barcode sequence, wherein each entry in the interaction data structure comprises a count of DNA barcode sequences corresponding to the respective molecular binder.

65. The method of any one of claims 1-64, further comprising:DB2 / 651687418.2 103Attorney Docket No.: 139643-5001-WOdetermining a respective interaction data structure for each protein analyte in the plurality of protein analytes, thereby generating a plurality of interaction data structures, and using the plurality of interaction data structures to obtain an interaction matrix dimensioned by the plurality of protein analytes and the plurality of molecular binders, wherein each entry in the interaction matrix comprises a count of DNA barcode sequences corresponding to a respective molecular binder for a respective protein analyte.

66. The method of any one of claims 1-65, wherein the indication of the identity for the first protein analyte is a protein identity of the first protein analyte.

67. The method of any one of claims 1-65, wherein the indication of the identity for the first protein analyte is a probability of each candidate identity, in a set of candidate identities.

68. The method of any one of claims 1-65, wherein the indication of the identity for the first protein analyte is a probability of a candidate identity, further comprising: applying an identification threshold to the probability to assign the candidate identity to the first protein analyte.

69. The method of any one of claims 1-68, further comprising receiving, as output from the model, an abundance of the first protein analyte.

70. The method of any one of claims 1-69, wherein the indication of the identity for the first protein analyte is obtained based on a sampling from a reference distribution, wherein:the reference distribution comprises, for each respective protein analyte in the plurality of protein analytes, for each respective molecular binder in the plurality of molecular binders, a corresponding rate of production for the DNA barcode sequence corresponding to the nucleic acid barcode sequence that is unique to the respective molecular binder.

71. The method of any one of claims 1-70, further comprising using the indication of the identity to determine a presence or abundance of the first protein analyte responsive to a target condition.DB2 / 651687418.2 104Attorney Docket No.: 139643-5001-WO72. The method of any one of claims 1-71, further comprising using the indication of the identity to determine a change in a presence or abundance of the first protein analyte responsive to a target condition.

73. A method for detecting protein analytes, comprising:A) obtaining a plurality of protein analytes from a sample of a subject, wherein: the plurality of protein analytes comprises a first class of protein analytes and a second class of protein analytes, andeach protein analyte in the first class of protein analytes is attached to a first initiator tag comprising a first DNA initiator sequence; andB) contacting the plurality of protein analytes with a plurality of molecular binders and a DNA polymerase, thereby forming a plurality of binder-analyte interactions, wherein:each protein analyte in the first class of protein analytes forms a binder-analyte interaction with at least one molecular binder in the plurality of molecular binders, each respective molecular binder in the plurality of molecular binders comprises a corresponding nucleic acid barcode sequence that is unique to the respective molecular binder, andone or more molecular binders in the plurality of molecular binders interacts nonspecifically with each protein analyte in a corresponding subset of the first class of protein analytes;C) generating, for a first protein analyte in the first class of protein analytes, a DNA concatemer comprising a corresponding set of DNA barcode sequences, wherein each DNA barcode sequence in the corresponding set of DNA barcode sequences:i) corresponds to a nucleic acid barcode sequence of a respective molecular binder, in the plurality of molecular binders, that forms a binder-analyte interaction with the first protein analyte, andii) is appended to the first DNA initiator sequence of the first protein analyte using the DNA polymerase, responsive to formation of the binder-analyte interaction between the respective molecular binder and the first protein analyte, thereby causing a spatial order of DNA barcode sequences in the corresponding set of DNA barcode sequences to correspond to a temporal order of formation of binder-analyte interactions DB2 / 651687418.2 105Attorney Docket No.: 139643-5001-WObetween respective molecular binders, in the plurality of molecular binders, and the first protein analyte;D) determining an interaction data structure for the first protein analyte based on a sequencing of the DNA concatemer, wherein the interaction data structure comprises an indication of a presence or abundance of each respective molecular binder in the plurality of molecular binders that formed a binder-analyte interaction with the first protein analyte; and E) receiving, responsive to inputting the interaction data structure to a model, an indication of an identity for the first protein analyte, as output from the model.

74. The method of claim 73, wherein each protein analyte in the second class of protein analytes is free of an initiator tag comprising the first DNA initiator sequence.

75. The method of claim 73 or 74, wherein each protein analyte in the second class of protein analytes is attached to a second initiator tag comprising a second DNA initiator sequence, further comprising:generating, for a second protein analyte in the second class of protein analytes, a DNA concatemer comprising a corresponding set of DNA barcode sequences, wherein each DNA barcode sequence in the corresponding set of DNA barcode sequences:i) corresponds to a nucleic acid barcode sequence of a respective molecular binder, in the plurality of molecular binders, that forms a binder-analyte interaction with the second protein analyte, andii) is appended to the second DNA initiator sequence of the second protein analyte using the DNA polymerase, responsive to formation of the binder-analyte interaction between the respective molecular binder and the second protein analyte, thereby causing a spatial order of DNA barcode sequences in the corresponding set of DNA barcode sequences to correspond to a temporal order of formation of binderanalyte interactions between respective molecular binders, in the plurality of molecular binders, and the second protein analyte;D) determining an interaction data structure for the second protein analyte based on a sequencing of the DNA concatemer, wherein the interaction data structure comprises an indication of a presence or abundance of each respective molecular binder in the plurality ofDB2 / 651687418.2 106Attorney Docket No.: 139643-5001-WOmolecular binders that formed a binder-analyte interaction with the second protein analyte; andE) receiving, responsive to inputting the interaction data structure to a model, an indication of an identity for the second protein analyte, as output from the model.

76. The method of claim 75, wherein each protein analyte in the first class of protein analytes is free of an initiator tag comprising the second DNA initiator sequence.

77. The method of claim 75 or 76, further comprising using the first DNA initiator sequence and the second DNA initiator sequence to determine a corresponding class for the first protein analyte and a corresponding class for the second protein analyte.

78. The method of any one of claims 75-77, wherein the first initiator tag preferentially binds to a first epitope on each protein analyte in the first class of protein analytes, and the second initiator tag preferentially binds to a second epitope on each protein analyte in the second class of protein analytes.

79. A composition, comprising:a DNA polymerase;a plurality of dNTPs;one or more protein analytes, wherein each respective protein analyte in the one or more protein analytes is attached to an initiator tag comprising a DNA initiator sequence; and a plurality of molecular binders, wherein:each protein analyte in the plurality of protein analytes forms a binder-analyte interaction with at least one molecular binder in the plurality of molecular binders, each respective molecular binder in the plurality of molecular binders comprises a corresponding nucleic acid barcode sequence that is unique to the respective molecular binder, andone or more molecular binders in the plurality of molecular binders interacts nonspecifically with each protein analyte in a corresponding subset of the plurality of protein analytes.DB2 / 651687418.2 107Attorney Docket No.: 139643-5001-WO80. The composition of claim 79, wherein a corresponding nucleic acid barcode sequence for a first molecular binder in the plurality of molecular binders comprises a first nucleotide sequence that is reverse complementary to a second nucleotide sequence of at least one other molecular binder in the plurality of molecular binders.

81. The composition of claim 79 or 80, wherein:a first nucleic acid barcode sequence for a first molecular binder in the plurality of molecular binders comprises a first portion positioned at a 5’ end of the nucleic acid barcode sequence,a second nucleic acid barcode sequence for a second molecular binder in the plurality of molecular binders comprises a second portion positioned at a 3’ end of the nucleic acid barcode sequence, andthe first portion of the first nucleic acid barcode sequence hybridizes to a reverse complement of the second portion of the second nucleic acid barcode sequence.

82. The composition of any one of claims 79-81, wherein, for each respective molecular binder in the plurality of molecular binders, the corresponding nucleic acid barcode sequence further comprises a first constant portion and a second constant portion, wherein the first constant portion is positioned at a 5’ end of the nucleic acid barcode sequence, the second constant portion is positioned at a 3’ end of the nucleic acid barcode sequence, and the first constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the second constant portion.

83. The composition of any one of claims 79-81, wherein the plurality of molecular binders comprises a first subset of molecular binders and a second subset of molecular binders, and wherein:for each molecular binder in the first subset of molecular binders, the corresponding nucleic acid barcode sequence comprises a first constant portion and a second constant portion, wherein the first constant portion is positioned at a 5’ end of the nucleic acid barcode sequence, and the second constant portion is positioned at a 3’ end of the nucleic acid barcode sequence,DB2 / 651687418.2 108Attorney Docket No.: 139643-5001-WOfor each molecular binder in the second subset of molecular binders, the corresponding nucleic acid barcode sequence comprises a third constant portion and a fourth constant portion, wherein the third constant portion is positioned at a 5’ end of the nucleic acid barcode sequence, and the fourth constant portion is positioned at a 3’ end of the nucleic acid barcode sequence,the first constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the fourth constant portion, andthe second constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the third constant portion.

84. A kit, comprising:a DNA polymerase;a plurality of dNTPs; anda plurality of molecular binders, wherein:each protein analyte in a plurality of protein analytes forms a binder-analyte interaction with at least one molecular binder in the plurality of molecular binders, each respective molecular binder in the plurality of molecular binders comprises a corresponding nucleic acid barcode sequence that is unique to the respective molecular binder, andone or more molecular binders in the plurality of molecular binders interacts nonspecifically with each protein analyte in a corresponding subset of the plurality of protein analytes.

85. The kit of claim 84, wherein a corresponding nucleic acid barcode sequence for a first molecular binder in the plurality of molecular binders comprises a first nucleotide sequence that is reverse complementary to a second nucleotide sequence of at least one other molecular binder in the plurality of molecular binders.

86. The kit of claim 84 or 85, wherein:DB2 / 651687418.2 109Attorney Docket No.: 139643-5001-WOa first nucleic acid barcode sequence for a first molecular binder in the plurality of molecular binders comprises a first portion positioned at a 5’ end of the nucleic acid barcode sequence,a second nucleic acid barcode sequence for a second molecular binder in the plurality of molecular binders comprises a second portion positioned at a 3’ end of the nucleic acid barcode sequence, andthe first portion of the first nucleic acid barcode sequence hybridizes to a reverse complement of the second portion of the second nucleic acid barcode sequence.

87. The kit of any one of claims 84-86, wherein, for each respective molecular binder in the plurality of molecular binders, the corresponding nucleic acid barcode sequence further comprises a first constant portion and a second constant portion, wherein the first constant portion is positioned at a 5’ end of the nucleic acid barcode sequence, the second constant portion is positioned at a 3’ end of the nucleic acid barcode sequence, and the first constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the second constant portion.

88. The kit of any one of claims 84-86, wherein the plurality of molecular binders comprises a first subset of molecular binders and a second subset of molecular binders, and wherein: for each molecular binder in the first subset of molecular binders, the corresponding nucleic acid barcode sequence comprises a first constant portion and a second constant portion, wherein the first constant portion is positioned at a 5’ end of the nucleic acid barcode sequence, and the second constant portion is positioned at a 3’ end of the nucleic acid barcode sequence,for each molecular binder in the second subset of molecular binders, the corresponding nucleic acid barcode sequence comprises a third constant portion and a fourth constant portion, wherein the third constant portion is positioned at a 5’ end of the nucleic acid barcode sequence, and the fourth constant portion is positioned at a 3’ end of the nucleic acid barcode sequence,the first constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the fourth constant portion, andDB2 / 651687418.2 110Attorney Docket No.: 139643-5001-WOthe second constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the third constant portion.

89. A method for identifying protein analytes, comprising:A) obtaining a plurality of protein analytes from a sample of a subject;B) contacting the plurality of protein analytes with a plurality of molecular binders and a DNA polymerase, thereby forming a plurality of binder-analyte interactions, wherein:each respective protein analyte in the plurality of protein analytes comprises a corresponding nucleic acid barcode sequence that is unique to the respective protein analyte, andone or more molecular binders in the plurality of molecular binders interacts nonspecifically with each protein analyte in a corresponding subset of the plurality of protein analytes;C) generating, for a first molecular binder in the plurality of molecular binders, a DNA concatemer comprising a corresponding set of DNA barcode sequences, wherein each DNA barcode sequence in the corresponding set of DNA barcode sequences:i) corresponds to a nucleic acid barcode sequence of a respective protein analyte, in the plurality of protein analytes, that forms a binder-analyte interaction with the first molecular binder, andii) is appended to the first molecular binder using the DNA polymerase, responsive to formation of the binder-analyte interaction between the respective protein analyte and the first molecular binder, thereby causing a spatial order of DNA barcode sequences in the corresponding set of DNA barcode sequences to correspond to a temporal order of formation of binder-analyte interactions between respective protein analytes, in the plurality of protein analytes, and the first molecular binder; D) determining an interaction data structure for the first molecular binder based on a sequencing of the DNA concatemer, wherein the interaction data structure comprises an indication of a presence or abundance of each respective protein analyte in the plurality of protein analytes that formed a binder-analyte interaction with the first molecular binder; and E) receiving, responsive to inputting at least the interaction data structure to a model, an indication of an identity for a first protein analyte in the plurality of protein analytes, as output from the model.DB2 / 651687418.2 111Attorney Docket No.: 139643-5001-WO90. The method of claim 89, wherein each respective molecular binder in the plurality of molecular binders is attached to an initiator tag comprising a DNA initiator sequence, and wherein the DNA concatemer is appended to the DNA initiator sequence of the first molecular binder using the DNA polymerase.

91. The method of claim 90, further comprising attaching the initiator tag to the first molecular binder.

92. The method of claim 90 or 91, wherein the attaching the initiator tag to a respective molecular binder in the plurality of molecular binders comprises functionalizing the respective molecular binder using a linker that attaches the respective molecular binder to the initiator tag.

93. The method of claim 92, wherein:the linker is a bifunctional linker, andthe respective molecular binder is functionalized with the bifunctional linker using primary amine modification by a process comprising:modifying a primary amine of the respective molecular binder, thereby obtaining a first attachment moiety comprising a modified primary amine on the respective molecular binder,attaching the bifunctional linker to the respective molecular binder, wherein the bifunctional linker comprises a second attachment moiety and a third attachment moiety, and wherein the second attachment moiety of the bifunctional linker binds to the first attachment moiety of the respective molecular binder, andattaching the initiator tag to the bifunctional linker, wherein the initiator tag comprises a fourth attachment moiety that binds to the third attachment moiety.

94. The method of claim 93, wherein the initiator tag is attached to the respective molecular binder using azide-alkyne cycloaddition.DB2 / 651687418.2 112Attorney Docket No.: 139643-5001-WO95. The method of claim 93, wherein the initiator tag is attached to the bifunctional linker using click chemistry or trans-cyclooctene (TCO) click chemistry.

96. The method of any one of claims 89-95, wherein a corresponding nucleic acid barcode sequence for a first protein analyte in the plurality of protein analytes comprises a first nucleotide sequence that is reverse complementary to a second nucleotide sequence of at least one other protein analyte in the plurality of protein analytes.

97. The method of any one of claims 89-96, wherein:a first nucleic acid barcode sequence for a first protein analyte in the plurality of protein analytes comprises a first portion positioned at a 5’ end of the nucleic acid barcode sequence, a second nucleic acid barcode sequence for a second protein analyte in the plurality of protein analytes comprises a second portion positioned at a 3’ end of the nucleic acid barcode sequence, andthe first portion of the first nucleic acid barcode sequence hybridizes to a reverse complement of the second portion of the second nucleic acid barcode sequence.

98. The method of any one of claims 89-97, wherein the nucleic acid barcode sequence comprises a first constant portion and a second constant portion, wherein the first constant portion is positioned at a 5’ end of the nucleic acid barcode sequence, the second constant portion is positioned at a 3’ end of the nucleic acid barcode sequence, and the first constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the second constant portion.

99. The method of any one of claims 89-97, wherein the plurality of protein analytes comprises a first subset of protein analytes and a second subset of protein analytes, and wherein:for each protein analyte in the first subset of protein analytes, the corresponding nucleic acid barcode sequence comprises a first constant portion and a second constant portion, wherein the first constant portion is positioned at a 5’ end of the nucleic acid barcode sequence, and the second constant portion is positioned at a 3’ end of the nucleic acid barcode sequence,DB2 / 651687418.2 113Attorney Docket No.: 139643-5001-WOfor each protein analyte in the second subset of protein analytes, the corresponding nucleic acid barcode sequence comprises a third constant portion and a fourth constant portion, wherein the third constant portion is positioned at a 5’ end of the nucleic acid barcode sequence, and the fourth constant portion is positioned at a 3’ end of the nucleic acid barcode sequence,the first constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the fourth constant portion, andthe second constant portion has a nucleotide sequence that is identical to a nucleotide sequence of the third constant portion.

100. The method of any one of claims 89-99, wherein generating C) comprises, for each respective protein analyte in the plurality of protein analytes that forms a binder-analyte interaction with the first molecular binder:hybridizing a 5’ portion of the corresponding nucleic acid barcode sequence for the respective protein analyte to a 3’ portion of the DNA concatemer, or a DNA initiator sequence appended to the first molecular binder,extending the DNA concatemer, or the DNA initiator sequence, using the DNA polymerase by generating a reverse complement of the nucleic acid barcode sequence, andreleasing the corresponding nucleic acid barcode.

101. The method of any one of claims 89-100, wherein the generating C) comprises adding each nucleic acid barcode sequence to the DNA concatemer using auto-cycling proximity recording (APR).

102. The method of any one of claims 89-101, wherein the indication of the presence or abundance of each respective protein analyte in the plurality of protein analytes is a count of DNA barcode sequences, in the DNA concatemer, that correspond to the respective protein analyte.

103. The method of any one of claims 89-102, wherein the interaction data structure is determined by a procedure comprising:DB2 / 651687418.2 114Attorney Docket No.: 139643-5001-WOmapping a plurality of sequence reads obtained from the sequencing of the DNA concatemer to a plurality of reference sequences comprising DNA barcode sequences associated with each protein analyte in the plurality of protein analytes, and determining, for each respective protein analyte in the plurality of protein analytes, a corresponding count of the corresponding DNA barcode sequence, wherein each entry in the interaction data structure comprises a count of DNA barcode sequences corresponding to the respective protein analyte.

104. The method of any one of claims 89-103, further comprising:determining a respective interaction data structure for each molecular binder in the plurality of molecular binders, thereby generating a plurality of interaction data structures, andusing the plurality of interaction data structures to obtain an interaction matrix dimensioned by the plurality of protein analytes and the plurality of molecular binders, wherein each entry in the interaction matrix comprises a count of DNA barcode sequences corresponding to a respective protein analyte for a respective molecular binder.

105. The method of any one of claims 89-104, wherein the indication of the identity for the first protein analyte is obtained based on a sampling from a reference distribution, wherein:the reference distribution comprises, for each respective protein analyte in the plurality of protein analytes, for each respective molecular binder in the plurality of molecular binders, a corresponding rate of production for the DNA barcode sequence corresponding to the nucleic acid barcode sequence that is unique to the respective protein analyte.

106. A method for identifying protein analytes, comprising:A) obtaining a plurality of protein analytes from a sample of a subject;B) contacting the plurality of protein analytes with a plurality of first molecular binders and a DNA polymerase, thereby forming a plurality of binder-analyte interactions;C) generating a DNA concatemer comprising a corresponding set of DNA barcode sequences, wherein a spatial order of DNA barcode sequences in the corresponding set of DNA barcode sequences corresponds to a temporal order of formation of binder-analyte interactions between respective first molecular binders, in the plurality of first molecular DB2 / 651687418.2 115Attorney Docket No.: 139643-5001-WObinders, and a respective protein analyte in the plurality of protein analytes, wherein each DNA barcode sequence in the corresponding set of DNA barcode sequences:i) corresponds to a nucleic acid barcode sequence of a respective first molecular binder, in the plurality of first molecular binders, that forms a binderanalyte interaction with the respective protein analyte, andii) is appended to the respective protein analyte using the DNA polymerase, responsive to formation of the binder-analyte interaction between the respective first molecular binder and the respective protein analyte;D) determining an interaction data structure for the respective protein analyte based on a sequencing of the DNA concatemer, wherein the interaction data structure comprises an indication of a presence or abundance of each respective first molecular binder in the plurality of first molecular binders that formed a binder-analyte interaction with the respective protein analyte; andE) receiving, responsive to inputting the interaction data structure to a model, an indication of an identity for the respective protein analyte, as output from the model.

107. The method of claim 106, wherein the contacting B) further comprises contacting the plurality of protein analytes with one or more additional molecular binders, other than the plurality of first molecular binders.

108. The method of claim 107, wherein the contacting B) further comprises contacting the plurality of protein analytes with a plurality of second molecular binders, wherein:each respective protein analyte in a first subset of the plurality of protein analytes forms a binder-analyte interaction with a second molecular binder in the plurality of second molecular binders, andeach respective protein analyte in a second subset of the plurality of protein analytes does not form binder-analyte interactions with the plurality of second molecular binders.

109. The method of claim 108, wherein each respective first molecular binder in the plurality of first molecular binders comprises a peptide.DB2 / 651687418.2 116Attorney Docket No.: 139643-5001-WO110. The method of claim 108 or 109, wherein each respective second molecular binder in the plurality of second molecular binders comprises a small molecule.

111. The method of any one of claims 108-110, wherein each respective second molecular binder satisfies any two or more rules, any three or more rules, or all four rules of the Lipinski's rule of Five:(i) not more than five hydrogen bond donors,(ii) not more than ten hydrogen bond acceptors,(iii) a molecular weight under 500 Daltons, and(iv) a LogP under 5.

112. The method of any one of claims 108-111, further comprising attaching, to each respective second molecular binder in the plurality of second molecular binders, an initiator tag comprising a DNA initiator sequence, wherein the DNA concatemer is appended to the DNA initiator sequence of the second molecular binder using the DNA polymerase.

113. The method of claim 112, further comprising, prior to contacting the plurality of protein analytes with the plurality of second molecular binders, functionalizing each respective second molecular binder in the plurality of second molecular binders with a corresponding linker comprising an attachment moiety for attachment with the initiator tag.

114. The method of claim 113, wherein the initiator tag is attached, via the corresponding linker, to each respective second molecular binder in the plurality of second molecular binders prior to contacting the plurality of protein analytes with the plurality of second molecular binders.

115. The method of claim 113, wherein the initiator tag is attached, via the corresponding linker, to each respective second molecular binder in the plurality of second molecular binders after contacting the plurality of protein analytes with the plurality of second molecular binders.DB2 / 651687418.2 117Attorney Docket No.: 139643-5001-WO116. The method of any one of claims 113-115, wherein:the linker is a bifunctional linker, andeach respective second molecular binder in the plurality of second molecular binders is functionalized with the bifunctional linker using primary amine modification by a process comprising:modifying a primary amine of the respective second molecular binder, thereby obtaining a first attachment moiety comprising a modified primary amine on the respective second molecular binder,attaching the bifunctional linker to the respective second molecular binder, wherein the bifunctional linker comprises a second attachment moiety and a third attachment moiety, and wherein the second attachment moiety of the bifunctional linker binds to the first attachment moiety of the respective second molecular binder, andattaching the initiator tag to the bifunctional linker, wherein the initiator tag comprises a fourth attachment moiety that binds to the third attachment moiety.

117. The method of claim 116, wherein the initiator tag is attached to the respective second molecular binder using azide-alkyne cycloaddition.

118. The method of claim 116, wherein the initiator tag is attached to the bifunctional linker using click chemistry or trans-cyclooctene (TCO) click chemistry.

119. The method of any one of claims 112-118, wherein a respective second molecular binder in the plurality of second molecular binder comprises one or more binding moieties.

120. The method of claim 119, further comprising attaching the initiator tag to a binding moiety of the respective second molecular binder, wherein the initiator tag is specific to the binding moiety.

121. The method of claim 119 or 120, wherein a binding moiety in the one or more binding moieties is an epitope.DB2 / 651687418.2 118Attorney Docket No.: 139643-5001-WO122. The method of any one of claims 112-121, wherein a respective second molecular binder in the plurality of second molecular binders comprises one or more initiator tags.

123. The method of any one of claims 108-122, wherein each respective second molecular binder in the plurality of second molecular binders is the same.

124. The method of any one of claims 108-111, wherein the plurality of second molecular binders comprises a plurality of different second molecular binders, and wherein each respective second molecular binder in the plurality of different second molecular binders comprises a corresponding unique initiator tag comprising a DNA initiator sequence that is unique to the respective second molecular binder.

125. The method of claim 124, wherein the unique initiator tag is attached to each respective second molecular binder in the plurality of different second molecular binders prior to contacting the plurality of protein analytes with the plurality of second molecular binders.

126. The method of any one of claims 108-111, further comprising attaching, to each respective protein analyte in the plurality of protein analytes, an initiator tag comprising a DNA initiator sequence, wherein the DNA concatemer is appended to the DNA initiator sequence of the respective protein analyte using the DNA polymerase.

127. The method of claim 126, wherein the initiator tag is attached to the respective protein analyte using azide-alkyne cycloaddition.

128. The method of claim 126 or 127, wherein the attaching the initiator tag to a respective protein analyte in the plurality of protein analytes comprises functionalizing the respective protein analyte using a linker comprising an attachment moiety for attachment with the initiator tag.DB2 / 651687418.2 119Attorney Docket No.: 139643-5001-WO129. The method of claim 128, wherein the initiator tag is attached to the linker using click chemistry or trans-cyclooctene (TCO) click chemistry.

130. The method of any one of claims 126-129, wherein a respective protein analyte in the plurality of protein analytes comprises one or more binding moieties, further comprising attaching the initiator tag to a binding moiety of the respective protein analyte in the plurality of protein analytes, wherein the initiator tag is specific to the binding moiety.

131. The method of claim 130, wherein a binding moiety in the one or more binding moieties is an epitope.

132. The method of any one of claims 126-131, wherein a respective protein analyte in the plurality of analytes comprises one or more initiator tags.

133. The method of any 126-132, wherein each respective second molecular binder in the plurality of second molecular binders comprises a corresponding nucleic acid barcode sequence that is unique to the respective second molecular binder.

134. The method of any one of claims 106-133, wherein each respective first molecular binder in the plurality of first molecular binders comprises a corresponding nucleic acid barcode sequence that is unique to the respective first molecular binder.

135. The method of any one of claims 106-134, wherein a respective first molecular binder in the plurality of first molecular binders has a dissociation constant (KD) of at least 1 micromolar (pM).

136. The method of any one of claims 106-135, wherein a respective first molecular binder in the plurality of first molecular binders binds to one or more protein interaction features in a plurality of protein interaction features.DB2 / 651687418.2 120Attorney Docket No.: 139643-5001-WO137. The method of any one of claims 106-136, wherein a respective first molecular binder in the plurality of first molecular binders binds to two or more protein interaction features in the plurality of protein interaction features.

138. The method of any one of claims 106-137, wherein a respective molecular binder in the plurality of molecular binders binds to two or more protein analytes in the plurality of protein analytes.DB2 / 651687418.2 121