Method of generating and screening peptide aptamer libraries from naturally occurring proteins
A method using a peptide aptamer library derived from natural peptides and MS/MS query spectra with signal-to-noise filters identifies peptide aptamers that bind targets, addressing the challenge of characterizing protein interactions within cells and on cell surfaces for diagnostic and therapeutic applications.
Patent Information
- Authority / Receiving Office
- AU · AU
- Patent Type
- Applications
- Current Assignee / Owner
- YYZ PHARMATECH INC
- Filing Date
- 2024-12-27
- Publication Date
- 2026-07-23
AI Technical Summary
Existing methods lack a systematic approach to identify peptide aptamers that bind targets at the proteome level for diagnostic and therapeutic purposes, particularly for characterizing protein interactions within cells and on cell surfaces.
A method involving a peptide aptamer library derived from natural peptides, where the library is contacted with a target, unbound aptamers are removed, and bound aptamers are analyzed using MS/MS query spectra, applying signal-to-noise filters and likelihood indicators to identify the peptide aptamer that binds the target.
This method enables the unambiguous identification of peptide aptamers that bind targets, overcoming uncertainties in mass spectrometry by using statistical analysis and experimental controls to enhance the accuracy of peptide sequence determination.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS:
[0001] This application claims priority from U.S. Provisional Application No. 63 / 616,398, filed December 29, 2023, the contents of which are incorporated herein by reference. FIELD:
[0002] The present disclosure generally relates to identifying peptide aptamers that bind a target, specifically by using a peptide aptamer library generated from naturally occurring proteins or peptides. BACKGROUND:
[0003] Methods to identify protein-protein interactions at the proteome level is known in the art (see e.g. Elhabashy et al (2022) Exploring protein-protein interactions at the proteome level, Structure, 30: 462-475; Richards et al (2021) Mass spectrometrybased protein-protein interaction networks for the study of human diseases. Mol Syst Biol. 17(1):e8792). LC-ESI-MS / MS can be used to identify the ligands of plasma and identify their receptor complexes on the surface of live cells. However, there remains a need for systematic affinity ligand-receptor systems to characterize protein interaction within the cell and on the cell surface, for the discovery of new circulating ligands and receptor drug targets for diagnostic and therapeutic purposes. SUMMARY:
[0004] In one aspect, provided herein is a method of identifying a peptide aptamer, from a library comprising a plurality of peptide aptamers, that binds a target, comprising: providing the peptide aptamer library comprising the plurality of peptide aptamers, the plurality of peptide aptamers being derived or adapted from one or more natural peptides; contacting the peptide aptamer library with the target; removing unbound peptide aptamers after contacting the peptide aptamer library with the target; analyzing the bound peptide aptamers, comprising: generating a MS / MS query spectrum of the bound peptide aptamers; receiving one or more parameters of the query spectrum; providing one or more candidate spectra; generating a plurality of query samples of the query spectrum; selecting at least one query sample from the plurality of query samples for comparison with the one or more candidate spectra; determining a likelihood indicator for each of the one or more candidate spectra based on a comparison with the at least one query sample; applying a signal to noise filter to the one or more candidate spectra based on the likelihood indicators for the candidate spectra; selecting at least one candidate spectrum as a proposed spectrum; and determining a peptide sequence of the proposed spectrum; thereby identifying the peptide aptamer of the plurality of peptide aptamers that binds the target.
[0005] In another aspect, provided herein is a method of identifying a peptide aptamer, from a library comprising a plurality of peptide aptamers, that binds a target, comprising: providing the peptide aptamer library comprising the plurality of peptide aptamers, the plurality of peptide aptamers being derived or adapted from one or more natural peptides; contacting the peptide aptamer library with the target; removing unbound peptide aptamers after contacting the peptide aptamer library with the target; analyzing the bound peptide aptamers; thereby identifying the peptide aptamer of the plurality of peptide aptamers that binds the target.
[0006] In some embodiments, analyzing the bound peptide aptamers comprises: generating a MS / MS query spectrum of the bound peptide aptamers; receiving one or more parameters of the query spectrum; providing one or more candidate spectra; generating a plurality of query samples of the query spectrum; selecting at least one query sample from the plurality of query samples for comparison with the one or more candidate spectra; determining a likelihood indicator for each of the one or more candidate spectra based on a comparison with the at least one query sample; applying a signal to noise filter to the one or more candidate spectra based on the likelihood indicators for the candidate spectra; and selecting at least one candidate spectrum as a proposed spectrum; and determining a peptide sequence of the proposed spectrum.
[0007] In some embodiments, analyzing the bound peptide aptamers does not involve determining an in-frame amino acid sequence of the one or more candidate spectra.
[0008] In some embodiments, determining the likelihood indicator comprises determining if more than one query sample fit to a candidate spectrum.
[0009] In some embodiments, applying the signal to noise filter comprises (i) determining an observation frequency of at least two query samples that fit to the candidate spectrum; (ii) determine an observation frequency of at least one control; (iii) if the observation frequency of the at least two samples that fit to the candidate spectrum is higher than the observation frequency of the at least one control, then the candidate spectrum is selected as a proposed spectrum.
[0010] In some embodiments, the method further comprises releasing bound peptide aptamers from the target after removing unbound peptide aptamers.
[0011] In some embodiments, generating the MS / MS query spectrum comprises generating a MS / MS query spectrum of the released bound peptide aptamers.
[0012] In some embodiments, the method further comprises digesting the bound peptide aptamers after removing unbound aptamers and before analyzing the bound peptide aptamers.
[0013] In some embodiments, generating the MS / MS query spectrum comprises generating a MS / MS query spectrum of the digested released bound peptide aptamers.
[0014] In some embodiments, deriving the plurality of peptide aptamers from the one or more natural peptides comprises digesting the one or more natural peptides.
[0015] In some embodiments, deriving the plurality of peptide aptamers from the one or more natural peptides comprises modification of the one or more natural peptides.
[0016] In some embodiments, the plurality of peptide aptamers comprises undigested natural peptides.
[0017] In some embodiments, the plurality of peptide aptamers comprises untreated natural peptides.
[0018] In some embodiments, the one or more natural peptides comprise an immunoglobulin superfamily member.
[0019] In some embodiments, the one or more natural peptides comprise an immunoglobulin, a B cell antigen receptor (BCR), a T cell receptor (TCR) or any combination thereof.
[0020] In some embodiments, the immunoglobulin comprises IgG, IgA, IgM, IgE, and / or IgD.
[0021] In some embodiments, the one or more natural peptides comprise IgG.
[0022] In some embodiments, the peptide aptamer library comprises Fab isolated from immunoglobulins.
[0023] In some embodiments, the target comprises a receptor, a ligand, an enzyme, a protein, an antibody, a variable domain, or a drug.
[0024] In some embodiments, the one or more natural peptides comprises one or more natural peptides isolated from a host animal. In some embodiments, the host animal has contacted an immunogen. In some embodiments, the one or more candidate spectra comprise spectra predicted from a transcriptome of immune cells isolated from the host animal. In some embodiments, the predicted spectra are generated by: (1) isolating immune cells from the host animal; (2) isolating RNAfrom the immune cells; (3) obtaining sequences of the RNA. In some embodiments, obtaining sequences of the RNA comprises reverse transcription of the RNA to generate DNAand sequencing said DNA. In some embodiments, the method further comprises generating amino acid sequences from the DNA. In some embodiments, generating amino acid sequences from the DNA comprises reading a plurality of reading frames of the DNA sequence.
[0025] In some embodiments, the method further comprises determining a variable domain repertoire of immune cells of the host animal. In some embodiments, determining the variable domain repertoire of immune cells comprises DNA sequencing. In some embodiments, the method further comprises determining a sequence of a variable domain of the isolated immunoglobulin, BCR orTCR.
[0026] In some embodiments, the one or more natural peptides originate from a biological sample. In some embodiments, the biological sample comprises tissue cells or biofluids. In some embodiments, the biofluid is plasma. In some embodiments, the one or more candidate spectra comprise spectra predicted from genomic DNA sequence of the biological sample.
[0027] Other features and advantages of the present disclosure will become apparent from the following detailed description. It should be understood, however, that the detailed description and the specific examples while indicating embodiments of the disclosure are given by way of illustration only, since various changes and modifications within the spirit and scope of the disclosure will become apparent to those skilled in the art from this detailed description. BRIEF DESCRIPTION OF THE DRAWINGS:
[0028] The embodiments of the application will now be described in greater detail with reference to the attached drawings in which:
[0029] Figure 1 is a general scheme for creating new peptide aptamer drugs against drug targets according to one embodiment.
[0030] Figure 2 is a general flow diagram for creating new peptide aptamer drugs against drug targets according to one embodiment.
[0031] Figure 3 is a flow diagram of identifying peptide aptamers that bind a target according to one embodiment.
[0032] Figure 4 is a flow diagram illustrating a method of identifying peptide aptamers that bind a target according to one embodiment.
[0033] Figure 5 is a flow diagram illustrating a method of identifying peptide aptamers that bind a target according to one embodiment.
[0034] Figure 6 is a flow diagram illustrating a method of identifying peptide aptamers that bind a target according to one embodiment. DETAILED DESCRIPTION OF THE DISCLOSURE:
[0035] The following is a detailed description provided to aid those skilled in the art in practicing the present disclosure. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting of the disclosure. All publications, patent applications, patents, figures and other references mentioned herein are expressly incorporated by reference in their entirety.
[0036] All of the features disclosed in this specification may be combined in any combination. Each feature disclosed in this specification may be replaced by an alternative feature serving the same, equivalent, or similar purpose. Thus, unless 5 expressly stated otherwise, each feature disclosed is only an example of a generic series of equivalent or similar features. I. Definitions
[0037] In understanding the scope of the present disclosure, the term "comprising" and its derivatives, as used herein, are intended to be open ended terms that specify the presence of the stated features, elements, components, groups, integers, and / or steps, but do not exclude the presence of other unstated features, elements, components, groups, integers and / or steps. The foregoing also applies to words having similar meanings such as the terms, "including", "having" and their derivatives.
[0038] All definitions, as defined and used herein, should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.
[0039] The term “consisting” and its derivatives, as used herein, are intended to be closed ended terms that specify the presence of stated features, elements, components, groups, integers, and / or steps, and also exclude the presence of other unstated features, elements, components, groups, integers and / or steps.
[0040] Further, terms of degree such as "substantially", "about" and "approximately" as used herein mean a reasonable amount of deviation of the modified term such that the end result is not significantly changed. These terms of degree should be construed as including a deviation of at least ±5% of the modified term if this deviation would not negate the meaning of the word it modifies.
[0041] More specifically, the term “about” means plus or minus 0.1 to 20%, 5-20%, or 10-20%, 10%-15%, preferably 5-10%, most preferably about 5% of the number to which reference is being made.
[0042] As used in this specification and the appended claims, the singular forms “a”, “an” and “the” include plural references unless the content clearly dictates otherwise. Thus, for example, a composition containing “a compound” includes a mixture of two or more compounds. It should also be noted that the term “or” is generally employed in its sense including “and / or” unless the content clearly dictates otherwise.
[0043] The definitions and embodiments described in particular sections are intended to be applicable to other embodiments herein described for which they are suitable as would be understood by a person skilled in the art.
[0044] The recitation of numerical ranges by endpoints herein includes all numbers and fractions subsumed within that range (e.g. 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.90, 4, and 5). It is also to be understood that all numbers and fractions thereof are presumed to be modified by the term "about."
[0045] Further, the definitions and embodiments described in particular sections are intended to be applicable to other embodiments herein described for which they are suitable as would be understood by a person skilled in the art. For example, in the following passages, different aspects of the disclosure are defined in more detail. Each aspect so defined may be combined with any other aspect or aspects unless clearly indicated to the contrary.
[0046] Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present disclosure, examples of methods and materials are now described. II. Methods
[0047] Peptide aptamers are protein molecules that bind to targets, for example, target proteins. A proteomic approach may be used to identify peptide aptamers that bind a specific target. It is disclosed herein that the problems associated with mass spectrometry of peptide aptamers can be overcome by using digests of naturally occurring proteins and / or peptides from cells, fluids or plasma as aptamers.
[0048] A peptide may be unambiguously identified using high resolution mass spectrometry and de novo sequencing. However since 4 amino acids have ambiguous masses and given the uncertainty introduced by low resolution mass spectrometry the identification of peptide from the statistical fit of MS / MS spectra alone is uncertain and frequently several peptide sequences are obtained with nearly identical statistical probability.
[0049] In such an instance where the peptide sequence itself is uncertain from the fit of MS / MS spectra by total fragment ions accounted for, goodness of fit or cross correlation or regression or heuristic algorithms, the observation frequency of the potential peptide, or the observation frequency of more than one independent peptide or peptides from the same polypeptide, protein, aptamer, or variable domain may be compared to statistical and experimental controls.
[0050] For example, a specifc high observation frequency peptide can refer to a peptide nested within a parent polypeptide where the peptide was observed with a high freqeuncy in the affintiy chromatography experiment that was not observed in the noise and random statistical controls or the non-specific binding experimental controls such that the peptide may be further tested for binding activity.
[0051] The antigen specific peptides, aptamers, or variable domains from affinity chromatography may be identified with a certain probability using statistical analysis of observation frequency of matches of predicted versus observed MS / MS compared to controls of blanks, condition blanks, washes, and / or control affinity columns.
[0052] In some embodiments, the method can comprise the use of experimental controls. The use of experimental controls can comprise the use of: (1) naive blank columns (never used) or conditioned blank columns; (2) control affinity support resin without the specific antigen, ligand, receptor, or diagnostic or therapeutic target; (3) a support resin with a control antigen, ligand, receptor, or diagnostic or therapeutic target; (4) support resin with the specific antigen, ligand, receptor, or diagnostic or therapeutic target for use with pre-immune serum; or (5) some other controls from tissues, cells or bodily fluids.
[0053] The observation frequency of the source polypeptide may be computed from the observation frequency of the peptides nested within the polypeptide or compute the likelihood of the full-length polypeptide sequences. After correction against the control electromagnetic noise and random MS / MS the full-length polypeptides specific to the antigen affinity chromatography may be selected for downstream applications, such as testing binding to the antigen.
[0054] The method may comprise the use of statistical controls, which can comprise the use of naturally occurring electromagnetic noise or random MS / MS spectra or computer generated random MS / MS spectra to select potential peptide sequences with a higher probability of being true positive peptide identification from MS / MS spectra.
[0055] Accordingly, in one aspect, provided herein is a method of identifying a peptide aptamer, from a library comprising a plurality of peptide aptamers, that binds a target, comprising: providing the peptide aptamer library comprising the plurality of peptide aptamers, the plurality of peptide aptamers being derived or adapted from one or more natural peptides; contacting the peptide aptamer library with the target; removing unbound peptide aptamers after contacting the peptide aptamer library with the target; analyzing the bound peptide aptamers, comprising: generating a MS / MS query spectrum of the bound peptide aptamers; receiving one or more parameters of the query spectrum; providing one or more candidate spectra; generating a plurality of query samples of the query spectrum; selecting at least one query sample from the plurality of query samples for comparison with the one or more candidate spectra; determining a likelihood indicator for each of the one or more candidate spectra based on a comparison with the at least one query sample; applying a signal to noise filter to the one or more candidate spectra based on the likelihood indicators for the candidate spectra; selecting at least one candidate spectrum as a proposed spectrum; and determining a peptide sequence of the proposed spectrum; thereby identifying the peptide aptamer of the plurality of peptide aptamers that binds the target.
[0056] In one aspect, provided herein is a method of identifying a peptide aptamer that binds a target. The peptide aptamer can be from a library comprising a plurality of peptide aptamers. The method can comprise providing the peptide aptamer library. The peptide aptamer library can comprise the plurality of peptide aptamers. The plurality of peptide aptamers can be derived or adapted from one or more natural peptides. The method can comprise contacting the peptide aptamer library with the target. The method can comprise removing unbound peptide aptamers after contacting the peptide aptamer library with the target. The method can comprise analyzing the bound aptamers. The method can comprise generating a MS / MS query spectrum of the bound peptide aptamers. The method can comprise receiving one or more parameters of the query spectrum. The method can comprise providing one or more candidate spectra. The method can comprise generating a plurality of query samples of the query spectrum. The method can comprise selecting at least one query sample from the plurality of query samples for comparison with the one or more candidate spectra. The method can comprise determining a likelihood indicator for each of the one or more candidate spectra based on a comparison with the at least one query sample. The method can comprise applying a signal to noise filter to the one or more candidate spectra based on the likelihood indicators for the candidate spectra. The method can comprise selecting at least one candidate spectrum as a proposed spectrum. The method can comprise determining a peptide sequence of the proposed spectrum, thereby identifying the peptide aptamer of the plurality of peptide aptamers that binds the target.
[0057] In another aspect, provided herein is a method of identifying a peptide aptamer, from a library comprising a plurality of peptide aptamers, that binds a target, comprising: providing the peptide aptamer library comprising the plurality of peptide aptamers, the plurality of peptide aptamers being derived or adapted from one or more natural peptides; contacting the peptide aptamer library with the target; removing unbound peptide aptamers after contacting the peptide aptamer library with the target; analyzing the bound peptide aptamers; thereby identifying the peptide aptamer of the plurality of peptide aptamers that binds the target.
[0058] In some embodiments, identifying the peptide aptamers comprises de novo sequencing. In some embodiments, identifying the peptide aptamers does not comprise de novo sequencing.
[0059] The one or more candidate spectra can be spectra predicted from amino acid sequences derived from DNA sequencing data. The reading frame of the DNA sequences may not be known. In such case, the DNA sequences can be read in a plurality of reading frames to generate a plurality of amino acid sequences and therefore, a plurality of predicted spectra can be generated for each DNA sequence.
[0060] In some embodiments, identifying the peptide aptamers does not comprise determining an in-frame amino acid sequence of the one or more candidate spectra.
[0061] More than one query sample can be fitted to a candidate spectrum. When more than one query sample is fitted to a candidate spectrum, the likelihood that the candidate spectrum corresponds to a peptide aptamer can be higher than when only one query sample is fitted to the candidate spectrum.
[0062] In some embodiments, determining the likelihood indicator comprises determining if more than one query sample fit to a candidate spectrum. A positive likelihood indicator can refer to fitting more than one query sample to a candidate 10 spectrum. In some embodiments, a positive likelihood indicator comprises fitting more than one query sample to a candidate spectrum. In some embodiments, a positive likelihood indicator comprises fitting more than two query samples to a candidate spectrum. In some embodiments, a positive likelihood indicator comprises fitting more than three query samples to a candidate spectrum. In some embodiments, a positive likelihood indicator comprises fitting more than four query samples to a candidate spectrum. In some embodiments, a positive likelihood indicator comprises fitting more than five query samples to a candidate spectrum. In some embodiments, a positive likelihood indicator comprises fitting more than six query samples to a candidate spectrum. In some embodiments, a positive likelihood indicator comprises fitting more than seven query samples to a candidate spectrum. In some embodiments, a positive likelihood indicator comprises fitting more than eight query samples to a candidate spectrum.
[0063] If a positive likelihood indicator is determined, then the candidate spectrum may be selected as a proposed spectrum, or the query samples that fit to the candidate spectrum can be further analyzed.
[0064] A signal to noise filter can be applied. Noise can be generated from various sources. For example, noise can come from a statistical control and / or a non-specific binding control. A statistical control can include, for example, an electromagnetic noise control, a random MS / MS spectra control, or both. There are various ways to perform a non-specific binding control. For example, a pre-immune serum may be used in binding. An affinity column without an antigen attached, an affinity column with a control antigen attached, a naive affinity column that has not been conditioned (naive no sample for electromagnetic noise), and / or a conditioned affinity column (irrelevant proteins injected to adsorb to the exposed irreversible binding sites on the resin) can also be a non-specific binding control.
[0065] When applying a signal to noise filter, an observation frequency of at least one control and an observation frequency of at least one query sample can be determined. If the observation frequency of the at least one query sample is higher than the observation frequency of the at least one control, then the at least one query sample can be further analyzed and / or the candidate spectrum can be selected as a proposed spectrum.
[0066] A threshold of the difference between the observation frequencies may be set to allow the at least one query sample to be further analyzed.
[0067] In some embodiments, applying the signal to noise filter comprises (i) determining an observation frequency of at least one query sample; and (ii) determine an observation frequency of at least one control. In some embodiments, if the observation frequency of the at least one query sample is higher than the observation frequency of the at least one control, then the at least one query sample is further analyzed. In some embodiments, if the observation frequency of the at least one query sample is higher than the observation frequency of the at least one control, then the candidate spectrum is selected as a proposed spectrum.
[0068] The signal to noise filter can be applied before or after determining the likelihood indicator.
[0069] In some embodiments, the signal to noise filter is applied where a positive likelihood indicator is determined. For example, if two query samples are fitted to a candidate spectrum, then an observation frequency of the two query samples, and an observation frequency of at least one control can be determined.
[0070] In some embodiments, applying the signal to noise filter comprises (i) determining an observation frequency of at least two query samples that fit to the candidate spectrum; and (ii) determine an observation frequency of at least one control. In some embodiments, if the observation frequency of the at least two samples that fit to the candidate spectrum is higher than the observation frequency of the at least one control, then the candidate spectrum is selected as a proposed spectrum. Fitting of spectra
[0071] The bound peptide aptamers can be identified by mass spectrometry, including but not limited to top down mass spectrometry, or electrospray ionization, or MALDI ionization or chemical ionization or electron impact ionization or LC-ESI-MS / MS.
[0072] Identifying the peptide aptamer that binds the target can be performed using a computational method described in WO2024000077, the content of which is incorporated by reference herein in its entirety.
[0073] In some embodiments, the method can comprise analyzing the bound peptide aptamers by MS / MS or MSn. In some embodiments, the MS / MS or MSn can be fitted to peptide sequences with or without a database or using statistical methods to fit the peptides to randomly generated or other peptides. Optionally, the methods can also involve using the principles of peptide library design and de novo sequencing to constrain and guide the creation of a peptide library based on the amino acid composition determined from MS, MS / MS, and / or MSn. In some embodiments, the MS / MS or MSn can be fitted to candidate spectra with or without a database or using statistical methods to fit the peptides to candidate spectra derived from randomly generated or other peptides. Optionally, the candidate spectra can be predicted from nucleic acid sequences. The methods disclosed herein can further involve: generating the in silico peptide library to match the characteristics of the physical library used in the experiment and reducing the dimension for each MS / MS or MSn spectra searched using the observed amino acids from the MS / MS or MSn spectra; fitting the observed MS / MS or MSn spectra from the physical peptide that bound the target to the possible peptide in the in silico library using de novo, or goodness of fit, or regression or linear models or correlation or heuristic algorithms to determine the amino acid sequence and the molecular composition of matter of the physical peptide that bound the target.
[0074] In some embodiments, analyzing the bound peptide aptamers does not involve determining an amino acid sequence of the one or more candidate spectra.
[0075] The methods disclosed herein can match spectra from a sample to peptide sequences and / or candidate spectra. In some embodiments, a combination of scoring algorithms can be used, for example, based on a spectra line count, a count of sequential ions or total number of matched ions, and fitting algorithms based on chi square, linear regression, amino acid matching and a signal to noise intensity score to identify peptide sequences. The methods can use a search engine and the search engine can toggle back and forth between spectra generated from a reference library and a random number or random MS or random MS / MS generator given certain parameters depending on the nature of the experiment. The search engine can employ a series of nested loops and can map fragments onto their respective precursor masses for multiple levels of fragmentation events (MSn). The disclosed embodiments can validate results by for example, counting a number of spectra in the sample that match each identified peptide sequence and 13 comparing this count with a count for randomly simulated and / or experimentally derived (e.g., blank noise) spectra.
[0076] The method can comprise generating a MS / MS query spectrum of the bound peptide aptamer. The query spectrum can comprise one or more parameters. The one or more parameters can be received from user input. The one or more parameters can be derived from the query spectrum by a processor. For example, the one or more parameters can be one or more of a neutral loss - including ammonia losses and water losses, a post-translational modification shift - which may only be located on a small group of amino acids, an immonium ion (e.g., methionine oxidation), a subtraction of a B ion, or a subtraction of a Y ion.
[0077] The method can further comprise generating one or more candidate peptide sequences or candidate spectra based on the one or more parameters. In some embodiments, generating one or more candidate peptide sequences or candidate spectra can involve selecting, from the memory, at least one of a plurality of peptide sequences or candidate spectra stored in memory to use as at least one candidate peptide sequence or candidate spectrum. For example, the plurality of peptide sequences or candidate spectra can be stored in a peptide / spectra database or library. In some embodiments, the plurality of peptide sequences can include naturally occurring peptide sequences and / or synthetic peptide sequences. In some embodiments, the plurality of candidate spectra can include spectra predicted from naturally occurring peptide sequences and / or synthetic peptide sequences. Selecting candidate peptide sequences / spectra from the memory can be suitable when natural aptamers are from a known source and / or where the possible peptide sequences are known (e.g., but not limited to digested plasma).
[0078] In some embodiments, whether to select candidate peptide sequences from memory or by randomly generating candidate peptide sequences can be determined based on the one or more parameters.
[0079] In some embodiments, the one or more parameters can include a parameter specified by the user for selecting whether to select candidate peptide sequences from memory or by randomly generating candidate peptide sequences. In some embodiments, the one or more parameters can be compared to certain thresholds to determine whether to select candidate peptide sequences from memory or by randomly generating candidate peptide sequences.
[0080] In some embodiments, candidate peptide sequences can be selected from memory and randomly generated. For example, a subset of amino acids can be identified from a protein database and then the candidate peptide sequence can be randomly generated to identify post translational modifications (PTM) by including the amino acid masses with mass shifts, depending on the nature of the PTM. In some embodiments, candidate spectra can be predicted from the aforementioned candidate peptide sequences.
[0081] In some embodiments, candidate peptide sequences can initially be selected from memory. Depending on the likelihood indicators, the same query spectrum can be analyzed again but with candidate peptide sequences randomly generated.
[0082] In some embodiments, the method comprises providing one or more candidate spectra.
[0083] A candidate spectrum can be predicted from a nucleic acid sequence. The nucleic acid sequence can be a sequence derived from DNA or RNA, such as genomic DNA, a cDNA reverse transcribed from total RNA, mRNA, etc. For example, starting from a nucleic acid sequence, a reading frame can be shifted to generate a plurality of corresponding peptide sequences. Candidate spectra can be predicted from those peptide sequences corresponding to different frames.
[0084] In some embodiments, identifying natural peptide aptamers that bind a target further comprises obtaining nucleic acid sequences from the host animal. In some embodiments, the nucleic acid sequences are nucleic acid sequences derived from B cells isolated from the host animal. For example, total RNA or mRNA can be isolated from B cells of the host animal and reverse transcribed into cDNA. The V(D)J regions can be amplified and sequenced using techniques known in the art.
[0085] A plurality of samples of the query spectrum for which a peptide is being identified can be generated. In some embodiments, the samples of the query spectrum can be generated experimentally or by simulation. In some embodiments, the simulation can be a Monte Carlo random simulation. Generating a plurality of samples can involve generating spectra lines of the samples.
[0086] At least one query sample from the plurality of query samples can be selected for comparison with the one or more candidate peptide sequences or the one or more candidate spectra. Selecting query samples of the plurality of query samples for comparison can reduce the computational burden of fitting a large dataset to a large search space. In order to select query samples for comparison, one or more signal processing filters can be applied. Signal processing filters can include minimum spectra counting, maximum spectra lines filtering, precursor mass testing, base peak testing, and / or delta mass filtering.
[0087] Minimum spectra counting involves discarding spectra with line counts lower than a given threshold as there is likely not enough information below the threshold of spectra lines to make a reliable peptide match. In some embodiments, a number of spectra lines of a sample can be counted. In some embodiments, the number of matching sequential ion (e.g., b and y ions) spectra lines can be counted. If the number of spectra lines of the sample or the number of matching sequential ion spectra lines is less than a pre-determined minimum number of spectra lines, that sample can be excluded from comparison with the one or more candidate peptide sequences. In some embodiments, a correction is applied to the number of matching sequential ion spectra lines. For example, a ratio of the number of matching sequential ion spectra lines and the total number of spectra lines can be determined and the pre-determined minimum number of spectra lines can correspond to a pre-determined minimum ratio. In some embodiments, the predetermined minimum number of spectra lines can be specified by the user. For example, the pre-determined minimum number of spectra lines can be a parameter received from the user.
[0088] Maximum spectra lines filtering involves using the most intense spectra. For maximum spectra lines filtering, a spectra intensity of a sample can be determined. In some embodiments, the sum of the spectra intensities of matching ion spectra lines, the total sum of the spectra intensity (i.e., the sum of the spectra intensities of all of the spectra lines), the ratio of the sum of the spectra intensities of the matching ion spectra lines and the total sum of the spectra intensity can be determined. If the ratio is less than a predetermined threshold, that sample can be excluded from comparison with the one or more candidate peptide sequences or the one or more candidate spectra. In some embodiments, the pre-determined spectra intensity can be based on a candidate peptide 16 sequence, such as a percentage of the spectra intensity of a candidate peptide sequence. The percentage of the spectra intensity can be specified by the user. In this manner, with the pre-determined spectra intensity being based on a candidate peptide sequence, the maximum spectra lines filtering can be applied dynamically. The max spectra relates to the intensity values. Assuming a max spectra value of 50 the engine would only examine the 50 most intense spectra lines. This significantly reduces the computation time as a single MS2 spectra can contain thousands of spectra lines.
[0089] Precursor mass testing involves comparing a precursor mass of a sample to the mass of a candidate peptide sequence. Samples having a precursor mass matching the mass of any candidate peptide can be further processed. Samples having a precursor mass that does not match any candidate peptide can be discarded from further consideration, and the next sample is assessed.
[0090] Precursor mass testing can involve determining a precursor mass for a sample. If the precursor mass is substantially equal to a mass of a candidate peptide sequence of the one or more candidate peptide sequences, that sample can be selected for comparison with the one or more candidate peptide sequences. The precursor mass can be considered to be substantially equal to a mass of a candidate peptide sequence within a pre-determined error range of the mass of the candidate peptide sequence. In some embodiments, the pre-determined error range can be specified by the user.
[0091] In some embodiments, a precursor mass for a sample can be determined at different charge states. For example, precursor masses at charge states of 1,2, and 3 can be determined.
[0092] In some embodiments, a precursor mass for a sample can be determined with consideration of the presence of one or more post-translational modifications (PTMs). For example, the precursor mass can include a mass shift from one or more post-translational modifications.
[0093] Base peak testing involves comparing a mass of a theoretical ion of the sample to the mass of a base peak. Samples having theoretical ions having a mass that matches the mass of the base peak can be further processed. Samples having theoretical ions having a mass that does not match the base peak can be discarded from further consideration, and the next sample is assessed. In some embodiments, a mass of a theoretical ion of the sample is determined. If the mass of the theoretical ion is substantially equal to a mass of a base peak, that sample can be selected for comparison with the one or more candidate peptide sequences or the one or more candidate spectra.
[0094] In some embodiments, if the precursor mass is not present in the observed spectra, the precursor mass can be added to the theoretical spectra. By adding the precursor mass to the theoretical spectra, the likelihood of identifying the first amino acid in the peptide sequence can be increased and the Amino Acid Match Ratio (AAMR) score can be improved.
[0095] In some embodiments, a theoretical ion can be added to the theoretical spectra, that is the low end of the spectra can be added to. The theoretical ion can represent a terminal mass. Addition of the theoretical ion can improve the peptide match mass ratio score.
[0096] In some embodiments, a theoretical peptide’s MH value at charge state 1, or 2 or 3 could be added to the theoretical spectra. Addition of the theoretical peptide’s MH value at charge state 1, or 2 or 3 can apply to the peptide match score. Once a random peptide is generated, both the theoretical spectra and MH value can be calculated. The MH mass can be compared to the precursor mass so as to calculate the peptide's charge. The mass spectrometer however, regardless of the precursor mass may well be reporting the spectra for a peptide at each of the three charge states within a single MS2 scan. If this is the case, then the real 1 + spectra representing the unknown peptide can generally provide the most usable spectra. In some embodiments, the 1 + spectra can be used as the primary identification signal followed by examination of the 2+ and 3+ spectra for additional evidence of a good identification.
[0097] In some embodiments, the MH of a particular peptide can be injected into the observed spectra. Addition of the MH of a particular peptide can be used to calculate the peptide match score.
[0098] In some embodiments, the difference in theoretical mass and the observed peptide mass can be calculated using the estimated charge of a theoretical peptide, the known theoretical mass of the same theoretical peptide, and the known observed precursor mass. The delta mass can be used to identify a modification which is an extra chemical element on the peptide. This is also known as a modification mass. For example, phosphorylation is a commonly observed post translational modification and has a known mass shift of 79.99 Da; if the delta mass is about this value, then it can be determined that the peptide is phosphorylated. Such modifications are relevant to the role a peptide may play with respect to human health as the presence of this element shows in the precursor mass but is not part of the theoretical mass, and also may not appear in the MS2 spectra. This delta mass value is then highly suggestive of some modification. To test for a modification, the modification mass is subtracted from the precursor mass. If the results of this calculation match the theoretical mass of the random peptide and the real MS2 spectra matches the theoretical spectra, it can be assumed that there is a theoretical peptide and a modification.
[0099] In some embodiments, application of signal processing filters can vary. For example, in some embodiments, only minimum spectra counting can be used. In other embodiments, minimum spectra counting, maximum spectra lines filtering, sequential ion (e.g., b and y ions) counting, and precursor testing can be used. Other combinations are possible. Furthermore, the signal processing filters can also depend on whether a candidate peptide sequence or a candidate spectrum is selected from memory or randomly generated.
[00100] Candidate peptide sequences and candidate spectra can be selected from memory. In some embodiments, one theoretical peptide or candidate peptide sequence, or candidate spectrum at a time is compared to all available scans by tests and scoring algorithms until the peptide or spectrum list is exhausted. In this embodiment, each theoretical peptide and spectrum pair must pass a precursor mass test and base peak test before they are scored by at least one of chi square, multiple / linear or nested regression, amino acid match ratio, and ion intensity match ratio. A user defined combination of these scores can be used to identify the best match for each scan. In some embodiments, one sample can be compared to each candidate peptide sequence or candidate spectrum before another sample is compared to each of the candidate peptide sequences or candidate spectra.
[00101] A likelihood indicator for each candidate peptide sequence or each candidate spectrum can be determined based on a comparison with the at least one sample.
[00102] In some embodiments, determining a likelihood indicator for each candidate peptide sequence or candidate spectrum can involve applying scoring techniques, filtering techniques or a combination thereof. For example, scoring techniques can first be applied. The scoring techniques can be analogous to the signal processing filters, including spectra line filtering.
[00103] The likelihood indicator can alternatively or additionally be determined based on additional filtering techniques including one or more of a chi-square score, a regression score, and a cross correlation score. Other scoring functions can be used to generate a likelihood indicator.
[00104] In some embodiments, a chi-square score can be used to rank candidate peptide sequence or candidate spectrum matches.
[00105] Chi-square can be applied to the m / z values of observed and expected spectra lines for all matching ions as described by Equation (1) where the smallest score is highest ranked (i.e., wins): (ObsMzl-ExpMzl)2 ExplMz (0bsMz2-ExpMz2)2 Exp2Mz (ObsMzn-ExpMzn)2 ExpNMz Equation (1)
[00106] Alternatively, the chi-square score can be calculated by summing the number of expected theoretical ions (e.g., b and y ions) for the expected peptide that fall within the M / z range of the mass spectrometer; summing the number of observed ions that match the expected M / z values, within the mass resolution limit of the mass spectrometer and applying Equation (2). _ _ CObsMz-ExpMz)2 ExpMz+1 Equation (2)
[00107] In some embodiments, prior to applying Equation (2), a correction can be applied to the expected theoretical ion count. A correction can be applied to compensate for the M / z range and / or mass error of the mass spectrometer. In some embodiments, applying the correction involves multiplying the expected theoretical ion count by the total number of observed ions and dividing the product by the number mass spectrometer bins. The number of bins can be determined according to Equation (3). Number of bins = Ran9e Equation (3) ’ Resolution '
[00108] For example, for a mass spectrometer with a range of 50 to 2000 M / z units and a mass resolution of + / - 0.5 M / z units, the number of bins is equal to 2000-50 / (0.5-(-0.5)=1950.
[00109] Various regression methods can be used to generate a likelihood indicator in some embodiments. Linear regression can be used to generate a simple linear regression model for each candidate peptide sequence to identify the best fitting model. Multiple linear regression can also be used to incorporate additional explanatory variables into the model such as intensity.
[00110] In some embodiments, cross correlation can be used to generate a likelihood indicator. A cross correlation score of a sample relative to a candidate peptide sequence can be determined. For example, a scoring function can measure the similarity of an MS / MS spectra relative to a theoretical spectrum. Cross correlation can use a sliding dot product.
[00111] A method of determining a cross correlation score can involve first assigning an intensity to theoretical b ions and y ions. The square root of the ion intensities of the observed b and y ions is then calculated. The mean of the cross correlation over 500 lags is then calculated and subtracted from the cross correlation between the observed and theoretical ions at a lag of 0. Other methods of determining a cross correlation score can be used, such as the method described in Eng, J. K., McCormack, A. L., & Yates, J. R. (1994). “An approach to correlate tandem mass spectral data of peptides with amino acid sequences in a protein database. Journal of the American society for mass spectrometry”, 5(11), 976-989.
[00112] Other scoring functions can also be used, such as an ion intensity match ratio (IIMR), or an Amino Acid Match Ratio (AAMR).
[00113] As described, the ion intensity match ratio (IIMR) can generate a signal to noise score for each theoretical peptide match. The ion intensity match ratio can be calculated from the sum of ion intensities where there is a match to the theoretical spectra and is divided by the total sum of the ion intensities in the observed spectra. The IIMR value ranges from 0 to 1, in which 1 represents a perfect score and corrects for duplicate ion matches.
[00114] It is noted that the signal processing filters previously mentioned may have already reduced the noise. Since the ion intensity match ratio is determined after the signal processing filters have been applied, the ion intensity match ratio can be based on the Total Ion Current (TIC) for that spectrum. That is, the IIMR can be determined by dividing by the TIC for that spectrum.
[00115] The Amino Acid Match Ratio can be used to generate an ordered list of suspected amino acids in the peptide by comparing the m / z differences in pairs of observed spectra lines.
[00116] The Peptide Match Ratio can be calculated by taking the sum of the correctly ordered amino acids and dividing that sum by the total number of amino acids used to generate the theoretical spectra. Since the Peptide Match Ratio score is computationally expensive, the Peptide Match Ratio may be only run against the best scoring results.
[00117] Independent scoring algorithms can also be used to generate a single vector using Equation (4). The single vector can be more computationally efficient. c = ^{a2 + b2} Equation (4)
[00118] Neutral Loss ions and A ions can be considered diagnostic or confirmatory ions which contribute to a score. Ion pairs, which can identify an Amino Acid as described in the Peptide Match ion, can represent a secondary scoring mechanism related to these confirmatory ions.
[00119] A signal-to-noise filter can be applied to the one or more candidate sequences / spectra based on the likelihood indicators for the candidate peptide sequences / spectra.
[00120] Mass spectrometers, even without a sample, continuously generate source noise spectra. A signal-to-noise filter can be applied using a Monte Carlo random simulation to separate real data from random spectra by chi-square or other statistical means. The signal-to-noise filter can eliminate samples that show minimal significance when compared to random or source noise spectra. The random source noise can be generated either with blank solvents applied to naive columns or using random number generators. Peptide sequences matching to random MS / MS spectra at high frequency compared to real data by chi square or other statistical means are not carried forward for further analysis.
[00121] In some embodiments, random source noise spectra at high frequencies can be generated. A difference between the random source noise spectra at high frequencies and the one or more candidate peptide sequences / spectra can be determined. If the difference exceeds a pre-determined threshold difference, the sample can be excluded from selection as a proposed peptide sequence for the query spectrum. Otherwise, the sample can be selected as a proposed peptide sequence.
[00122] A candidate peptide sequence / spectrum can be selected as a proposed peptide sequence / spectrum for the query spectrum based on the filtered candidate peptide sequences / spectra. The selection of proposed peptide sequences / spectra from filtered candidate peptide sequences / spectra can be based on a best fit. The fit can relate to a peptide, accession, or gene symbol. In some embodiments, a machine learning model can be used to select the best fitting filtered candidate peptide sequence / spectrum to the proposed peptide sequence / spectrum.
[00123] In some embodiments, a signal to noise ratio filter can be applied after a proposed peptide sequence / spectrum for the query spectrum has been selected. In such embodiments, a signal to noise ratio filter may not be applied to the candidate peptide sequence / spectrum. For example, the number of spectra in the sample that match the proposed peptide sequence / spectrum can be counted and compare this count to a count for randomly generated and / or experimentally obtained spectra (e.g., inauthentic). Samples associated with counts that are similar to those of inauthentic samples can be excluded from further analysis.
[00124] A challenge with mass spectrometry relates to generating non-redundant datasets for analysis. One or more strategies can be employed to reduce the redundancy of the final data without a loss of information at the level of peptides, accessions, and gene symbols. For example, one strategy can involve assigning a unique identifier to each sample or MS / MS spectrum (e.g., SpectralD) and selecting the best fit per spectrum (BFPS) for each sample or MS / MS spectrum. The best fit can be determined by the likelihood indicator. In some embodiments, the best fit can be determined using a machine learning model.
[00125] Candidate peptide sequences can be generated by multiple combinations of methods, including multiple libraries and multiple random generation methods. A unique identifier can be assigned to each combination (e.g., search definition ID). There can be different approaches to calculating a best fit. In some embodiments, a best fit can be calculated by search definition ID, where for each search definition ID, a best fit per spectra is chosen independently of the other search definition IDs.
[00126] In some embodiments, the best fit per spectrum (BFPS) can be selected by Library and Search Engine. In the case where there are multiple static search methods for one library engine combination, the counts may be split off into the method where the SpectralD scored the best, or had the highest likelihood indicator. In this case there is minimal overlap between the two methods except where the SpectralD scored equally highest in both methods. Once the best fit per spectrum is identified, a non-redundant list can be made for each level that the data may be analyzed.
[00127] At the level of gene symbols, only unique SpectralD- SearchDefinitionlD-Genesymbol combinations are carried forward to create counts for all the subgroups for the experiment and are used for further statistical analysis where groups are compared by gene symbol.
[00128] The same is done at the level of peptides, where only SpectralD-SearchDefinitionlD-Peptide combinations are carried forward for count tables and analysis at the level of peptides. SearchDefinitionlD corresponds to the library and search settings. Library of peptide aptamers
[00129] The method disclosed herein can comprise providing a peptide aptamer library comprising a plurality of peptide aptamers, the plurality of peptide aptamers being derived or adapted from one or more natural peptides.
[00130] As used herein, the terms “protein”, “peptide”, “polypeptide” and the like are interchangeable and refer to any chain of two or more natural or unnatural amino acid residues. The terms encompass modifications, such as modifications to the backbone and modifications to side chains. [00131 ] As used herein, the term “natural peptide” refers to a peptide originated from a natural source, such as a biological sample. Examples of biological samples include but are not limited to cells, cultured cells, tissues such as blood, biofluids such as plasma, etc. A natural peptide can be or can originate from any protein, including ligands, receptors, variable domains, antibodies.
[00132] In some embodiments, the one or more natural peptides originate from a biological sample. In some embodiments, the biological sample comprises a tissue, a cell, and / or a biofluid. In some embodiments, the biological sample comprises a cell. In some embodiments, the biological sample comprises a biofluid. In some embodiments, the biofluid is plasma.
[00133] In some embodiments, the one or more natural peptides comprises an immunoglobulin superfamily (Ig superfamily) member, such as an immunoglobulin, a B cell receptor (BCR), a T cell receptor (TCR), etc. In some embodiments, the one or more natural peptides comprises an immunoglobulin, a B cell antigen receptor (BCR), a T cell receptor (TCR) or any combination thereof. In some embodiments, the one or more natural peptides comprises an immunoglobulin. In some embodiments, the one or more natural peptides comprises a B cell antigen receptor (BCR). In some embodiments, the one or more natural peptides comprises a T cell antigen receptor (TCR). In some embodiments, the immunoglobulin comprises IgG, IgA, IgM, IgE, and / or IgD. In some embodiments, the immunoglobulin comprises IgG.
[00134] In some embodiments, the one or more natural peptides are isolated from a host animal. In some embodiments, the host animal has contacted an immunogen. In some embodiments, an immune response has been induced in the host animal.
[00135] As used herein, the term “immune response” refers to activation of the adaptive immune system.
[00136] As used herein, the term "immunogen" means a substance which provokes an immune response, which can cause production of an antibody. The immunogen may comprise for example a protein, a protein fragment, and / or a nucleic acid, and may be conjugated to for example a carrier protein, an immunogenicity enhancing agent, and / or a particle such as a microparticle or nanoparticle. The protein or protein fragment may be, for example, a binding site or an ectodomain of a receptor. The immunogen may comprise a multiple antigenic peptide (MAP).
[00137] Contacting the immunogen with the host animal to induce an immune response can comprise injecting the immunogen into the host animal.
[00138] The host animal can be, for example, a mouse, a rabbit, a goat, a donkey, or a camelid.
[00139] In some embodiments, the host animal is a mouse, a rabbit, a goat, a donkey, or a camelid.
[00140] The immunoglobulins can be any type, for example, IgG, IgA, IgM, IgE, and / or IgD. In one embodiment, the immunoglobulin is IgG. IgG can be isolated from blood plasma, for example, by protein A / G chromatography, or immuno-affinity chromatography.
[00141] The one or more natural peptides can be an antibody fragment (Fab). In some embodiments, the method further comprises isolating Fab from the isolated immunoglobulin.
[00142] In some embodiments, the one or more natural peptides comprises proteins extracted from a biological sample. In some embodiments, the one or more natural peptides comprises proteins originate or extracted from tissue cells. In some embodiments, the one or more natural peptides comprises proteins originate or extracted from body fluids.
[00143] The plurality of peptide aptamers can be generated from peptides derived or adapted from natural peptides. For example, a natural peptide may be modified chemically or enzymatically into a peptide from which the peptide aptamers can be generated. For example, the native peptide may be reduced with DTT or mercaptoethanol, or alkylated with iodoacetic acid or iodoacetamide, or oxidized with performic acid, or modified with other modifications. The peptide aptamers may also be synthesized, recombinantly or chemically, to have sequences of the natural peptides.
[00144] In some embodiments, deriving the plurality of peptide aptamers from the one or more natural peptides comprises modification of the one or more natural peptides. In some embodiments, the modification comprises one or more of: (i) reduction; ii) alkylation; and (ill) oxidation. In some embodiments, deriving the plurality of peptide aptamers from the one or more natural peptides comprises reduction of the one or more natural peptides. In some embodiments, reduction comprises reducing with DTT and / or mercaptoethanol. In some embodiments, deriving the plurality of peptide aptamers from the one or more natural peptides comprises alkylation of the one or more natural peptides. In some embodiments, alkylation comprises alkylating with iodoacetic acid and / or iodoacetamide. In some embodiments, deriving the plurality of peptide aptamers from the one or more natural peptides comprises oxidation of the one or more natural peptides. In some embodiments, oxidation comprises oxidating with performic acid.
[00145] As used herein, the term “contacting the peptide aptamer library with the target” means allowing the peptide aptamer library and the target to interact by any means. In some embodiments, the target can be immobilized on a surface, and the peptide aptamer library can be allowed to come into contact with the surface. Contacting the peptide aptamer library with the target can cause one or more peptide aptamers to bind the target. In some embodiments, peptide aptamers that are not bound to the target after contacting the peptide aptamer library with the target are removed. In some embodiments, the bound peptide aptamers can be released from the target for analysis.
[00146] In some embodiments, deriving the plurality of peptide aptamers from the one or more natural peptides comprises digesting the one or more natural peptides. In some embodiments, digesting comprises digesting with one or more proteases. The proteolytic digest may be performed on native proteins, or proteins after modification, such as reduction with DTT or mercaptoethanol, alkylation with iodoacetic acid or iodoacetamide, oxidization with performic acid, and other modifications.
[00147] Non-specific and / or sequence specific proteases can be used. Examples of proteases that can be used include but are not limited to Arg-C, Asp-N, Asp-N (N-terminal Glu), BNPS or NCS / urea, Chymotrypsin, Chymotrypsin (low specificity), Clostripain, Glu-C (AmAc buffer), Glu-C (Phos buffer), Lys-C, Lys-N, Lys-N (Cys modified), Pancreatic elastase, Pepsin A, Pepsin A (low specificity), Prolyl endopeptidase, Proteinase K, Thermolysin, Trypsin, Trypsin (Arg blocked), Trypsin (Cys modified), Trypsin (Lys blocked). Examples of sequence-specific protease include but are not limited to Caspase-1, Caspase-2, Caspase-3, Caspase-4, Caspase-5, Caspase-6, Caspase-7, Caspase-8, Caspase-9, Caspase-10, Enterokinase, Factor Xa, Granzyme B, HRV3C protease, TEV protease, and / or thrombin.
[00148] In some embodiments, digesting comprises digesting with trypsin or chymotrypsin, or any other endo protease. In some embodiments, digesting comprises digesting with any exopeptidase, including C-terminal and N-terminal peptidases.
[00149] In some embodiments, the one or more native peptides comprise an amino acid sequence with at least one tryptic (R / K) or chymotryptic (W / Y / F) cleavage site.
[00150] Deriving the plurality of peptide aptamers from the one or more natural peptides can comprise chemical digest. Examples of chemical digest include but are not limited to CNBr, CNBr (methyl-Cys), CNBr (with acids), Formic acid, Glu-C (AmAc buffer), Glu-C (Phos buffer), Hydroxylamine, lodosobenzoic acid, Mild acid hydrolysis, NBS (long exposure), NBS (short exposure), and / or NTCB.
[00151] Digesting the one or more natural peptides or the one or more natural peptides that have been modified is optional. For example, as demonstrated herein, an undigested biofluid, such as an undigested serum, can be used as an aptamer library for screening of peptide aptamers that bind a target. In some embodiments, the plurality of peptide aptamers can comprise one or more undigested natural peptides. Furthermore, the plurality of peptide aptamers can comprise one or more natural peptides that have not been treated enzymatically or chemically. In some embodiments, the plurality of peptide aptamers can comprise one or more untreated natural peptides. Digesting of bound peptide aptamers prior to analysis
[00152] The method disclosed herein comprises removing unbound peptide aptamers after contacting the peptide aptamer library with the target. Before analyzing the bound aptamers, the bound aptamers can be digested. In some embodiments, the method further comprises digesting the bound aptamers after removing the unbound aptamers and before analyzing the bound aptamers. In some embodiments, generating a MS / MS query spectrum of the bound peptide aptamers comprises generating a MS / MS query spectrum of the digested bound peptide aptamers.
[00153] In some embodiments, the method further comprises releasing the bound aptamers from the target after removing the unbound aptamers. In some embodiments, the method further comprises digesting the released bound peptide aptamers. In some embodiments, generating a MS / MS query spectrum of the bound peptide aptamers comprises generating a MS / MS query spectrum of the digested released bound peptide aptamers.
[00154] In some embodiments, digesting comprises digesting with one or more proteases. Non-specific and / or sequence specific proteases can be used. Examples of proteases that can be used include but are not limited to Arg-C, Asp-N, Asp-N (N-terminal Glu), BNPS or NCS / urea, Chymotrypsin, Chymotrypsin (low specificity), Clostripain, Glu-C (AmAc buffer), Glu-C (Phos buffer), Lys-C, Lys-N, Lys-N (Cys modified), Pancreatic elastase, Pepsin A, Pepsin A (low specificity), Prolyl endopeptidase, Proteinase K, Thermolysin, Trypsin, Trypsin (Arg blocked), Trypsin (Cys modified), Trypsin (Lys blocked). Examples of sequence-specific protease include but are not limited to Caspase-1, Caspase-2, Caspase-3, Caspase-4, Caspase-5, Caspase-6, Caspase-7, Caspase-8, Caspase-9, Caspase-10, Enterokinase, Factor Xa, Granzyme B, HRV3C protease, TEV protease, and / or thrombin.
[00155] In some embodiments, digesting comprises digesting with trypsin or chymotrypsin, or any other endo protease. In some embodiments, digesting comprises digesting with any exopeptidase, including C-terminal and N-terminal peptidases.
[00156] In some embodiments, digesting comprises chemical digest. Examples of chemical digest include but are not limited to CNBr, CNBr (methyl-Cys), CNBr (with acids), Formic acid, Glu-C (AmAc buffer), Glu-C (Phos buffer), Hydroxylamine, lodosobenzoic acid, Mild acid hydrolysis, NBS (long exposure), NBS (short exposure), and / or NTCB. One or more candidate spectra
[00157] The one or more candidate spectra can be spectra predicted from nucleic acid sequences, such as genomic DNA sequences, or sequences of expressed RNA in a cell. For example, where the peptide aptamer library is generated from immunoglobulins of a host animal that has been exposed to an antigen, the one or more candidate spectra can be spectra predicted from the transcriptome, optionally the variable domain repertoire of immune cells of the host animal.
[00158] The one or more natural peptides or peptides derived or adapted from natural peptides comprise proteins extracted from a biological sample, such as tissue cells 29 or body fluids. In such case, the one or more candidate spectra can be spectra predicted from genomic DNA sequences derived from the biological sample.
[00159] In some embodiments, the method can further comprise sequencing the transcriptome or the variable domain repertoire of immune cells at the level of nucleic acids. This can be done, for example, by isolating immune cells (e.g. B cells) from the host animal, extracting total RNAfrom the isolated immune cells, reverse-transcribing the extracted RNA, amplifying the variable domain, cloning into a vector, and performing DNA sequencing. Deep sequencing techniques can also be used.
[00160] In some embodiments, the one or more candidate spectra comprise spectra predicted from a transcriptome of immune cells isolated from the host animal. In some embodiments, the one or more candidate spectra comprise spectra predicted from a variable domain repertoire of immune cells of the host animal. In some embodiments, the immune cells comprise B cells. [00161 ] In some embodiments, the one or more candidate spectra comprise spectra predicted from genomic DNA sequences.
[00162] The variable domains of the immunoglobulins, BCRs and / or TCRs can be identified against the variable domain repertoire by performing liquid chromatography electrospray ionization tandem mass spectrometry (LC-ESI-MS / MS) on the digested products.
[00163] Proteomic analysis can be performed by 64 bit computation using for example XTANDEM, SEQUEST, regression, goodness of fit, cross correlation (X-corr), the count of fragment ion matches, or heuristic algorithms.
[00164] Proteomic analysis can be performed in SQL Server and R using classical statistics. Target
[00165] In some embodiments, the target comprises a receptor or a fragment thereof. In one embodiment, the target comprises a ligand or a fragment thereof. In one embodiment, the target comprises an enzyme or a fragment thereof. In one embodiment, the target comprises a protein or a fragment thereof. In one embodiment, the target comprises an antibody or a fragment thereof. In one embodiment, the target comprises a variable domain or a fragment thereof. In one embodiment, the target comprises a drug.
[00166] In some embodiments, the target comprises a receptor, a ligand, an enzyme, a protein, an antibody, a variable domain, or a drug.
[00167] The target can be immobilized on a suitable surface prior to or after incubation with the peptide aptamer library, for example, immobilization on microbeads, nanobeads, a 2-dimensional surface, a 3-dimensional scaffold, and / or a 3-dimensional fiber. The target can be immobilized on the surface using any suitable method. The target can be immobilized with or without the use of a linker, optionally, a cleavable linker.
[00168] In some embodiments, the target is immobilized on microbeads, nanobeads, a 2-dimensional surface, a 3-dimensional scaffold, and / or a 3-dimensional fiber.
[00169] In some embodiments, immobilization comprises use of a cleavable linker.
[00170] In some embodiments, the target is immobilized prior to contacting with the peptide aptamer library. In one embodiment, the target is immobilized after contacting with the peptide aptamer library.
[00171] After incubation of the target with the peptide aptamer library, unbound peptide aptamers can be removed, for example, by performing one or more washing steps.
[00172] Bound peptide aptamers can be eluted from the target. Alternatively, bound peptide aptamers with the target can be released from the surface, for example, by cleaving the cleavable linker.
[00173] In some embodiments, the method further comprises removing unbound peptide aptamers after contacting the peptide aptamer library with the target and releasing bound peptide aptamers after removing unbound peptide aptamers.
[00174] In some embodiments, removing unbound peptide aptamers comprises one or more washing steps. In one embodiment, the one or more washing steps comprise washing with an aqueous buffer, a weak salt solution, a weak mixture of organic solvent with water, a weak acid or base close to neutral pH, and / or a mass spec compatible detergent.
[00175] In some embodiments, releasing bound peptide aptamers comprises eluting with a strong salt solution, a strong mixture of an organic solvent with water, and / or a strong acid or base far from neutral pH.
[00176] In some embodiments, releasing the bound peptide aptamers comprises cleaving the cleavage linker.
[00177] The amino acid sequence can be derived from the MS / MS spectra by de novo sequencing. Alternatively, the observed MS / MS spectra can be fitted to a predicted library of MS / MS spectra by 64 bit computation using de novo sequencing using cross correlation, XTANDEM, SEQUEST, regression, goodness of fit, cross correlation (X-corr), the count of fragment ion matches, heuristic algorithms, other algorithms, or combinations thereof.
[00178] In some embodiments, identifying the aptamers that bind the target comprises identifying by mass spectrometry, or top down mass spectrometry, or electrospray ionization, or MALDI ionization or chemical ionization or electron impact ionization or LC-ESI-MS / MS.
[00179] In some embodiments, identifying the aptamers comprises de novo sequencing.
[00180] In some embodiments, identifying the peptide aptamers comprises fitting of observed MS / MS spectra to a predicted library.
[00181] In some embodiments, fitting of observed MS / MS spectra to the predicted library comprises 64 bit computation.
[00182] In some embodiments, the 64 bit computation comprises use of XTANDEM, SEQUEST, or regression, or goodness of fit, cross correlation (X-corr), the count of fragment ion matches, or heuristic algorithms.
[00183] The identified peptide aptamers can be used, for example, as binding reagents against drug targets for therapeutic or biomarkers for diagnostic purposes. The drug target may be a ligand or a receptor or an enzyme or other.
[00184] The identified peptide aptamers can be used in screening assays, such as binding assays, enzyme linked immune-mass spectrometric assay (ELiMSA), and live cell affinity chromatograph (LARC), to identify new aptamer drugs.
[00185] The above disclosure generally describes the present application. A more complete understanding can be obtained by reference to the following specific examples. These examples are described solely for the purpose of illustration and are not intended to limit the scope of the application. Changes in form and substitution of equivalents are contemplated as circumstances might suggest or render expedient. Although specific terms have been employed herein, such terms are intended in a descriptive sense and not for purposes of limitation. EXAMPLES Example 1 - Identification of plasma peptides using mass spectrometry
[00186] Liquid chromatography electrospray ionization tandem mass spectrometry (LC / ESI-MS / MS) can be used to identify the ligands of plasma and identify their receptor complexes on the surface of live cells. We have found that results from HPLC LC-ESI-MS / MS linear quadrupole ion trap (LTQ) mass spectrometers and results from the orbitrap show excellent agreement. Examining orbitrap data we have discovered that in contrast to small molecules almost all ionizing peptides have heavy isotopes and hydrogen rearrangements and that only a minority of peptides are found at the monoisotopic mass but instead the peptide MS / MS spectra from precursor with MH delta mass values of -1, -2, -3 and the heavy isotopes at MH delta mass of +1, +2, +3, +4 and +5 all match and far outnumber the peptides at the monoisotopic mass at 0. Thus, peptides ionize as an envelope that is 9 Daltons wide and so there is essentially no point to using mass spectrometers that have much higher mass accuracy than that of Nature for the precursor. Instead, a wide range of precursor masses need to be considered and the MS / MS spectra computed to reveal the identity of the peptides. The LTQ is entirely sufficient to identify and quantify the observation frequency of peptides and proteins from plasma protein ligands or their receptor complex on the surface of live cells. However commercial software packages often re-use MS / MS spectra and this redundancy needs to be removed in a relational database management system. In addition, all analytical experiments require a blank and so we maintain a database of peptides from blank noise injections (Albumin and antibodies from the skin) to correct experimental observation frequency. Example 2 - Statistical approach
[00187] The decoy library method for proteomics does not rely on any known statistical distribution or test but is rather an arbitrary competition for significance against and arbitrarily chosen library that does not generate an FDR q-value. We have reevaluated 18 non-human protein standards used for the decoy library method and shown that the so called empirical statistical model is based on the incorrect assumption that sigma protein standards are pure when they in fact contain hundreds or thousands of proteins and so the empirical statistical model showed a 99.7% false negative rate. In agreement with the known hydrogen rearrangement and heavy isotope or peptides the optimal search conditions for fitting peptide MS / MS fragmentation spectra are -4.5 to + 5.5 Da. Accordingly, we have instead applied classical statistics to the problem of proteomics using the SQL Server and R statistical system.
[00188] Classical statistical methods like ANOVA, regression or Chi Square used in biomedicine, agriculture and engineering applications all require that the data is randomly and independently sampled from a population with a known distribution. We have shown that intensity may be log transformed to yield a normal distribution and the observation frequency is a discrete distribution that follows the alpha n=1 distribution and so is amenable to the Chi square (x2) test. Classical statistical methods like ANOVA, regression or Chi Square all have a null hypothesis (Ho) and test statistic (F, x2) and a critical value to generate a p-value from the known distribution. The p-values are generated and corrected by the FDR method of Benjamini and Hochberg. We employed two independent statistical methods to fit the peptides by MS / MS spectra to the theoretical fragmentation spectra and the Monte Carlo simulation versus millions of random MS / MS spectra to determine the type I error rate of our plasma ligand and cell surface receptor experiments. We have shown that the identity of peptides and proteins can be computed entirely in SQL Server and R using classical statistics and showed excellent agreement with the XITANDEM and SEQUEST algorithm after correcting for redundant use of MS / MS spectra to ensure that each MS / MS spectra is never fit to more than one single peptide sequence that eliminated virtually all error in proteomics. Thus, there is no need for high resolution orbitraps for the proteomic analysis of proteins from the fit of MS / MS spectra to tryptic peptides. In contrast to peptides, small molecules are in high abundance, have few hydrogen rearrangements and heavy isotopes and few fragments that can be fit by regression or Chi Square and so insensitive high resolution mass spectrometry like Orbitraps or Qq-TOF is entirely appropriate for small molecule analysis but are not required for the analysis of observation frequency of tryptic peptides for plasma and receptor discovery. The observation that classical statistics in SQL Server and the R statistical system may be used to describe blood plasma ligands and their receptors of the surface of cells with great sensitivity to the pg range and defined type I and type II error established from regression, ANOVA, or Chi Square. Example 3 - Creating peptide aptamer library from naturally occurring proteins and peptides
[00189] The advent of large molecule drugs such as proteins (insulin), peptides (octreotide, glucagon), antibodies and nanobodies borrow from the products of nature to create new diagnostic and drugs treatments. The immune system can create highly selective and specific molecules from the Ig superfamily in response to disease that may serve as diagnostic and therapeutic agents.
[00190] New binding reagents to target or ectodomain of the target human Fc receptor or other receptors can be created by using digests of naturally occurring proteins and peptides from cells, fluids or plasma as aptamers, and / or expressing the receptor ectodomain or ectodomain binding sites or peptides from ectodomain binding sites on MAP peptides of the ligand binding site, and presenting the protein, peptide or MAP peptides to a host animal.
[00191] Natural proteins from cells, biofluids or plasma or derivatives thereof may be digested by a protease.
[00192] Receptor, ligand, enzyme, protein, antibody, variable domain or drug will be immobilized on microbeads or nanobeads or a flat of 2 dimensional surface or 3 dimensional scaffold or fiber.
[00193] The receptor, ligand, enzyme, protein, antibody, variable domain or drug will be immobilized prior to or after incubation with a digest.
[00194] Natural peptides may be obtained from extracts of tissues, cells, biofluids or plasmas.
[00195] Natural peptides may be obtained from digests of tissues, cells, biofluids or plasmas.
[00196] Natural peptides may be obtained from isolated and then digested IgG, IgA, IgM, IgE , IgD or TCR or BCR from blood plasma especially where the T-cell or B-cell VDJ variable domain repertoire has been sequenced at the level of nucleic acids.
[00197] Proteases include non specific proteases such as Arg-C, Asp-N, Asp-N (N-terminal Glu), BNPS or NCS / urea, Chymotrypsin, Chymotrypsin (low specificity), Clostripain, Glu-C (AmAc buffer), Glu-C (Phos buffer), Lys-C, Lys-N, Lys-N (Cys modified), Pancreatic elastase, Pepsin A, Pepsin A (low specificity), Prolyl endopeptidase, Proteinase K, Thermolysin, Trypsin, Trypsin (Arg blocked), Trypsin (Cys modified), Trypsin (Lys blocked) or sequence specific proteases such as Caspase-1, Caspase-2, Caspase-3, Caspase-4, Caspase-5, Caspase-6, Caspase-7, Caspase-8, Caspase-9, Caspase-10, Enterokinase, Factor Xa, Granzyme B, HRV3C protease, TEV protease, Thrombin. Chemical digests may be used, such as CNBr, CNBr (methyl-Cys), CNBr (with acids), formic acid, hydroxylamine, iodosobenzoic acid, mild acid hydrolysis, NBS (long exposure), NBS (short exposure), NTCB.
[00198] MAP peptides can be injected into mice, followed by harvesting the B cells and PCR amplifying the variable domains and creating a variable domain library for mass spectrometry. Subsequently the IgG from the mice plasma will be isolated by protein A / G chromatography, digested and the variable domains identified against the variable domain library using LC-ESI-MS / MS. Example 4 - Using the aptamer library to identify binding agents to a target
[00199] The aptamer library from natural peptides will be incubated with the receptor, ligand, enzyme, protein, antibody, variable domain or drug to induce binding.
[00200] The receptor, ligand, enzyme, protein, antibody, variable domain or drug will be immobilized on microbeads or nanobeads or a flat 2 dimensional surface or 3 dimensional scaffold or fiber. [00201 ] The receptor, ligand, enzyme, protein, antibody, variable domain or drug will be immobilized prior to or after incubation with the aptamer library.
[00202] Unbound peptides will be washed away with aqueous buffer, weak salt solutions, weak mixtures of organic solvents with water, weak acids or bases close to neutral pH, and / or mass spec compatible detergents.
[00203] Bound peptides will be eluted with strong salt solutions, strong mixtures of organic solvents with water, strong acids or bases far from neutral pH.
[00204] The eluted peptides will be identified by mass spectrometry, or top down mass spectrometry, or electrospray ionization, or MALDI ionization or chemical ionization or electron impact ionization or LC-ESI-MS / MS.
[00205] The amino acid sequence will be derived from the MS / MS spectra by de novo sequencing.
[00206] Alternatively, the observed MS / MS spectra will be fitted to a predicted library of MS / MS spectra by 64 bit computation using de novo sequencing using XTANDEM, SEQUEST, or regression, or goodness of fit, or heuristic algorithms, or other algorithms. Example 5 - MS / MS computation
[00207] Real peptides and their fragments frequently contain heavy isotopes and so a filter to look for spectrum lines where there is a presence of isotopes of hydrogen rearrangements or hydrogen loss or hydrogen rearrangements to use isotope filtering to remove the noise.
[00208] Isotopes and hydrogen re-arrangements or losses may occur in precursor peptides or fragments.
[00209] The de novo of the MS / MS spectra or the fit of the a,b,c & x,y,z fragment series, or regression, goodness of fit or cross correlation to generate the top fits.
[00210] The observation frequency of peptides or polypeptides will be compared to those of random MS / MS spectra to identify true positive peptides. [00211 ] Selecting the top fitting 10-50 peptides will be filtered from the noise by the use of observation frequency versus the Monte Carlo at the level of proteins, polypeptides or peptides to select the true positive results.
[00212] Blank injections with no aptamer injected would serve as an analytical control.
[00213] True positive results will accumulate in a subset of peptides or variable domains while false positive hits will show a random distribution.
[00214] Thus, a series of signal to noise filters along with classical statistical analysis and specifically designed aptamer libraries to make the computation of synthetic random aptamers feasible.
[00215] The receptor or receptor ectodomain may be from the family of innate immune receptors including the Fc receptors or scavenger receptors; or adaptive immune system including B-cell or T-cell receptors; or growth factor receptors including the insulin receptor; or metabolic receptors including insulin or glucagon receptor; or growth factor receptors such as IGFR, or EGFR or FGFR or LEDGF or FLT3; or other receptors. Example 6 - Aptamers generated from IgG or immunoglobulin variable domains
[00216] Tryptic or chymotryptic or other proteolytic digests of proteins, blood, cells, naturally occurring variable domains etc that bind against ligands or receptors can be analyzed using affinity chromatography with low resolution LC-ESI-MS / MS and statistical analysis of the observation frequencies of the potential peptides versus control affinity columns and statistical controls such as blanks, electromagnetic noise or random MS / MS spectra.
[00217] Methods
[00218] A total of 15 mice (three for each receptor ectodomain target) were assigned and the pre-immune serum (or plasma) collected. The receptor ectodomain targets are listed in Table 1. Table 1. Receptor ectodomain targets Receptor MAP Receptor ectodomain antigen FCGR1B MAPI TQTSTPSYRITSASVN EGFR MAP2 SMDFQNHLGSCQKCDP FLT3 MAP3 VSESPEDLGCALRPQ SS FGFR4 MAP4 SNDDEDPKSHRDPSNR INSR MAP5 LHHKCKNSRRQGC
[00219] Polypeptides from the ectodomain of receptors were injected into mice and after immune boosting injections a sample of serum (or plasma) of the mice was collected. After strong immune response was detected the serum or plasma was collected and B-cells were obtained and the RNA extracted for reversed transcription, PGR amplification of the BCR variable domain followed by DNA sequencing that resulted in more than a million DNA sequences from each mouse.
[00220] Immunoglobulins from the immunized mice were purified against the antigenic polypeptide immobilized on an affinity (or bispecific) chromatography resin.
[00221] The following experimental treatments and controls were performed: 1. Chromatography support presenting protein or polypeptide target (resin) with immune serum 2. Chromatography support presenting protein or polypeptide target (resin) with Fab fragment 3. IgG receptor (CD64) ecto domain resin + mouse immune serum 4. EGFR ecto domain resin + mouse immune serum 5. FLT3R ecto domain resin + mouse immune serum 6. FGFR4 ecto domain resin + mouse immune serum 7. INSR ecto domain resin + mouse immune serum 8. Chromatography support presenting ectodomain but with pre-immune serum 9. Chromatography support without ectodomain or target with immune or pre-immune serum.
[00222] After washing away the unbound fraction, immunoglobulins specifically bound to the antigen were eluted and digested with trypsin or chymotrypsin or otherwise proteolytically digested. The resulted digested peptides were separated by chromatography for electrospray ionization tandem mass spectrometry with a linear quadrupole ion trap (FIG. 5).
[00223] The MS / MS spectra from peptide digests of immunoglobulins that bound the antigen was matched to the predicted MS / MS spectra from the millions of BCR variable domains (-100 amino acids in length) that may contain many potential tryptic, chymotryptic or other proteolytic peptides.
[00224] Since the reading frame and read direction of the millions of BCR variable domains is not clear, all 6 reading frames from both directions of the theoretically complementary DNA reads (12 possible full length polypeptides with potential binding activity) of the PCR product were obtained.
[00225] The observation frequency of polypeptides from the experiment were compared to (1) those of electromagnetic noise or computer random MS / MS spectra; (2) analytical controls of naive C18 columns or conditioned C18 columns; (3) statistical controls of random MS / MS spectra; (4) experimental controls of Control resin (no antigen) + serum; and (5) experimental controls ectodomain resin + pre-immune serum.
[00226] Representative polypeptides identified from the above experiments are listed in Tables 2-6. Table 2. Representative polypeptides identified as potential binders of MAP2 (EGFR) ecto domain from affinity purified mouse IgG variable domains Peptides (Query samples) Trypsin digest, DTT treated Trypsin digest, no DTT Wash contro 1 Naive colum n Conditione d column Rando mSpec tra Candidate polypeptide* CSGSER,CV RYGK 23 2 2 3 3 1 RCSGSERCVRYGKCSYIEP WWWAD LFQE PTSMXCVD V MXVRRVVPSHLQPDPLXXG SRSRWTADFYNVEXXQMSY YXTCASLPPGSTFILPDYDV VDVGR ACDGDT,VR STVK 13 0 0 1 2 2 DGCSWCSISPFYFQLGPPS ERVRSTVKLLTVITYKIFSLX AADLERKFCARSTTTEPXW DPICQAGCSINQELRRFPWF LLIPCXPCTNGLTCRACDGD T MAKNPPGLS K,VPSIQGK 10 25 3 0 0 2 TFRSVSXHVSSLVLINHTMD VVGAVRGFEYEMFQRMAK NPPGLSKSPPPRPCLKWGV XSSHLRLLDPQIITKVPSIQG KCKPSPCFNLYFARLRRGX RRX KKTTQQTKK ,TTQQTK 8 0 2 0 2 1 MDAVGAASARFSSSLVPAP NVSHSEDYDMPLSTPVEAV PLSKIGRAHVXTPVTPENLV CRLLLEKKKTTQQTKKTNKT NDTTEXYITDLVYTIDTSVRD STLGNR,TTK KTK 8 0 2 1 2 0 NXHTRKXTTTKKTTKKTKKK KVRLLPYALRGLTDLKSAHE KARSTLGNRMPLLSWSTTG VSKTQPKISNTIQKPPPWFQ XQDPVGFRXGXEWHKGGR RL PSSSGIGR,Y GSGSR 8 0 0 0 0 1 PSSSGIGRHGRAHIQCTXP WSSSVLE EG LRTNYRVGWT EDLRDERSLRHPDPGKYGS GSRHTSDFYDGXXXRXSPH QXPASHLRASTFSFPDYDV VDVGR ALPEDSVRC TRHTGGRP GRITTYPVS R,HTGGRPG RITTYPVSR 6 0 0 0 0 0 ALPEDSVRCTRHTGGRPGR ITTYPVSRLWSDDDPSGPET QRVGLDPYIPTPKTVGRWT DEYLVGRSHRNFLVKVCFX DLERLWAVAFTTIFPWTVQG GES HSDSCYSSS SRSGVPDIC VGGRVGCT SSR,TVDSC DDR 6 0 0 0 0 0 IXPTSSSHLSLNSPNLPKYD LRHSDSCYSSSSRSGVPDI CVGGRVGCTSSRRTVDSCD DRXRHVLPXXIXXYDTXXPQ FLGVSGRGVHYFSLDCTRR RV CTGGNR,YK CTGGNRRR 5 3 1 0 0 0 RLYVYVLDRTEHGCKKRSP SSPLPVFLAKYKCTGGNRR RKDLFCHXMTGFCSGRLRS XLIAPETSLTIXPFSFNLVSEV PSRVCSYKVTVPDYDVVDV GR EGGGGGR,E GREGGEGK 5 1 0 0 0 1 VSVVVNVPFGSVIVHRANH GGGRKGERSRRGGKGERE GREGGEGKGKGKQKREKK EGGGGGRRMVIKGNNNRC DTIRHLQATSVVRPLVCPTT TWLTXV FCCGGRR,S RFCCGGR 5 0 1 0 0 0 RDSTPTLYPVELLNSPQVPR RRVRPLSYTLLLNLPTKNTL WSLTVLQTQRVPTKISLRPQ KQPISPFXVITLMIRYLMTPV PWSQWQRSRFCCGGRRXX RLSVGK,VCI LMLTK 5 0 1 0 0 0 SXRPYPHWTXAVGSPFPDL TDRCXTTLLLLAMLGRDFSS ECQRFLWRLLVRKDLXXXHL XRLXRYMMTRVCILMLTKR MTPVPXDQXQRRLSVGKGL QK LPNQKKK,LP NQKKKK 5 0 0 1 0 1 CEYSAFLKIHKLPNQKKKKV RLLPYALRGLTDLKSAHEKA RVPAFTYXMPPLLSWSTTG VSKTQPKISNTIQKPPPWFQ XQDPVGFRXGXEWHKGGR RL FCCGGRR,S RFCCGGR 5 0 1 0 0 1 SSIRYRTQAVXGLFSDLTQR WXSSPPSMWMIGLSHFPAK WXRSLLRFLWDMDVYSSDS RLLCRYIMTRSVQLPMKLLM TPVPWXECQRSRFCCGGR RXVT SPSHRSCLK ,WTSRSLSQ LSIQLSK 5 0 0 0 0 0 WTSRSLSQLSIQLSKYVTIV VFGPVGGFEEXIERRLDLRP QGRSKSPSHRSCLKWEWX LGHLRLLRRWIMTVVLLLLG MCKPPPWFDLYFARLRRGX RRX *“X” refers to where an amino acid cannot be identified Table 3. Representative polypeptides identified as potential binders of MAPI (FCGR1B) ecto domain from affinity purified mouse IgG variable domains Peptides (Query samples) Observation frequency of control** Observation frequency of query Candidate sequence* GLGGIAA 7 84 SIAIGDLRGVGSCPEFPNLHIQKSA LSPVTMKLWLNWVFLLTLLHGIQCE VKLVESGGGLVQPGGSLRLSCATS GFTFSAFYMEWVRQPPGKGLGGIA A AGGTSGS 20 76 GRHLGRTDSLRRRXLRFLDPSSPX HSPRYYRSHGFPWNSSTWLCPHL SQSSAGXRICSEMCLWRWCCGIX WMDYRILYLHNICAPSTLTPFLRAG GTSGS AAASEP 10 68 EDIWEGLTLXGDGERGALAPVVKS YATTTVLHGPVLNANVFRSXGQDL QVDATFESNSCGDEPSFQRVSIVT TIVIDLPXPFIYTLTNLLXSKAAASEP SSSGIGRHGR 0 66 SSSGIGRHGRVHIQCTXPRSSSVR EEGLRPKYRVGWTEDLRDERSHR HLDPGEXESVSRLTSDFYDGEXXR SFPHQWASASHDPGSTSTLPDYDV VDVGR APVGGR 4 61 APVGGRIYGTSVQTWTGLLLHSCC XLSLHMSCPKLLXKSLALGYXSPH RPSVXLVLSLGFHXALLVWVXAGFV SLQGRVWSGWPTFGWMMISTITH PXRP STATEANR 0 55 DGCSWCSISPFYFQLCPRAEREW VTTPLLTVISGSIFSLHAADCERIRG PRSTATEANRDSRSQVGCAINPGF GGGFXLLLVPVQVAGTYTXADTAG DGD IPSSAR 5 37 GRHLGRTDSLQRQXPESLGPSGY PLSHSNTCLCLQFSSCSFADRACF VNDLWRWXTCLSLPXILXVLIRDRK STCLNSSHITPSRIPSSARNKKSDS SNN SDSGRHGR 0 35 SDSGRHGRVHIQGXXTXPXSSSVL RGDLRTKYPVGXTEDLRVERSHRH PDPGEXESVNRGTSDFYNGEXXR SSHHHMACASLPPGSTFILPDYDV VDVGR APGGAE 8 35 GDTSVGLRGSSXGPKFKDKMDFQ VQIFSFLLISASVILSRGQIVLTQYPA SMSASQGEKVTMTCRASSRRRXK HGYQQKPGNAPKRWSKDTPKRAP GGAE TVVLISRGK 0 32 SQWYWTFRSVSHSLLHRTMVVFG PVRGFDDYMIRRLAMXPQGLAKXP SPIPCLKXKWXSXHVRLLDRQIKTV VLISRGKCKPSPCFNLYFARLRRGX RRX LPGGGR,YTADGK GK 4 63 DGCSWCSISPFDFQLGASTERPRG EVLMTAINCQDFSRYTADGKGKISP RCTAGEETRETRLPGGGRVDQQN RRRPWCLLVTGQEVVRIXSVENTLT GLT ARGGDA,VTNNLTR 8 58 DGCSWCSISPFYFQLGPPSERARR ITIMLSTIIVHIFIRHNADGERRIFPRT TAAEAPWGASKQYGGSVEEQARR SSWWLLLQAEKVTNNLTRRARGG DA ASEPSCCVGGRJP SSAR 7 37 EDIGEGLTLCRDSDQSPLAPVSKP GARAVSLXQXLAQNPLRASEPSCC VGGRHGXVSPPVSVRELLSDRKST RLNSSHITRSRIPSSARIKXLMLSIH P ARGSGT,ASNLEAG IPARARGSGT 3 18 KSSPEXYGEMETDTLLLWVLLLWV PGSTGDIVLTQSPASLAVSLGQRATI SCRASERVDSYGNSFMHWYQQKP GQPPKLLSYRASNLEAGIPARARG SGT AAGGRA,RAAGGR 2 9 GXCSVSTGISXGPKFKDKMDFQVQ IFSFLLISASVILSRGQIVLTQSPAIM SASPGEKVTMTCSASSSISYMHWY QQKPGTAPKRWMYDTSKRAAGGR A KYAGDSEGTS,YAG DSEGTS 0 9 RMQLVQHQPGSAPAWYQHRTXAD ILDVAYSNKHPHPQPPLYXFSVXKQ FLTHCHXTCLGLLRQGWTSDISGA EETGLASGGTKTRKWGKKYAGDS EGTS TPGLAK,TPGLAKC PSPR 2 8 PSQWNLRSGSHPRSHRVIMVVVS ARGWCDTMSRCRANRTPGLAKCP SPRRCLKREWXSLPVRLLQRLIKTV VISSLGDQCKPSPCFNLYFARLRRG XRRX GSASPARPGRR,R KAGTNR 0 7 SVRVESTRGKDIRTACASRWSHRH MSLYTCRCGCLVLRETVXXPTLKNT CPHEXERGSASPARPGRRWERRK AGTNRNQGKQPKERGNRQRSGR GERQSE ASGYTFTSYNMHW VKQTPGQGLEWIG YIYPGNGGTNYNQ K,SGASVK 0 7 RKGKQSIRGKTLTLTMGFSRIFLFLL SITTGVHSQAYLQQSGAELVRSGA SVKMSCKASGYTFTSYNMHWVKQ TPGQGLEWIGYIYPGNGGTNYNQK FKG CSCQEAR,TPRCS CQEAR 0 6 RRPWVESHFSDPLKPXAWCRDRS TVH RSTEAQPHTQTRGVQRGILPF RENLPATQPPSPAPGLGTQGKDAR SPRTPRCSCQEARIGCSIGCQSVR ARAA *“X” refers to where an amino acid cannot be identified ** Wash control Table 4. Representative polypeptides identified as potential binders of MAP3 (FLT3) ecto domain from affinity purified mouse IgG variable domains Peptides (Query samples) Observation frequency of control** Observation frequency of query Candidate sequence* SSSGIGR 0 116 SSSGIGRHGRVHIQCTXPRSSSVREEGLRPKYRV GWTEDLRDERSHRHLDPGEXESVSRLTSDFYDG EXXRSFPHQWASASHDPGSTSTLPDYDVVDVGR EVPDPLPLK 0 112 MDAVGAASARLISSLVPPPNVHGXLLRXWQXXVA ASSASMLLIVREXEVPDPLPLKRAGTPEASLDVSX IHLLGEVPGFCWYQCMXLILELALQVMVTF TVVILSIGK 0 110 SQSQWTFRSVLHPXLHRTIVVFGPVRGFRDXMS RRMAMSPQGLAKCPSPRPCLKXEWXSLHVRLLN RLIKTVVILSIGKCKPSPCFNLYFARLRRGXRRX GLPGGR 5 54 GLPGGRXEXNFSSVWGVRALEIADKTLFSNQGL SSLLTVXPLVLWVXDTDHTPNSESQGHQRLTTFQ NXFASWCLYDTXXPQFLGVSGRGVSQSGRVYRR ARGGMG 3 53 HLISKSTGGIEVTLELDDEPXHPYMNVGFSKQVW RGGGGQGVNLXLFLCSGVRXKGKLKNFFIPEPXK EREWREDQNMNRGRGAGCAGQDRRRARGGM G AAGADS 11 43 DGCSWCSISPFYFQLGPPSERVRRVVVLLTKILCX DFXRQIADGGGGVFCRSTACGAMSGSSVLVGCR VNHQCCRXLWLLLILACIITNILPSRAAGADS TCTEHVYFESR 0 39 RWMQLVQHQPVLFPAWSPLRTCTEHVYFESRNK LPDPQPPLCXSXVXNLSLIHCHXTCLGPQKIGWK LCRSGALETGLASAGTNVNRCFHYCVQGSDXIC RAGTPEAR 0 38 MDAVGAASARFISSLVPPPNVYGLLLHCWQXXVA ASSASMLLIVREXEVPDPLPLKRAGTPEARLDVR XIQGLGEDLGFCWYQYMXLTLELALQVMVTF AASGSR 2 38 AASGSRRHRRVHIQCTCPSSSSVRXGHCLPKYR VGLTEDLRTVRSHRHPDPXEXEGISRRTSDFYDG EXXQSSPPHRAGASHTPGSTSTLPDYDVVDVGR SGTPMYR 0 32 MDAVGAASARFSSSLVPAPNVRGXLXNCAQXXC VRSSTCTLLMVRVKSVPDPLPVKRSGTPMYRLDA PYISSLGDCSGFFWYQAKXCTLCLLEXRLSLA EVPDPLPLK,MD AVGAASARFSS SLVPAPNVSG 0 176 MDAVGAASARFSSSLVPAPNVSGXLLLCWQXXVA ASSASIRLIVREXEVPDPLPLKRAGTPEARLDVLXI QSLGEVPGFCWNQCMXLTLELALQVMVTF DAVGAASARFIS NFVPEPNVNGE R,EVPDPLPLK 0 22 MDAVGAASARFISNFVPEPNVNGERXYWWQXXV AASSASMLLIVREXEVPDPLPLKRAGTPEARLDVL XIQSLGEDPGFCWYQCKXLELTLELAVQVMV DSTVPVGCRVN qcfrr,rlpwfl LIPGYISTHILTCL AGDADP 0 16 DGCSWCSISPFYFQLGPPSERVRIAVILLTEILCQV FRLHIADGESEICPRSTACEAIRDSTVPVGCRVNQ CFRRLPWFLLIPGYISTHILTCLAGDADP RVSQKGS,VSQ KGS 3 13 ANAGCWIYGDCSQDSAWTXGFLLTFLASCCSGF QVPDVTSRXPSLHPPYLPLWEKESVSLVGQVRKL GVTXAGFSRNQMELLKAXTTPHTRERRVSQKGS GKDVTS,SSHHY LRGKDVTS 3 11 RWMQLVQHQPVXFPAWCLHRTSEENVYLASNNK LPNPQPPLCXFSVXNLSLIHCQXTCQGLQSPVXT PDRSGAXETGLASVTTNSNKSSHHYLRGKDVTS LCFHHSISLPAT PD,LCFHHSISLP ATPDPFLEAGEP SVHHSWLMRTL 0 9 GRHLGRTDSLRRRXLRFLDPSSPXRNHVWHSST WLCHQFGDCSFLRKLGSWSCPCXCSVWIXELNY RLCFHHSISLPATPDPFLEAGEPSVHHSWLMRTL GGLGPK,VGGT EDPRDER 0 9 SGSGRHSRVHIQGRXTXPWSSSVLRGGLGPKYR VGGTEDPRDERSHRHLDPGXRESVSRRTSDFYD GEXXQLSPHQGASASHDPGSTSTLPDYDVVDVG R METVTSGCDTE YHDARTTKCSG, TTKCSG 2 8 RWMQLVQHQPVXFPAWCLHRTSTEDYHTVDNN KXQHLQAPGCXCXRNNPFQTHFHXTLMGCLARN XWXLCLPELQTVRMETVTSGCDTEYHDARTTKC SG LSGRGLCYGSC GTKTFTR,LSGR GLCYGSCGTKT FTRQILSVR 0 8 NFSHSTSTTSDPLRITSDLPGPLRGHVGDLSESH RYGTEPRRSEASSPTSPSVGNHDYHQSRXLVCDI FRLSGRGLCYGSCGTKTFTRQILSVRKGLQK QVRGTF,TSSW GECLDSAGNK 0 7 MDAVGAASARFSSSLVPAPNVRGEYAIGGRNKVP DTQPPLGWMXGGKESQNPCHXNEPGIQRAGLK GYKTSSWGECLDSAGNKAWXLTLECDRQVRGTF *“X” refers to where an amino acid cannot be identified ** Wash control Table 5. Representative polypeptides identified as potential binders of MAP4 (FGFR4) ecto domain from affinity purified mouse IgG variable domains Peptides (Query samples) Observation frequency of control** Observation frequency of query Candidate sequence* AVGAASARFSS SLVPAPNVSG 0 177 MDAVGAASARFSSSLVPAPNVSGXLAYCRQX XSAKSSDSRLLMVREXSDPDLLPLNLFGTPE SKVDAAXIRRLIVPSGFCXSQLKXPLISXLARQ VRLT PSGTRR 0 102 PSGTRRSGHSYTHEYIGPXLSLVSSEDLTTIC PVGWPCDPRGXRSVRHLDVVXSETGSRHTS DFWNVXXXHLSQCRXVSASRAPVSTFILPDY DVVDVGR RTSKTLNNDSC K 0 86 TVVVYSSFTLISVNKSYHSRLCPFXRSRGPDI TFLESSTTRXFQVTVTXSLCYKRVLVVGRRTS KTLNNDSCKNLMRHLQATSVVRPLVCPTTTW LTXV GGSGIK 3 77 WPFLVXLSNSXIIYSSVXCXLLGILPSCSFNFV HFXKFEDGVGTFIXXTSFFYFGGXXHVRPIYQ PXKCSINVFLDSSIIKEDCFSNYVCKGGSGIKR R GGSGEGG 4 68 KFVLPTYGEMNTVFSTVTESQGPYNEMQLG HLLPDGSGYRNQFRGSAAAVWGRACEVRGL SQVVLHSFWRQHERQQKARGEAEAGKGRE GEGRGGSGEGG LNPQAA 9 67 GKTFGKDXLSEETVRVVPWPQXSQFAQXYM SVSSDLRLVICRNRVFFLLSLXMLNLPFTVSA XXMVLLPLLLYATHPISFSRAXRTHCISQLLKL NPQAA APGGPR 4 64 APGGPRSFPKPYLTSVFFIFIINTLYNTAKSLSV QVVEVSTKCFVRDRRLLVELLTPVSVHNDTLC GPNDAIIDGTIPDDPSSLESVAEESLSQEGFT EG SGTGRPGR 0 56 SGTGRPGRVHIQGQXTXPWSSSVHGGGLRP KYRVGXTEDLRDERSHRHPDPGEXESVSRH TSDFYDGEXXRSSCHQWVSASHDPGSTSTL PDYDVVDVGR ASLGGGS 1 50 NSRDRTRRRKAESRKKKKLRGRGWLXMXYR MXERDVLREGXIPSFYHPNPTGSCHWNQGG GFCIVLLILGWVLDTPVVDQDKRVSTRLSLMR ASLGGGS DTMSWVYK 0 33 PRQCXSLIRLTQVLFGLVNKXPDYPPWSLAR GPQGQSKSPRDXPLFRREWXCPRVXLLLRYI KTRDTMSWVYKSRHLGSSDRIQSGSGEGES GTKVEGDS TWAGYYSTAG SK,VGAAGK 16 52 RWMQLVQHQPVSAPAWSQHRTWAGYYSTA GSKKWQHLQPPCCXWXESKRHQTHCHRSE QGLQKPGWRXEKSREWGRKVGAAGKSTSN XHLSWHCRSWXPS RGGSGC,RRG GSGC 4 21 VKMXIDTGRGGQSWNXFPVPHVQXXAVNTD PSPXTSGSDXFSLSLLXKVSSVTXSWWSLGE AXXSLEGPXNSPVQPLHSLSVAIPCLGFARLR RRGGSGC QAPPTNTGRR K,QGAAGG 0 17 FLKRLRYGDQNSKTKWIFKCRFSAYCXSERQ SXCPEEKMFSPSLQQSRLHPQGKRATXPAG PAQGXGRGTGTSRRQAPPTNTGRRKQPNR QTERQGAAGG GAGTAR,GAGT ARYR 0 14 MDAVGAASARFISSLVPPPNVYRXLLCCRQX DSANSXDWTLMVGRVKAVPDPLLMKGAGTA RYRXDAEXSSSLGDCAGCCXXYDTLLTTCLP ALQVATI VTKSGK,VTKS GKSSQSLRES SNQK 0 12 YTYYVDLRGGRGSKMDSQAQVLMLLLLWVS GTCGDIVMSQSPSSLAVSVGEKVTKSGKSSQ SLRESSNQKNYLAWYQQKQGQKPKLQKEGE STREEGDR GLNAAK,SGQR VR 0 11 VXMRLETRVITLVNECNHGRLRSGQRVRGLN AAKVGXKTPGSVQVTVTXSLSKVXVLVVSPP TPRPSNKDESHMCTRRVQATTLVRPRLCPTT TWLTXV GPRTGR,VTW GKFSR 0 8 KESGEGRTGKGPRTGRVEQGIEEXRLGLSRV HSKSLGGPTWSGGGLCSSKWRVTWGKFSR CYFFXSGFRYHSLHTSHNXXLTLRARLGETK GAXPXMLLY KNNSYLKSSDP K,YVVSTV 0 8 SSEEGGNTVRVGVDLGXPRTVTLVPPPKTXH NDXNPTQNLWGQHLVLFSRRSHKCLYXIGRA HVXTPVTRFHLVCRLLLEKKKNNSYLKSSDPK YVVSTV GPRTGR,VTW GKFSR 0 8 TNXRGEGRAGKGPRTGRGRWVTEDCSVGR RRGLPKSLGGPTWAGGGLSSSKWRVTWGK FSRCYFFXSGFRYHSLHTSHNXXLTLRARLG ETKGEXVYPGW CTGSGAGTE,Y TGVPERCTGS GAGTE 0 8 SRASRPYGEIXNASHQHGHQNGVTDSGVDG DIVMTQSHKFMSTSVGDRVSITCKASQDVST AVAWYQQKPGQSPKLLMYSASYRYTGVPER CTGSGAGTE *“X” refers to where an amino acid cannot be identified ** Wash control Table 6. Representative polypeptides identified as potential binders of MAP5 (INSR) ecto domain from affinity purified mouse IgG variable domains Peptides (query samples) Observation frequency of control** Observatio n frequency of query Candidate sequence* PSGTRR 0 1152 PSGTRRSGHSYTHEYIGPXLSLVSSEDLTTICPVG WPCDPRGXRSVRHLDVVXSETGSRHTSDFWNV XXXHLSQCRXASASHDPGSTSTLPDYDVVDVGR TIVVFGPVRGFR 0 222 SQSQWTFRSVLHPXLHRTIVVFGPVRGFRDXMS RRMAMSPQGLAKCPSPRPCLKXEWXSLHVRLLN RLIKTVVILSIGMCKPPPWFDLYFARLRRGXRRX GGGGEP 62 198 DGCSWCSISPFYFQLCPRAEREWIAVILLTEILCQ VFRLHIADGESEICPRSTACEAIRDSTVPVGCRVN QCFRRLPWFLVITGYISTHILTGRGGGGEP GLAGTS 34 178 RWMQLVQHQPVXFPAWCLHRTSDSYNIADSNKL PGLQPSHCXWXEXNLSQIHCLXSDQGPQIPXWM PSKSAVXGTARVFPGTGPRSSFICSLXKGLAGTS EEGGGG 1 163 YSPCGNLRGTEEAGPGFDSQFLTFSQHXTQTPH HELRAQLDFPCPYFKRCPVXSAAGGVWGRISEA WRVPETLRCSLWSHVQELGHVRGAPGPREEGG GG GGSGIK 11 148 XPWPFFVXLSNSXIIYSVVXCXLLGILPSCFFNFVH FXNFQDGVGTFIXXTSFFYFGGXXHVTYRARKCS INVFLVSSIIKGGCFPNYVCKGGSGIKRR CAGGGK 10 133 CAGGGKIYGGTYENSSNXLGTKIQRQNGFSSAD FQLPANQCFSHNVQRTNCSLPVSSNPVCISRGE GHNDLQGQLKCKLHALVPAEARLLPQTLDLCHIQ AASGTR 1 116 AASGTRRPGRVHIXCTXPWSSSVYGGGLIPKXCV GWTEDLRVERSHRHPDPXEXESVSRRTSHFYDG EXXRSSNDQGVPASHLRGSTFSLPDYDVVDVGR GSDLTC 0 112 RWMQLVQHQPVXFPAWCLHRTSEENVYLASNNK LPNPQPPLCXFSVXNLSLIHCQXTCQGLQSPVXT PDRLGALETGLASVTTNSNMSFHHYLRGSDLTC GGPINR,GGPIN RSSVV 0 62 RDGHAPIRGTNIEKNRSGLXIMAWISLIISFLAPHS GLISXALLTQEFALTTSLAQIVTLTCRPSTGAFTTR HYANWVQEQPDYLFTVLKGGPINRSSVV RVLSQSLGGST RSGGGLCSSK,V TWGKFSR 0 46 SSHGSVXARMCRRTAWIEMILDDVCVGLRRVLSQ SLGGSTRSGGGLCSSKWRVTWGKFSRCYFFXS GFRYHSLQTSHNXXLTLRARLGETKGASRLVREG VLSKSLGRPTW CGGGLSSSKWR ,VTWGKFSR 0 46 NERGEVRSGQGPGTGWVEXVIEDVSVGRQRVL SKSLGRPTWCGGGLSSSKWRVTWGKFSRCYFF XSGSRYHSLHTSHNXXLTLRARLGETKGASSHCI R GLSKSRGGPTW AGGGLSSSK,VT WGKFSR 0 46 TNERGEGRAGKGPRTDGVEEVIDDGSXGLRRGL SKSRGGPTWAGGGLSSSKWRVTWGKFSRCYFF XSGFRYHSLHTSHNXXLTLRARLGETKGAYCYRN W GGSGDK,GGSG DKK 0 33 WPFFVXLSNSXIIYSVVXCXLLGILPSCFFNFVHFX NFQDGVGTFIXXTSFFYFGGXXHVRPVYRARKCS INVFLVSSIIKGGCFPNYVCKGGSGDKKKE GCLSASARAPN DPF,HRREGPSG NAGQD 0 32 GKTFGKDXLSAETVTRVPWPQXANQAHCDNND XTPDNKXSLARAPGLSQMEASEGQRKHRREGPS GNAGQDPYKSLAMKSGSXGCLSASARAPNDPFX LP CCCRGSQIGD,G SQIGD 0 30 SSMTYXTQAVRGPFPDLTXPLXLGLSPCYLICGR DFLFKXXRSPLRFSCDMDVYWFHSRLLCREIMTR STPMMPMKLQTPRPWCQWQRSRCCCRGSQIG D CCCRGSQIGD,G SQIGD 0 30 DFLQXVCPRFLRCPTLYPFTXTCVCCFCCXXRW WWWSQSEFLYYPIXSYGFYSTXTRQTFDRPGLS DSVVSSPPVSSFXRRRLXTSXLFRCCCRGSQIGD YGSGRK,YGSG RKR 0 27 KYGSGRKRQAXALTASLHTRRPELKASPARPLHS SIEVGNIDRSSNXVESTLSCPGPRPGYPKPNPKS VTQYKSRHLGSSDRIQSGSGEGESGTKVEGDS DQSYQST,KDQS YQST 0 26 ESGSRLSLGHLTELVGVSVGRWWDFSQTLRRLX EWCLGPSSQCRNLPILHSNRPQSPQMSGCLVSC RLCWRIFLQSMLPCPXTSDLTYFHLRKDQSYQST *“X” refers to where an amino acid cannot be identified ** Wash control Example 7 - Aptamers generated from natural proteins
[00227] The proteins from tissue cells, body fluids may contain polypeptides that have binding activity against ligands, receptors or other protein targets.
[00228] The intact polypeptides may be extracted from tissue cells, and / or body fluids.
[00229] Separately, the polypeptides or proteins from tissue cells, and / or body fluids may be digested with trypsin or chymotrypsin or other proteolytic digestion to yield peptides.
[00230] The intact polypeptides, or digested peptides, with binding activity against ligands receptors or protein that are diagnostic or therapeutic targets were affinity purified against the ligands (eg IgG), receptors, protein targets immobilized on an affinity (or bispecific) chromatography resin (FIG. 6).
[00231] The polypeptides of serum or cells that bind targets may be identified from digested peptides matched to the predicted tryptic, chymotryptic or proteolytic peptides from the organism that was the source of the natural polypeptides. For undigested serum or cells: MS / MS spectra with greater than 1000 (>E3) arbitrary counts were searched 49 against a mouse library where cells were applied to IgG and a human library was used where normal human plasma was applied to IgG. Fully-digested tryptic peptide [RK]|[P] search conditions were used where the samples were digested with trypsin and chymotryptic search conditions [FYW]|[P] where samples were digested with chymotrypsin. Search conditions considered potential phosphorylation of Serine, threonine, or tyrosine (79.966331), potential carbamidomethylation of cysteine (57.021464). A single best fit was accepted per spectra. Samples were compared to controls by observation frequency to determine the most likely binders.
[00232] Representative polypeptides identified from the above experiments are listed in Tables 7-11. Table 7. Representative polypeptides identified from trypsin-digested cells GeneSymbol Peptides (Query samples) Observation frequency of control* Observation frequency of query samples mCG_13775 GGPASC 4 105 Myh9 AKLQEMESAVK,AKQ.TLENERGELANEVK,ALEEAMEQK,A LEEAMEQKAELER,ALEQQVEEMK,ALEQQVEEMKTQLEE LEDELQATEDAK,ANLQIDQINTDLNLER,ASIAALEAK,DEL ADEIANSSGK,DFSALESQLQDTQELLQEENR,DFSALESQL QDTQE LLQE E N RQK, D LG E E LEALKTE LE DTLDSTAAQQE L R,DVLLQVEDER,DVLLQVEDERR,ELEDATETADAMNR,E MEAELEDER,EMEAELEDERK,FDQLLAEEK,GALALEEK,IA QLEEELEEEQGNTELINDR,IAQLEEQLDNETK,IAQLEEQLD NETKER,IRELETQISELQEDLESER,KANLQIDQINTDLNLER ,KEEELQAALAR,KFDQLLAEEK,KLKDVLLQVEDER,KMEDG VGCLETAEEAKR,LKDVLLQVEDERR,LKSMEAEMIQLQEEL AAAER,LQQELDDLLVDLDHQR,LQVELDSVTGLLSQSDSK, LRLEVNLQAMK,NAEQFKDQADK,NAEQFKDQADKASTR, NSFREQLEEEEEAK,NSFREQLEEEEEAKR,QAQQERDELAD EIANSSGK,QLEEAEEEAQR,QLEEAEEEAQRANASR,QTLE N E RG E LAN E VK, RG D LP FVVTR,S M EAE MIQLQE E LAAAE R ,TEMEDLMSSKDDVGK,THEAQIQEMR,TQLEELEDELQAT EDAK,TRLQQELDDLLVDLDHQR,VEAQLQELQVK,YKASIA ALEAK 0 91 A330017A19 Rik GLGASGSPWAVTGAGR,RGGGGG 0 84 Egflam KSGGTSS,SGGTSS 1 74 Ccdcl24 AAADAK 8 70 Myh9 AKLQEMESAVK,AKQ.TLENERGELANEVK,ALEEAMEQK,A LEEAM EQKAELER,ALEQQVEEM K,AN LQIDQINTDLN LER, AQQAADKYLYVDKNFINNPLAQADWAAKK,ASIAALEAK, DELADEIANSSGK,DFSALESQLQDTQELLQEENR,DFSALES QLQDTQE LLQE E N RQK, D LG E E LEALKTE LE DTLDSTAAQQ ELR,DVLLQVEDER,DVLLQVEDERR,ELEDATETADAMNR, EMEAELEDER,EMEAELEDERK,FDQLLAEEK, GALALEEK, G ELANEVK,IAQLEEELEEEQGNTELINDR,IAQLEEQLDNETK, IAQLEEQLDNETKER,IRELETQISELQEDLESER,KANLQIDQ INTDLNLER,KEEELQAALAR,KFDQLLAEEK,KKVEAQLQEL QVK,KLKDVLLQVEDER,KQELEEICHDLEAR,LKDVLLQVED ERR,LKNKHEAMITDLEER,LQQELDDLLVDLDHQR,LQVEL DSVTGLLSQSDSK,LRLEVNLQAMK,NAEQFKDQADK,NAE QFKDQADKASTR,NSFREQLEEEEEAKR,QAQQERDELADE IANSSGK,QIATLHAQVTDMKK,QLEEAEEEAQR,QLEEAEE EAQRANASR,Q.TLENERGELANEVK,RGDLPFVVTR,SMEA EMIQLQEELAAAER,TEMEDLMSSK,TEMEDLMSSKDDVG K,THEAQIQEMR,TQLEELEDELQATEDAK,TRLQQELDDLL VDLDHQR,VAEFTTNLMEEEEK,VEAQLQELQVK,YKASIAA LEAK 0 65 Tcfl9 GEAGGG 0 57 Tnfrsfl8 GGAVVS 0 53 Vim DNLAEDIMR,EEAESTLQSFR,EKLQEEMLQR,EMEENFALE AANYQDTIGR,FADLSEAANR,FADLSEAANRNNDALR,ILL AELEQLK,ILLAELEQLKGQGK,ISLPLPTFSSLNLR,KLLEGEES R,KVESLQEEIAFLK,LGDLYEEEMR,LQDEIQNMK,LQEEML QR,MALDIEIATYR,NLQEAEEWYK,QAKQESNEYRR,QDV DNASLAR,QVDQLTNDK,QVQSLTCEVDALK,QVQSLTCEV DALKGTNESLER,SLYSSSPGGAYVTR,TNEKVELQELNDR,V ELQELNDR 0 52 Vdac2 FGIAAK,MCIPPPYADLGKAAR 0 48 Myh9 AKLQEMESAVK,ANLQIDQINTDLNLER,ASIAALEAK,DELA DEIANSSGK,DVLLQVEDER,DVLLQVEDERR,ELEDATETA DAMNR,EMEAELEDER,EMEAELEDERK,GALALEEK,IAQL EEELEEEQGNTELINDR,IAQLEEQLDNETK,IAQLEEQLDNE TKER,KANLQIDQINTDLNLER,KLKDVLLQVEDER,LKDVLL QVEDERR,LKSMEAEMIQLQEELAAAER,NAEQFKDQADK, NAEQFKDQADKASTR,QAQQERDELADEIANSSGK,QLEE AEEEAQR,QLEEAEEEAQRANASR,RGDLPFVVTR,SMEAE MIQLQEELAAAER,YKASIAALEAK 0 48 mCG_15507 DVAGGS 0 45 Bcng-1 GGAAGK,KMYFIQHGVAGVITKSSK 0 44 Myh9 AKLQEMESAVK,ALEEAMEQK,ALEEAMEQKAELER,ALEQ QVEEMK,ANLQIDQI NTDLN LER,ASIAALEAK,DELADEIAN SSGK,DVLLQVEDER,DVLLQVEDERR,ELEDATETADAMN R,EMEAELEDER,EMEAELEDERK,FDQLLAEEK,GALALEEK 0 43 ,IAQLEEELEEEQGNTELINDR,IAQLEEQLDNETK,IAQLEEQ LDNETKER,KANLQIDQINTDLNLER,KFDQLLAEEK,KLKDV LLQVEDER,LKDVLLQVEDERR,LQQELDDLLVDLDHQR,LR LEVNLQAMK,NAEQFKDQADK,NAEQFKDQADKASTR,Q. AQQERDELADEIANSSGK,QLEEAEEEAQR,QLEEAEEEAQ. RANASR,RGDLPFVVTR,SMEAEMIQLQEELAAAER,TEME DLMSSK,TEMEDLMSSKDDVGK,TQLEELEDELQATEDAK, TRLQQELDDLLVDLDHQR,YKASIAALEAK Hspa5 ALSSQHQAR,DNHLLGTFDLTGIPPAPR,ELEEIVQPIISK,GV PQIEVTFEIDVNGILR,IDTRNELESYAYSLK,IEIESFFEGEDFS ETLTR,IEWLESHQDADIEDFK,IINEPTAAAIAYGLDK,ITITN DQN R, ITITN DQN RLTP E E1E R, KSQIFSTAS D N QPTVTIK, LS SEDK,LYGSGGPPPTGEEDTSEKDEL,NKITITNDQNR,NQLT SNPENTVFDAK,RALSSQHQAR,SQIFSTASDNQPTVTIK,T WNDPSVQQDIK,VEIIANDQGNR,VLEDSDLKK,VTHAVVT VPAYFNDAQR,VYEGERPLTK 0 43 Fbxl20 GCGGLK,GCLGVGDNALSTCTSLSKFCSKLR 0 43 mCG_21897 GCGGLK 0 42 Hspa5 ALSSQHQAR,DNHLLGTFDLTGIPPAPR,ELEEIVQPIISK,GV PQIEVTFEIDVNGILR,IDTRNELESYAYSLK,IEWLESHQDAD IEDFKJINEPTAAAIAYGLDKJTITNDQNRJTITNDQNRLTP E E1E R, KSQI FSTAS D N QPTVTI K, LSS E D K, LYG SG G P P PTG E EDTSEKDEL,NKITITNDQNR,NQLTSNPENTVFDAK,RALSS QH QAR,SQI FSTAS D NQPTVTI K,TW N D PSVQQD1K, VE11A ndqgnr,vledsdlkk,vthavvtvpayfndaqr,vyeger PLTK 0 42 Fbxl20 GCGGLK 0 42 Hist2h4 DAVTYTEHAK,DNIQGITKPAIR,ISGLIYEETR,ISGLIYEETRG VLK,KTVTAMDVVYALK,KTVTAMDVVYALKR,RISGLIYEET R,TLYGFG,TVTAMDVVYALK,TVTAMDVVYALKR,VFLENV IR 0 41 ‘Control is a conditioned blank control Table 8. Representative polypeptides identified from chymotrypsin-digested cells GeneSymbol Peptides (Query samples) Observation frequency of control* Observation frequency of query samples B3gnt7 AGGGGF 0 30 Rpa2 GGAGGY,SSSTYGGAGGY 2 25 Foxo6 GGGGGF,GSCKASAYGGGGGF 0 22 BC107364 STCCGS 0 16 Plcb2 IS F E FSAQKN RSYVVSS F, LGG LGS,TE LKAY 0 16 Svepl ECICEKGYYGKGLQHECTACPSGTY,EVSGIYGY,GSKVVY,G TMVSY,LCEAGKDCCDRMASCKCGTHTGQF,QCTSGY,QG NIRELNDMASTPKEEHCYLLHSFEEF,SAEDFHAGSTVTY,SC EPGY,SCKCPPGF,SGTVPRCEAISCSKPNPLW,TSTGGAW 0 16 Atp5sl VAGAVF 0 14 Eln GAGAGVPGFGAGAGVP,GAGGAGALGGLVPG,GFGAGA GVPGFGAGAGVPGF,GGGGALGPGGKPPKPGAGLLGT,GL GGAGGLGAGGLG,LGGGGGAL,MEGRRSRGGLREGVPGA VPGGLPGG,PGAGIGGLGGGGGALGPGGKPPKPG,PGAVP GGLPGGVPGGVYYPGAGIGGLGGGGGALG,PGGVYYPGA GIGGLGGGGGALG 0 14 mCG_12393 EVSGIYGY,GSKVVY,GTMVSY,QCTSGY,SAEDFHAGSTVTY ,SCEPGY,SCKCPPGF,SGTVPRCEAISCSKPNPLW,TSTGGA W 0 12 Lor CGGSSGGGSSGGCGGGSGGGKYSGGGGGSSCGGGY,GG GSSCGGGGGSGGGVKYSGGGGGSSCGGGY,SGGGGGSC GGGSSGGGGGY,SGGGGGSSCGGGYSGGGGGSSCGGGS Y,SGGGGSSCGGGGGY,SGGGGSSCGGGY,SGGGGSSGGS SCGGGY,SSGGGGSSGGCGGGY,SSQQTSQTSCAPQQSYG GGSSGGGGSCGGGSSGGGGGGGCY 0 12 Rttn LRFSVFPW 0 12 Eln AQPGGVPGAVPGGLPGGVPG,GAGAGVPGFGAGAGVP, GAGGAGALGGLVPG,GFGAGAGVPGFGAGAGVPGF,GLG GAGGLGAGGLG,PGAGIGGLGGGGGALGPGGKPPKPG,P GAVPGGLPGGVPGGVYYPGAGIGGLGGGGGALG,PGGVY YPGAGIGGLGGGGGALG 0 12 Klf5 VAGGAW 1 12 Rttn LRFSVFPW 0 12 Krt76 GGAGGF,GGPSGFGGAGGF,SCGSQGF 0 11 Tro CNTASISFGGAPSTSTSF,GASSSTSSDF,GDGLGSSTSF,GG AISTSF,GGALNSNAGFGGAISTSTNF,GSSLGTSTGF,GSSL GTSTGFGGSLGPSASF,GSSPYSGAGFGGTLSTSISF,NGGLG NSAGFNGGLNTNTDF,SGTPSTSAPF 0 11 Sema3f TSSGSVF,VASAVGVTHLSLHRCQAYGAACADCCLARDPY 0 10 mCG_127462 CALGDT 0 10 Sema3f TSSGSVF,VASAVGVTHLSLHRCQAYGAACADCCLARDPY 0 10 mCG_146485 CALGDT 0 10 ‘Control is a conditioned blank control Table 9. Representative polypeptides identified from trypsin-digested NHP (normal human plasma) GeneSymbol Peptides (Query samples) Observation frequency of control* Observation frequency of query samples ALB AACLLPK,ADDKETCFAEEGK,ADDKETCFAEEGKK,AEFAE VSK,AEFAEVSKLVTDLTK,AFKAWAVAR,ATKEQLKAVMD DFAAFVEK,AVMDDFAAFVEK,AWAVAR,CCTESLVNR,DD NPNLPR,ECCEKPLLEK,EQLKAVMDDFAAFVEK,FKDLGEE NFK,FQNALLVR,HPYFYAPELLFFAK,KLVAASQAALGL,KQ TALVELVK,KVPQVSTPTLVEVSR,LCTVATLR,LDELRDEGK, LDELRDEGKASSAK,LKCASLQK,LKECCEKPLLEK,LSQRFPK, LVAASQAALGL,LVNEVTEFAK,LVRPEVDVMCTAFHDNEE TFLKK,LVTDLTK,NECFLQHKDDNPNLPR,QIKKQ.TALVELV K,QNCELFEQLGEYK,QNCELFEQLGEYKFQNALLVR,QRLK CASLQK,Q.TALVELVK,RHPDYSVVLLLR,RHPYFYAPELLFFA K,RMPCAEDYLSVVLNQLCVLHEK,RPCFSALEVDETYVPK,S HCIAEVENDEMPADLPSLAADFVESK,SLHTLFGDK,SLHTL FGDKLCTVATLR,TCVADESAENCDK,TPVSDRVTK,TYETTL ek,vfdefkplveepqnlik,vhtecchgdllecaddradla K,VPQVSTPTLVEVSR,YICENQDSISSK,YKAAFTECCQAAD K,YLYEIAR 0 4872 FLG DKQSGDGSR,DKQSGDGSRHSGSR,SPRETGGKRHESSSEK ,SRSSDGKSSSQVNR,TGPSTGGR 43 3052 GTF3C2 AYFTAPR,GSTSGK 250 875 ALB DDNPNLPR,FKDLGEENFK,HPYFYAPELLFFAK,LCTVATLR, LVNEVTEFAK,LVRPEVDVMCTAFHDNEETFLKK,NECFLQ HKDDNPNLPR,RHPYFYAPELLFFAK,SLHTLFGDK,SLHTLF GDKLCTVATLR,TCVADESAENCDK,YLYEIAR 0 667 YTHDC2 GGGDIR,GKGANR 116 582 DNAH3 ETVPFIQK,QSGGSGK,VPAGAK,VQPHLK 95 513 ZFYVE26 ATASGK,GGGPPR,SERGSLGVPK,SPSTLC,YQPATRHPSLR 50 463 CHPT1 SHQNNMD 0 394 CPAMD8 CLTCVR,ELAAAK,TEKRKR 27 311 LTB4R GGSLGQTARSGPAAL,SGPAAL 0 303 YBEY GLFGGS 3 270 PRDM1 GSPEMP 42 270 SUMF2 ASASGK,QGSCKQPGGDK,QPGGDK 0 255 TGFBRAP1 ASASGK 0 253 PIM3 CLCGGR 0 225 INPP5J AGAGAK 0 225 SLITRK5 GPALPK 0 221 ZNF835 CG ECG K,CGQCAK, IP RTISSPAATQASVPD DSSSRR 0 219 RHOT2 TRSCGR 0 215 ‘Control is a conditioned blank control Table 10. Representative polypeptides identified from chymotrypsin-digested NHP (normal human plasma) GeneSymbol Peptides (Query samples) Observation frequency of control* Observation frequency of query samples CA7 GGGLST 0 88 RARG AAGALG 0 64 RPL13A APAGGQ. 0 64 KRT3 GGAGGF,GGGSSGF,GGPGGF,GGSGGF,GVSGGGF,SAGGG SGSGF 0 42 CSMD2 GCN AGY, N ISLTVEYF,QCEPGYALQG HAH ISCM PGTVRRW, QFQAQLMLICDPGYY,SDISVSAAGFHLEY,SSSIVY,TLIYSCQE G F,TVSTVCTAV,YP N N LNCTW 0 26 EMC10 SSTGGA 0 23 B3GNT3 CGGGGF 1 23 MC4R SDGGCY,YALQYHNIMTVKRVGISISCIW 0 21 TECPR1 AIDFPATYTKDKKWNSCVRRRKW,ASDFPASY,GIGGGW,TLS TAGQYW,VDVRLALEQFTGHDGVRDSILFIYY,VKTGALQW 0 21 LTK AAAGGF,AGGGGGGGGATYVF 0 21 MC4R SDGGCY,YALQYHNIMTVKRVGISISCIW 0 21 SYP GPQDSY,GPQGGY,GQPAGSGGSGY,LATAVF 0 20 DMRT3 MNGYGSPY,SVGSAF 0 20 MC4R SDGGCY 0 20 PYCR1 GAAVHG 0 20 CUBN GDRGSF,HSSECY,IVVTSPDLLVTF,SLEEAIGNYY,TDGSRPY,T FLAFDLEHHINCSTDYLELY,TGPSGY,TSDNQMFVQFISDHSN EGQGF,YSSMSTAMVIF,YTDFLEIRDGGY 0 17 TRO DGGLGTSAGFGGGPGTSTGF,GCAHSTSTSF,GGAHGTSLCF GGAPSTSLCF,GGGLNTSAGF,GGGLVTSDGF,GGGPGTSTGF ,GGPPSTSACFSGATSPSFCDGPSTSTGF,GGSPCTSTGF,GGT LSTSVSF,GSRPNASFDRGLSTIIGFGSGSNTSTGF,SGGASSGF ,SGG LSTSSG F,SSG PSSIVG F,SSG PSSIVG FSGG PSTGVG FCSG PSTSGF,STSAGF 0 16 ETV4 GPKGGY 0 16 DMPK APNPGF,VADFLQW 0 16 ETV4 GPKGGY 0 16 ‘Control is a conditioned blank control Table 11. Representative polypeptides identified from undigested serum vs digested serum Observation frequency Accessi on Descrip tion Blank_No Serumlg G Blank_NoS erumNoig G Digested Serumlg G Digested Se rumNoigG IgGInta ctSeru m Intactser umNoig G P02647. 1 Apolipop rotein A-l APOA1 0 5 46 0 74 27 F6KPG5 .1 Albumin (Fragme nt) 1 29 36 0 386 35 A8K9P0 .1 cDNA FLJ7841 3, highly similar to Homo sapiens albumin, mRNA 1 25 34 0 305 32 B4DPR2 .1 cDNA FLJ5083 0, highly similar to Serum albumin 1 26 33 0 329 30 B4DPP6 .1 cDNA FLJ5437 1, highly similar to Serum albumin 1 27 31 0 380 33 B7WNR 0.1 Serum albumin ALB 1 25 31 0 319 29 H0YA55. 1 Serum albumin 1 22 31 0 308 27 (Fragme nt) ALB D6RHD 5.1 Serum albumin ALB 1 22 29 0 297 27 C9JKR2 .1 Albumin, isoform CRA k ALB 1 22 25 0 231 25 P02671. 1 fibrinoge n alpha chain isoform alpha prepropr otein FGA 0 0 22 0 24 20 P01009. 3 Alpha-1 -antitryps in SERPIN A1 0 3 19 1 37 9 Q8IUK7. 1 ALB protein 0 19 19 0 211 21 D3DP16 .1 Fibrinog en gamma chain, isoform CRAa FGG 0 2 18 0 16 15 Q7Z664. 1 Putative unchara cterized protein DKFZp7 79N092 6 (Fragme nt) DKFZp7 79N092 6 0 2 18 0 16 15 P02675. 2 Fibrinog en beta chain FGB 0 0 16 0 17 13 Q32Q65 .1 Fibrinog en beta chain FGB 0 0 16 0 17 13 Q9P173 .1 PRO227 5 0 2 15 1 15 7 B4E1D3 .1 cDNA FLJ5395 2, highly similar to Fibrinog 0 0 15 0 16 12 en beta chain D3DP13 .1 Fibrinog en beta chain, isoform CRAe FGB 0 0 15 0 15 12 D6REL8 .1 Fibrinog en beta chain FGB 0 0 15 0 14 10 NP_001 171670. 1 fibrinoge n beta chain isoform 2 prepropr otein FGB 0 0 15 0 15 12 H7C013 .1 Serum albumin (Fragme nt) ALB 0 9 13 0 100 10 P00738. 1 Haptogl obin HP 0 0 11 0 12 4 C5J0G2 .1 Serpina 1 (Fragme nt) 0 1 9 0 6 4 H0Y300. 4 Haptogl obin HP 0 0 9 0 13 4 J3QR68 .1 Haptogl obin (Fragme nt) HP 0 0 9 0 12 4 Q3KRA 7.1 FGA protein (Fragme nt) FGA 0 0 9 0 5 9 Q6NSD 8.1 FGA protein FGA 0 0 9 0 5 9 XP_005 255979. 1 PREDIC TED: haptoglo bin isoform X1 HP 0 0 9 0 12 4
[00233] While the present application has been described with reference to what are presently considered to be the preferred examples, it is to be understood that the application is not limited to the disclosed examples. To the contrary, the application is 58 intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.
[00234] All publications, patents and patent applications are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference in its entirety. Specifically, the sequences associated with each accession number provided herein including for example accession numbers and / or biomarker sequences (e.g. protein and / or nucleic acid) provided in the Tables or elsewhere, are incorporated by reference in its entirely.
[00235] The scope of the claims should not be limited by the embodiments and examples, but should be given the broadest interpretation consistent with the description as a whole.
Claims
1. A method of identifying a peptide aptamer, from a library comprising a plurality of peptide aptamers, that binds a target, comprising:providing the peptide aptamer library comprising the plurality of peptide aptamers, the plurality of peptide aptamers being derived or adapted from one or more natural peptides;contacting the peptide aptamer library with the target;removing unbound peptide aptamers after contacting the peptide aptamer library with the target; andanalyzing the bound peptide aptamers, comprising:generating a MS / MS query spectrum of the bound peptide aptamers;receiving one or more parameters of the query spectrum;providing one or more candidate spectra;generating a plurality of samples of the query spectrum;selecting at least one query sample from the plurality of samples for comparison with the one or more candidate spectra;determining a likelihood indicator for each of the one or more candidate spectra based on a comparison with the at least one query sample;applying a signal to noise filter to the one or more candidate spectra based on the likelihood indicators for the candidate spectra; andselecting at least one candidate spectrum as a proposed spectrum; anddetermining a peptide sequence of the proposed spectrum;thereby identifying the peptide aptamer of the plurality of peptide aptamers that binds the target.
2. A method of identifying a peptide aptamer, from a library comprising a plurality of peptide aptamers, that binds a target, comprising:providing the peptide aptamer library comprising the plurality of peptide aptamers, the plurality of peptide aptamers being derived or adapted from one or more natural peptides;contacting the peptide aptamer library with the target;removing unbound peptide aptamers after contacting the peptide aptamer library with the target; andanalyzing the bound peptide aptamers;thereby identifying the peptide aptamer of the plurality of peptide aptamers that binds the target.
3. The method of claim 2, wherein analyzing the bound peptide aptamers comprises:generating a MS / MS query spectrum of the bound peptide aptamers;receiving one or more parameters of the query spectrum;providing one or more candidate spectra; generating a plurality of samples of the query spectrum;selecting at least one query sample from the plurality of samples for comparison with the one or more candidate spectra;determining a likelihood indicator for each of the one or more candidate spectra based on a comparison with the at least one query sample;applying a signal to noise filter to the one or more candidate spectra based on the likelihood indicators for the candidate spectra;selecting at least one candidate spectrum as a proposed spectrum; and determining a peptide sequence of the proposed spectrum.
4. The method of any one of claims 1-3, wherein analyzing the bound peptide aptamers does not involve determining an in-frame amino acid sequence of the one or more candidate spectra.
5. The method of claim any one of claims 1-4, wherein determining the likelihood indicator comprises determining if more than one query sample fit to a candidate spectrum.
6. The method of any one of claims 1 -5, wherein applying the signal to noise filter comprises (i) determining an observation frequency of at least two query samples that fit to the candidate spectrum; (ii) determine an observation frequency of at least one control; (iii) if the observation frequency of the at least two samples that fit to the candidate spectrum is higher than the observation frequency of the at least one control, then the candidate spectrum is selected as a proposed spectrum.
7. The method of any one of claims 1 -6, wherein selecting at least one sample from the plurality of query samples for comparison with the one or more candidate spectra comprises, for each sample: counting a number of matching sequential b and y spectra lines of that sample; and if the number of matching sequential b and y spectra lines of that sample is less than a pre-determined minimum number of spectra lines, excluding that sample from comparison with the one or more candidate spectra.
8. The method of any one of claims 1 -7, wherein selecting at least one sample from the plurality of query samples for comparison with the one or more candidate peptide sequences comprises, for each sample: determining a spectra intensity of that sample; determine a number of matching b and y spectra lines; for the matching b and y spectra lines, determining a corresponding sum of the spectra intensity; determining a total sum of the spectra intensity; determining a ratio of the sum of the spectra intensity for the matching b and y spectra lines and the total sum of the spectra intensity; and if the ratio is less than a pre-determined threshold, excluding that sample from comparison with the one or more candidate spectra.
9. The method of any one or claims 3-8, wherein selecting at least one query sample from the plurality of query samples for comparison with the one or more candidate spectra comprises, for each sample: determining a precursor mass forthat sample; and if the precursor mass is substantially equal to a mass of a candidate spectra, selecting that sample for comparison with the one or more candidate spectra.
10. The method of claim 9, wherein the precursor mass comprises at least one of a precursor mass at a charge state of 1,2, or 3.
11. The method of claim 9, wherein the precursor mass comprises a mass shift from one or more post-translational modifications.
12. The method of any one or claims 1-11, wherein selecting at least one query sample from the plurality of query samples for comparison with the one or more candidate spectra comprises, for each sample: determining a mass of a theoretical ion of the sample; and if the mass of the theoretical ion is substantially equal to a mass of a base peak, selecting that sample for comparison with the one or more candidate spectra.
13. The method of any one or claims 1-12, wherein determining a likelihood indicator for each of the one or more candidate spectra based on a comparison with the at least one query sample comprises using one or more of linear regression, multiple linear regression, nested regression or linear model.
14. The method of any one or claims 1-13, wherein determining a likelihood indicator for each of the one or more candidate spectra based on a comparison with the at least one query sample comprises for each query sample, using a chi square test to compare theoretical ions of that sample with corresponding theoretical ions of the candidate spectra.
15. The method of any one or claims 1-14, wherein determining a likelihood indicator for each of the one or more candidate spectra based on a comparison with the at least one query sample comprises, for each query sample, determining a cross correlation score relative to the candidate spectra.
16. The method of any one or claims 1-15, wherein applying a signal to noise filter to the one or more candidate spectra based on the likelihood indicators for the candidate spectra comprises: generating random source noise spectra at high frequencies; determining a difference between the random source noise spectra at high frequencies to the one or more candidate spectra; and if the difference exceeds a pre-determined threshold difference, excluding the query sample fromselection as a proposed spectrum for the query spectrum; otherwise including that query sample for selection as a proposed spectrum.
17. The method of any one of claims 1-16, wherein generating random source noise spectra at high frequencies comprises using a Monte Carlo random simulation.
18. The method of any one of claims 1-17, wherein selecting at least one candidate spectrum as a proposed spectrum for the query spectrum comprises selecting a best fit per spectrum at the level of one of a peptide, an accession, or a gene symbol.
19. The method of any one of claims 1-18, wherein generating one or more candidate spectra based on the one or more parameters of the peptide query comprises one or more of: storing a plurality of spectra in a computer-readable medium; selecting, from the computer-readable medium, at least one of the plurality of stored spectra to use as at least one candidate spectrum, based on the one or more parameters.
20. The method of claim 19, wherein the plurality of stored spectra comprise one or more of: a spectrum predicted from a naturally occurring peptide sequence and a spectrum predicted from a synthetic peptide sequence.
21. The method of any one of claims 1 -20, wherein randomly generating at least one candidate spectrum, based on the one or more parameters comprises: randomly generating a peptide sequence; determining whether the randomly generated peptide sequence satisfies the one or more parameters; and if the randomly generated peptide sequence satisfies the one or more parameters, use the randomly generated peptide sequence as to predict a candidate spectrum, otherwise discard the randomly generated peptide sequence.
22. The method of any one of claims 1 -21, wherein each of the at least one randomly generated peptide sequence has a pre-determined length.
23. The method of any one of claims 1 -22, wherein generating a plurality of query samples of the query spectrum comprises one or more of generating experimental spectra or generating simulated spectrum.
24. The method of claim 23, wherein generating simulated spectra comprises using a Monte Carlo random simulation.
25. The method of any one of claims 1 -24, wherein the target comprises a receptor, a ligand, an enzyme, a protein, an antibody, a variable domain, or a drug.
26. The method of any one of claims 1 -24, wherein the target is immobilized on microbeads, nanobeads, a 2-dimensional surface, a 3-dimensional scaffold, and / or a 3-dimensional fiber.
27. The method of claim 26, wherein immobilization comprises use of a cleavable linker.
28. The method of claim 26, wherein the target is immobilized prior to contacting with the peptide aptamer library.
29. The method of claim 26, wherein the target is immobilized after contacting with the peptide aptamer library.
30. The method of any one of claims 1-29, further comprises releasing bound peptide aptamers from the target after removing unbound peptide aptamers.
31. The method of claim 30, wherein generating the MS / MS query spectrum comprises generating a MS / MS query spectrum of the released bound peptide aptamers.
32. The method of any one of claims 1 -31, wherein removing unbound peptide aptamers comprises one or more washing steps.
33. The method of claim 32, wherein the one or more washing steps comprise washing with an aqueous buffer, a weak salt solution, a weak mixture of an organic solvent with water, a weak acid or base close to neutral pH, and / or a mass spec compatible detergent.
34. The method of claim 33, wherein releasing bound peptide aptamers comprises eluting with a strong salt solution, a strong mixture of an organic solvent with water, and / or a strong acid or base far from neutral pH.
35. The method of claim 34, wherein immobilization of the target comprises use of a cleavable linker and releasing bound peptide aptamers comprises cleaving the cleavable linker.
36. The method of any one of claims 1 -35, further comprises digesting the bound peptide aptamers after removing unbound aptamers and before analyzing the bound peptide aptamers.
37. The method of claim 36, wherein generating the MS / MS query spectrum comprises generating a MS / MS query spectrum of the digested released bound peptide aptamers.
38. The method of any one of claims 1 -35, wherein deriving the plurality of peptide aptamers from the one or more natural peptides comprises digesting the one or more natural peptides.
39. The method of any one of claims 36-38, wherein the digesting comprises digesting with one or more proteases.
40. The method of claim 39, wherein the one or more proteases are selected from the group consisting of: Arg-C, Asp-N, Asp-N (N-terminal Glu), BNPS or NCS / urea, Chymotrypsin, Chymotrypsin (low specificity), Clostripain, Glu-C (AmAc buffer), Glu-C (Phos buffer), Lys-C, Lys-N, Lys-N (Cys modified), Pancreatic elastase, Pepsin A, Pepsin A (low specificity), Prolyl endopeptidase, Proteinase K, Thermolysin, Trypsin, Trypsin (Arg blocked), Trypsin (Cys modified), Trypsin (Lys blocked), Caspase-1, Caspase-2, Caspase-3, Caspase-4, Caspase-5, Caspase-6, Caspase-7, Caspase-8, Caspase-9, Caspase-10, Enterokinase, Factor Xa, Granzyme B, HRV3C protease, TEV protease, and thrombin.
41. The method of claim 38, wherein the digesting comprises chemical digest.
42. The method of claim 41, wherein the chemical digest comprises digest withCNBr, CNBr (methyl-Cys), CNBr (with acids), Formic acid, Glu-C (AmAc buffer), Glu-C (Phos buffer), Hydroxylamine, lodosobenzoic acid, Mild acid hydrolysis, NBS (long exposure), NBS (short exposure), NTCB, or any combination thereof.
43. The method of any one of claims 1 -42, wherein deriving the plurality of peptide aptamers from the one or more natural peptides comprises modification of the one or more natural peptides.
44. The method of claim 43, wherein the modification comprises one or more of: (i) reduction; (ii) alkylation; and (iii) oxidation.
45. The method of any one of claims 1 -44, wherein the plurality of peptide aptamers comprise undigested natural peptides.
46. The method of any one of claims 1 -44, wherein the plurality of peptide aptamers comprise untreated natural peptides.
47. The method of any one of claims 1-46, wherein identifying the peptide aptamers that bind the target comprises identifying by mass spectrometry, or top down mass spectrometry, or electrospray ionization, or MALDI ionization or chemical ionization or electron impact ionization or LC-ESI-MS / MS.
48. The method of claim 47, wherein identifying the peptide aptamers comprises de novo sequencing.
49. The method of claim 47, wherein identifying the peptide aptamers comprises fitting of observed MS / MS spectra to a predicted library.
50. The method of claim 49, wherein the fitting of observed MS / MS spectra to the predicted library comprises 64 bit computation.
51. The method of claim 50, wherein the 64 bit computation comprises use of cross correlation, XTANDEM, SEQUEST, regression, goodness of fit, cross correlation (X-corr), the count of fragment ion matches, heuristic algorithms, or combination thereof.
52. The method of any one of claims 1 -51, wherein the one or more natural peptides comprise an immunoglobulin superfamily member.
53. The method of any one of claims 1 -51, wherein the one or more natural peptides comprise an immunoglobulin, a B cell antigen receptor (BCR), a T cell receptor (TCR) or any combination thereof.
54. The method of claim 53, wherein the immunoglobulin comprises IgG, IgA, IgM, IgE, and / or IgD.
55. The method of claim 54, further comprising isolating Fab from the isolated immunoglobulin56. The method of any one of claims 52-55, wherein the one or more natural peptides are isolated from a host animal.
57. The method of claim 56, wherein the host animal has contacted an immunogen.
58. The method of claim 57, wherein an immune response has been induced in thehost animal.
59. The method of any one of claims 56-58, wherein the one or more candidate spectra comprise spectra predicted from a transcriptome of immune cells isolated from the host animal.
60. The method of claim 59, wherein the predicted spectra are generated by: (1) isolating immune cells from the host animal; (2) isolating RNAfrom the immune cells; (3) obtaining sequences of the RNA.
61. The method of claim 60, wherein obtaining sequences of the RNA comprises reverse transcription of the RNA to generate DNAand sequencing said DNA.
62. The method of claim 61, further comprises generating amino acid sequences from the DNA.
63. The method of claim 62, wherein generating amino acid sequences from the DNA comprises reading a plurality of reading frames of the DNA sequence.
64. The method of any one of claims 56-63, wherein the host animal is a mouse, a rabbit, a goat, a donkey, or a camelid.
65. The method of any one of claims 56-64, further comprises determining a variable domain repertoire of immune cells of the host animal.
66. The method of claim 65, wherein determining the variable domain repertoire of immune cells comprises DNA sequencing.
67. The method of claim 53, further comprising determining a sequence of a variable domain of the isolated immunoglobulin, BCR or TCR.
68. The method of claim 67, wherein determining the sequence of the variable domain comprises performing liquid chromatography electrospray ionization tandem mass spectrometry (LC-ESI-MS / MS) on the digested products.
69. The method of any one of claims 67-68, wherein determining the sequence of the variable domain comprises 64 bit computation.
70. The method of any one of claims 67-69, wherein determining the sequence of the variable domain comprises using XTANDEM, SEQUEST, regression, goodness of fit, cross correlation (X-corr), the count of fragment ion matches, and / or heuristic algorithms.
71. The method of any one of claims 67-69, wherein determining the sequence of the variable domain comprises analysis performed in SQL Server and R using classical statistics.
72. The method any one of claims 1 -51, wherein the one or more natural peptides originate from a biological sample.
73. The method of claim 72, wherein the biological sample comprises tissue cells or biofluids.
74. The method of claim 72 or 73, wherein the one or more candidate spectra comprise spectra predicted from genomic DNA sequence of the biological sample.
75. The method of any one of claims 1 -74, wherein the peptide aptamer library is generated from a biological sample.
76. The method of claim 75, wherein the biological sample comprises a tissue, a cell, and / or a biofluid.
77. The method of claim 76, wherein the biofluid is plasma.