Next generation sequencing for protein measurement
By using hybridization capture technology and next-generation sequencing technology to replace the aptamers in SOMAmer elution buffer, the scalability and cost issues of proteomics detection and quantification methods have been solved, achieving efficient and accurate protein abundance quantification.
Patent Information
- Application Number
- CN202280089749.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-12-30
- Filing Date
- 2022-12-29
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2042-12-29
AI Technical Summary
In existing technologies, proteomics detection and quantification methods suffer from limited scalability, high cost, and limited commercial availability of microarrays, making it difficult to effectively replace SOMAmer molecule capture and elution quantification.
Using hybridization capture technology, the SOMAmer elution molecules characterizing protein capture are replaced with "reporter" DNA molecules containing SOMAmer-specific recognition markers. Next-generation sequencing technology is used to sequence the molecules and form a three-molecule complex for quantification.
This method enables efficient and accurate quantification of protein abundance in biological samples, reduces detection costs, and improves the scalability and reliability of the method.
Smart Images

Figure CN118591637B_ABST
Abstract
Description
[0001] Cross-references
[0002] The following application and materials are incorporated in full into this document and used for all purposes: U.S. Provisional Application No. 63 / 294,964, filed on December 30, 2021. Technical Field
[0003] This disclosure relates to systems and methods for the quantitative measurement of proteins in biological samples. More specifically, the disclosed embodiments involve capturing a target protein with a specially designed aptamer, forming an eluent containing the aptamer containing the captured target protein, and then replacing the aptamer in the eluent with a “reporter” DNA molecule that is easier to sequence than the aptamer itself. Background Technology
[0004] Typically, various attempts to assess gene activity and / or decode biological processes (including disease processes or pharmacological processes) focus on genomics. However, proteomics can provide further information about the biological functions of cells and organisms. Proteomics involves detecting and quantifying expression at the protein level, rather than the gene level, to achieve qualitative and quantitative measurements of gene activity. Proteomics also includes the study of non-gene-coding events, such as post-translational protein modifications and protein-protein interactions.
[0005] Currently, obtaining vast amounts of genomic information is possible. DNA microarrays, as molecular arrays, have been put into practical use, and the price of direct DNA sequencing continues to decline significantly. Similarly, the demand for high-throughput proteomics is constantly increasing. In health monitoring, proteomics is preferable to genomics because the genome is static, only indicating medical possibilities, while the proteome dynamically changes with a patient's medical status and can even be considered to define that status. However, detecting and quantifying proteins is difficult, while detecting and quantifying nucleic acids is relatively easy, at least in part because proteins are more complex and variable than DNA in terms of biological function. This has spurred many attempts to represent protein concentration by measuring messenger RNA (mRNA) concentration. However, it has been shown that mRNA concentration does not correlate well with protein concentration. Proteomics appears to rely on the ability to directly detect proteins.
[0006] One method for detecting and quantifying the presence of specific proteins in biological samples is to use protein capture and sustained-release modified aptamers. Reagents. SOMAmer reagents are constructed with chemically modified nucleotides that greatly expand the physicochemical diversity of large random nucleic acid libraries used to select SOMAmer reagents therefrom. Assays using SOMAmer reagents measure native proteins in complex matrices by converting each individual protein concentration to a corresponding SOMAmer reagent concentration, which is subsequently quantified by standard DNA techniques such as microarray or qPCR.
[0007] SOMAmer reagents are DNA-based single-stranded protein affinity reagents that include chemically modified nucleotides that mimic amino acid side chains, thereby increasing the chemical diversity of standard aptamers, enhancing the specificity and affinity of protein-nucleic acid interactions. These modified nucleotides are incorporated into nucleic acid libraries used in an iterative selection and amplification process called Systematic Evolution of Ligands by Exponential Enrichment (SELEX) in which SOMAmer reagents are selected. SOMAmer reagents can be generated using SELEX-type processes to capture proteins that are resistant to selection with unmodified nucleic acids (ACTG traditional aptamers). SOMAmer reagents can be customized to select for desired specificity and slow-release properties, as well as mimic the assay conditions in which the reagents will be used.
[0008] In SOMAmer-based assays, the presence of proteins in a sample is converted to a specific SOMAmer-based DNA signal. The SOMAmer-protein binding step is followed by a series of separation and washing steps to convert the relative protein concentration to a measurable nucleic acid signal that is quantified using DNA detection techniques, such as hybridizing fluorescently labeled SOMAmers to custom DNA microarrays. After laser scanning the microarray, the readout in relative fluorescence units (RFU) is directly proportional to the amount of target protein in the original sample.
[0009] There are several drawbacks to protein quantitation detection with microarray hybridization, including limited scalability, fixed assay cost, and limited commercial sources of microarrays. Therefore, there is a need to develop alternative quantitation methods for SOMAmer molecules in the post-capture eluate. SUMMARY
[0010] The present disclosure provides systems, devices, and methods related to the detection and quantitation of proteins using parallel sequencing technology known as next-generation sequencing or NGS. More specifically, the present disclosure relates to a hybrid capture (HC) technology in which the SOMAmer elution molecules that characterize the protein capture are replaced with "reporter" DNA molecules that contain a SOMAmer-specific identification tag or "SOMA ID" that is then sequenced using NGS technology.
[0011] In some embodiments, the present disclosure relates to a system and method for quantifying abundance of a target protein in a biological sample, comprising capturing the target protein by exposing the biological sample to a plurality of aptamers configured to capture the target protein, each of the plurality of aptamers configured to bind to a particular protein; isolating the aptamer that captured one of the plurality of target proteins from an aptamer-containing eluate; forming a plurality of trimolecular complexes by exposing the aptamer in the eluate to a plurality of capture probes, each of the plurality of capture probes configured to hybridize to a particular aptamer, each trimolecular complex comprising: one of the plurality of aptamers from the eluate, a first capture probe comprising a portion that hybridizes to a first portion of the aptamer, and a second capture probe comprising a portion that hybridizes to a second portion of the aptamer, a DNA primer region, and an aptamer ID sequence corresponding to the aptamer; separating the trimolecular complex from the capture probes that did not bind to the aptamer; dissociating the capture probes in the trimolecular complex from the corresponding aptamer; amplifying the aptamer ID sequence in the dissociated capture probes; sequencing the aptamer ID sequence by next-generation sequencing; and determining the abundance of the target protein in the biological sample based on data obtained by sequencing the aptamer ID sequence.
[0012] In some embodiments, the present disclosure relates to a system and method for quantifying abundance of two or more target proteins in a biological sample, comprising capturing the target proteins by exposing the biological sample to a plurality of aptamers, each of the plurality of aptamers configured to capture a particular protein; forming an aptamer-containing eluate by isolating the aptamer that captured one of the plurality of target proteins in the biological sample; forming a plurality of trimolecular complexes, each trimolecular complex comprising: a particular aptamer present in the aptamer-containing eluate, a first probe that hybridizes to a corresponding first portion of the particular aptamer, and a second probe comprising a portion that hybridizes to a corresponding second portion of the particular aptamer, at least one DNA primer region, and an aptamer ID sequence corresponding to the particular aptamer; amplifying the aptamer ID sequence; sequencing the aptamer ID sequence; and quantifying the abundance of the target protein based on the sequenced aptamer ID sequence.
[0013] In some embodiments, the present disclosure relates to a system and method for detecting a target protein in a biological sample, comprising capturing the target protein by an aptamer by combining the biological sample with a plurality of aptamers, each of which is configured to bind to a specific protein; forming a tri-molecular complex, the tri-molecular complex comprising: the aptamer that captures the target protein, a first probe, the first probe comprising a portion that hybridizes to a corresponding first portion of the aptamer that captures the target protein, and, a second probe, the second probe comprising a portion that hybridizes to a corresponding second portion of the aptamer that captures the target protein, at least one DNA primer region, and an aptamer ID sequence corresponding to the aptamer that captures the target protein; amplifying the aptamer ID sequence; and sequencing the aptamer ID sequence to identify the aptamer ID sequence, thereby identifying the aptamer that captures the target protein and the target protein.
[0014] In some embodiments according to aspects of the present disclosure, the aptamer used to capture the target protein can be sequenced directly, e.g., using next generation sequencing technology, without the need to convert the aptamer into a simpler sequence using a tri-molecular complex.
[0015] In some embodiments according to aspects of the present disclosure, the aptamer used to capture the target protein can be a SOMAmer.
[0016] In some embodiments according to aspects of the present disclosure, the aptamer-containing eluate can be divided into groups before and / or after exposure to the hybridization probes. In some cases, some or all of the eluate groups can be diluted to a desired degree.
[0017] In some embodiments according to aspects of the present disclosure, a quantitative spike-in reporter can be added at a desired assay stage to correct for compensatory changes in the proportion of analyte counts.
[0018] Features, functions, and advantages can be implemented independently in embodiments of the present disclosure, or can be combined with other embodiments, further details of which can be found in the following description and drawings. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 is a schematic of a tri-molecular complex according to aspects of the present disclosure comprising one aptamer and two probes hybridized to the aptamer.
[0020] Figure 2A is a flow chart of steps of an exemplary method of preparing a hybridization probe needed to make a tri-molecular complex according to aspects of the present disclosure. Figure 1
[0021] Figure 2B is a flow chart of steps of an exemplary method of preparing a tri-molecular complex according to aspects of the present disclosure. Figure 1 A flowchart of the steps of an exemplary method for preparing a three-molecule complex.
[0022] Figure 3A This is based on the use of the three-molecule complex in various aspects according to this specification, such as... Figure 1 A flowchart of the steps of an exemplary method for performing next-generation sequencing content determination using a complex of [a specific compound].
[0023] Figure 3B This is a flowchart illustrating the steps of an exemplary method for performing next-generation sequencing content determination according to various aspects of this specification, the method including capturing a target protein and then forming a three-molecule complex, such as... Figure 1 Complexes.
[0024] Figure 4 This is a flowchart of the steps and byproducts for exemplary next-generation sequencing content determination involving four hybrid groups, based on various aspects of this specification.
[0025] Figure 5 This is a flowchart illustrating the steps and byproducts of exemplary next-generation sequencing assays involving one hybridization group and four PCR groups, based on various aspects of this specification.
[0026] Figure 6 This is a flowchart of the steps and byproducts for exemplary next-generation sequencing content determination involving four hybridization groups and four PCR groups, according to various aspects of this specification.
[0027] Figure 7 It is a histogram depicting the hypothetical, simplified results of two-analyte assays for two different samples.
[0028] Figure 8 This is based on all aspects of this instruction manual. Figure 7 The graph shows the content determination results, where quantitative spike (qSpike) control reporter factors were added to both samples at the same concentration.
[0029] Figure 9 This is a histogram of the raw results of the determination of SOMAmer content in three analytes in eight different samples with different analyte concentrations according to various aspects of this specification, with four different qSpike reporter factors added to each sample.
[0030] Figure 10 This is based on all aspects of this instruction manual. Figure 9 Histogram of normalized results for content determination.
[0031] Figure 11 This is a flowchart of some steps and byproducts of an exemplary next-generation sequencing assay involving four PCR groups, each with the addition of a qSpike control reporter factor, according to various aspects of this specification.
[0032] Figure 12 is a plot of relative fluorescence units (RFU) versus temperature showing experimentally obtained single-phase thermal melting data for a SOMAmer-probe duplex, overlaid with a two-state theoretical model curve fit, according to aspects of the present description.
[0033] Figure 13 is a plot of relative fluorescence units (RFU) versus temperature showing experimentally obtained two-phase thermal melting data for a SOMAmer-probe duplex, overlaid with a two-phase theoretical model curve fit, according to aspects of the present description. DETAILED DESCRIPTION
[0034] Various aspects and examples of hybrid capture, next-generation sequencing content determination systems, and related methods for protein detection and quantification are described below and illustrated in the related drawings. Unless otherwise indicated, a protein content determination according to the present description, and / or various components thereof, can include at least one of the structures, components, functions, and / or variations described, illustrated, and / or incorporated herein. Furthermore, unless specifically excluded, process steps, structures, components, functions, and / or variations described, illustrated, and / or incorporated herein in connection with the present description can be included in other similar devices and methods, including interchangeability between disclosed embodiments. The following description of examples is merely exemplary in nature and is not intended to limit the disclosure, its application, or uses. Furthermore, the advantages described in relation to the examples and embodiments implemented below are merely exemplary in nature and not all examples and embodiments have the same advantages or the same degree of advantages.
[0035] This detailed description includes the following sections: (1) Definitions; (2) Overview; (3) Examples, Components, and Alternatives; (4) Advantages, Features, and Benefits; and (5) Conclusion. The Examples, Components, and Alternatives section is further divided into multiple subsections, each of which has been labeled accordingly.
[0036] Definitions
[0037] The following definitions apply herein, unless otherwise indicated.
[0038] “Include,” “includes” and “including” are interchangeable with “comprise,” “comprises” and “comprising” and mean including but not necessarily limited to.
[0039] Terms such as “first,” “second,” and “third” are used to differentiate or identify members of a group, and are not intended to illustrate sequential or chronological order, unless otherwise indicated.
[0040] “AKA” means “also known as” and can be used to indicate an alternative name or corresponding term for one or more given elements.
[0041] Directional terms, such as“upper,”“lower,”“vertical,”“horizontal,” and the like, are understood to be in the context of the particular object being discussed. For example, an object can be oriented about defined X, Y, and Z axes. In these examples, the X-Y plane is defined as horizontal, upward is defined as the positive direction of the Z axis, and downward is defined as the negative direction of the Z axis.
[0042] In the context of methods,“providing” can include receiving, obtaining, purchasing, making, generating, processing, pre-processing, and / or the like, such that the provided object or material is in a state and configuration that is usable for further steps to be performed.
[0043] “NGS” means“next generation sequencing.”
[0044] “HC” means“hybrid capture.”
[0045] “SOMAmer” refers to“Slow Off Rate Modified Aptamer” reagents developed and manufactured by SomaLogic Operating Co, Inc. (“SomaLogic”) of Boulder, Colorado.
[0046] “SOMAmer ID sequence” or“SOMA ID” or“reporter” refers to a portion of a tri-molecular complex that includes a SOMAmer-specific DNA strand that can be sequenced using NGS technology.
[0047] “Quantitative spike” or“reporter spike” or“qSpike” refers to an amplifiable reporter used to normalize compensatory reads across samples, allowing for identification of true signal changes.
[0048] In this disclosure, one or more publications, patents, and / or patent applications can be referred to by reference number. However, these materials are incorporated only to the extent that no conflict exists between the statements and drawings set forth herein and such incorporated material. In the event of any such conflict, the present disclosure controls.
[0049] SUMMARY
[0050] In general, the present disclosure relates to methods of detecting and quantifying target molecules, such as proteins, in a biological sample. The disclosed methods can include capturing a target molecule by an aptamer, replacing the aptamer with an aptamer identification sequence, and then sequencing the aptamer identification sequence by a next generation sequencing technology. Alternatively, the disclosed methods can include capturing a target molecule by an aptamer and then directly sequencing the aptamer.
[0051] Examples, Components, and Alternatives
[0052] The following sections describe some aspects of protein detection and quantification using aptamers (such as SOMAmer reagents) and related systems and / or methods, in which SOMAmers are replaced by reporter DNA molecules containing SOMAmer-specific fragments that can be sequenced using next-generation sequencing technologies, by hybridization capture. The examples in these sections are intended to be illustrative and should not be construed as limiting the scope of the disclosure. Each section can include one or more different embodiments or examples, and / or contextual or related information, functionality, and / or structure.
[0053] A. Exemplary aptamers
[0054] This section describes slow-release modified aptamers (SOMAmers), which are illustrative examples of aptamers suitable for use in conjunction with the example systems and methods described herein.
[0055] A method known as "Systematic Evolution of Ligands by Exponential Enrichment," sometimes also referred to as the SELEX process, has clearly demonstrated that nucleic acids have a three-dimensional structural diversity similar to proteins. The SELEX process is a method by which nucleic acid molecules are evolved in vitro to achieve a certain desired activity. Here, SELEX is described for the production of nucleic acid molecules that bind with high specificity to a target molecule. The SELEX process provides a class of products known as nucleic acid ligands or aptamers, each with a unique sequence and with the property of binding specifically to a desired target compound or molecule. Each SELEX-identified nucleic acid capture reagent is a specific ligand for a given target compound or molecule. The SELEX process is based on the unique perspective that nucleic acids have sufficient ability to form a variety of two- and three-dimensional structures, and that they have sufficient chemical versatility within their monomers to serve as ligands (form specific binding pairs) for virtually any compound, whether monomeric or polymeric. Molecules of any size or composition can be targeted.
[0056] SELEX methods applied to high affinity binding include selection from a mixture of candidate oligonucleotides, and iterative binding, separation and amplification using the same general selection protocol, to achieve almost any desired standard of binding affinity and selectivity. The SELEX method, starting with a mixture of nucleic acids, which preferably includes a stretch of random sequence, includes the steps of contacting the mixture with a target under conditions favoring binding, separating unbound nucleic acids from nucleic acids that have specifically bound to the target molecule, dissociating the nucleic acid-target complex, amplifying the nucleic acids dissociated from the nucleic acid-target complex to produce a mixture of nucleic acids enriched for ligands, and then repeating the steps of binding, separation, dissociation and amplification until a desired number of cycles is achieved, to produce a nucleic acid ligand that is highly specific and of high affinity for the target molecule. In this way, aptamers can be discovered that are suitable for binding to almost any target protein.
[0057] More specifically, SOMAmers are protein binding aptamers discovered by modification of the SELEX process, which dissociation rates (t 1 / 2 ) are typically between 30 and 240 minutes, which is the average time required for a protein-aptamer complex to dissociate by half. In addition, SOMAmers contain modified nucleosides that provide different intrinsic functionalities. These functionalities can include labels for immobilization, labels for detection, means to facilitate or control separation, side chains that provide better affinity to proteins, etc. Modifications that improve affinity to proteins are typically chemical groups attached to the 5 position of the pyrimidine base. By functionalizing this 5 position with protein-like groups such as phenyl, 2-naphthyl, the chemical diversity of SOMAmers is extended, allowing high affinity binding to a wider range of target molecules. In addition, some polymerases are still able to transcribe DNA modified at these positions, allowing the amplification required by the SELEX process.
[0058] It should be noted that while binding aptamers, including SOMAmers, are typically discovered by the SELEX process, it is possible that other methods can be used to select them. For example, as computer modeling of molecular interactions improves, it can later be possible to directly calculate the ideal nucleic acid sequence of an aptamer and the associated chemical modifications of a SOMAmer, to generate a capture reagent that is specific for a given target molecule. Other chemical techniques for screening aptamers and SOMAmers are possible in addition to SELEX.
[0059] Assays for detecting and quantifying the amount of physiologically significant molecules in biological and other samples are important tools in scientific research and health care. Each SOMAmer is capable of binding to a target molecule in a sample in a highly specific manner and with very high affinity. After appropriate washing and separation steps, first to remove unbound proteins and then to remove unbound SOMAmers, the SOMAmers are eluted from the resulting SOMAmer-protein complex. The SOMAmer eluate is then contacted with a microarray containing complements of the SOMAmers, enabling the determination of the absence, presence, quantity and / or concentration of the target molecule in the sample.
[0060] B. Exemplary hybrid capture content assay methods
[0061] This section describes a targeted hybrid capture (HC) assay in which the SOMAmer signal from an assay eluate is replaced with "reporter" DNA molecules containing SOMAmer-specific identification sequences or "SOMA IDs" to enable sequencing.
[0062] Prior to the HC assay described in this section, a SOMAmer binding step has been performed, resulting in an eluate containing SOMAmer reagents that indicate the presence of the corresponding target protein in the sample. For example, but not by way of limitation, the following steps can be performed to obtain an eluate containing SOMAmers:
[0063] (1) Protein-specific SOMAmer reagents labeled with a 5' fluorescent group, a photocleavable linker, and biotin are immobilized on streptavidin (SA)-coated beads and incubated with one or more samples containing a complex mixture of proteins;
[0064] (2) SOMAmer-target protein complexes are formed on the beads;
[0065] (3) The beads are washed to remove unbound proteins and the bound proteins are labeled with biotin;
[0066] (4) The SOMAmer-protein complexes are released from the beads by photocleavage of the linker by UV light;
[0067] (5) Incubation in a buffer containing a polyanion competitor prevents rebinding of dissociated proteins, resulting in a dynamic increase in complexes with slow off rates that are specific for the target protein compared to interactions of the target protein with corresponding SOMAmers that have fast off rates;
[0068] (6) SOMAmer-protein complexes were recaptured on the second group of streptavidin-coated beads by biotin-labeled proteins, followed by an additional washing step to facilitate further removal of non-specifically bound SOMAmer reagents; and
[0069] (7) Release the SOMAmer reagent from the beads in the denaturing buffer to form an eluent containing SOMAmer suitable for quantitative analysis.
[0070] Now let's turn to the focus of this section: hybrid capture methods. Figure 1 The diagram schematically illustrates a three-molecule complex, collectively denoted as 100, that can be used for next-generation sequencing assays to identify target proteins. Complex 100 includes SOMAmer 102, which is one of the SOMAmers retained in the eluent after assaying, i.e., after exposure to the biological sample and (e.g.) the other steps described above. In other words, the presence of SOMAmer 102 in the eluent after assaying indicates the presence of the corresponding target protein (or other target molecule) in the sample.
[0071] Complex 100 further includes a first probe 104 and a second probe 106. The first probe 104 includes a... Figure 1 The left-hand complementary hybridization region H1 of SOMAmer 102. The second probe 106 includes... Figure 1 The right-side complementary hybridization region H2 of SOMAmer 102 also includes universal primer regions P1 and P2, as well as the unique SOMAmer recognition sequence I corresponding to SOMAmer 102. S Or “SOMA ID”, as described in more detail below. The positions of H1 and H2 can also be reversed, i.e., accompanied by the unique SOMAmer recognition sequence I in the universal primer regions P1 and P2 and H2. S With proper inversion, H1 hybridizes to the right side of SOMAmer 102, and H2 hybridizes to the left side of SOMAmer 102.
[0072] Hybrid regions H1 and H2 are configured to specifically bind different complementary portions of the corresponding SOMAmers and can be designed to have similar melting temperatures (T0). m To achieve simultaneous hybridization under a given set of content determination conditions. For example, the hybridization region of complex 100 can be designed and formed according to the following steps.
[0073] First, the target region to be used as a probe is determined on the SOMAmer. For truncated SOMAmers, the entire SOMAmer sequence can be used as the target region, including the 5 bases of the fixed region used for amplification from each end of the random region in SELEX. For full-length SOMAmers, the SOMAmer can be truncated in silico (i.e., computationally) to any desired length, such as a 50-mer, and then the hybridization complement is determined.
[0074] Next, the boundaries for dividing the target region into two parts are determined. In some examples, to have similar melting temperatures for the two hybridization regions, the melting temperature of the 25-mer (for example) duplex between the SOMAmer and H1 and H2 can be determined by calculation, and then the boundary between the two regions is adjusted stepwise until the melting temperature between the H1- and H2-SOMAmer duplexes reaches a maximum equilibrium. In other examples, different melting temperatures can be intentionally chosen, for example, a first melting temperature of the H1 probe (e.g., 45°C) and a second melting temperature of the H2 probe (e.g., 35°C).
[0075] Length restrictions can also be imposed on the hybridization regions. For example, the H1 and H2 minimum length can be set to be 18-mers. Similarly, for example, the H2 maximum length can be set to be 30-mers to ensure that H2 plus the remaining reporter portion of the second probe is still shorter than the maximum length required for subsequent synthesis, such as 100 bases long. Under these restrictions, the hybridization regions H1 and H2 can be generated by calculation.
[0076] When generating the universal primer regions P1 and P2 and the SOMAmer ID sequence I S Various factors can be considered. For example, the sequencing amplification design for a counting application must strike a good balance between the need for short and inexpensive reads and the need for sequences with sufficient length and information content as identifier sequences for the SOMAmer IDs for counting and barcode sequences for multiplexing, etc. For these reasons, the true area of the reporter region, which ultimately becomes the largest part of the sequencing template when scaled, can be limited in length when scaling the content. For example, the length of the primer regions P1 and P2 can be limited to 24-mers, the length of the SOMAmer ID sequence I S may be limited to 15-mers, the length of the edit distance is at least 5, and the length of the homopolymer is no more than 2-mers. The primer regions and SOMAmer ID sequence can also have other length restrictions and choices.
[0077] Figure 2A is to generate hybridization probes: H1 (104) and H2 (106) (for formingFigure 1 FIG. 2 is a flowchart of steps of an exemplary method 200 for producing a set of SOMAmer probes for a content assay. In step 202, a set of SOMAmer sequences for a content assay is provided.
[0078] In step 204, hybridization probe regions H1 and H2 are generated. These hybridization probe regions can be computationally determined, for example, under various lengths and / or other constraints, as described above. Also as described previously, in some cases, hybridization regions can be split from a single SOMAmer complement structure based on factors such as balancing the melting temperature of each region.
[0079] In step 206, SOMAmer ID (I S ) regions are generated that uniquely correspond to each SOMAmer. SOMAmer ID regions can be designed by various methods. For example, I S regions can be designed "by eye" (e.g., with a maximum edit distance), or they can be computationally generated along with the computational generation of hybridization regions H1 and H2. After a library of SOMAmer ID regions is generated, SOMAmer IDs can be assigned to SOMAmers at random or in any other suitable manner to generate a unique reporter for each SOMAmer.
[0080] In step 208, universal primer regions P1 and P2 are generated. Universal primers can be designed to be stable and to reduce the risk of bias downstream. For example, in some cases, primer lengths can be 24 or 25-mers with an estimated melting temperature of about 70°C. In some cases, primers can be terminated at the 3' end with a guanine (G) to preserve stability. In some cases, primers can be evaluated with an oligo analyzer to reduce the potential risk of dimer formation. In some cases, primers can be further customized to avoid non-specific interactions with functional oligos used in known sequencing technologies.
[0081] In step 210, first and second SOMAmer-specific probes (sometimes referred to as "capture probes") are generated. Each first probe includes a hybridization region H1, as well as one or more elements suitable for binding to a content assay bead, such as biotin for binding to a streptavidin-coated bead. First probes can also include other elements such as a photolyzable linker. Each second probe includes a hybridization region H2, universal primer regions P1 and P2, and a SOMAmer ID sequence I S . As part of generating second probes, SOMAmer ID regions can be appended to universal primers to generate amplifiable reporters.
[0082] Figure 2B is an exemplary method 200 for producing a set of SOMAmer probes for a content assay. In step 202, a set of SOMAmer sequences for a content assay is provided. Figure 1A flowchart of the steps of an exemplary method 250 of the tri-molecular complex) is shown. In step 252, a SOMAmer-containing eluate is provided, the SOMAmers in the eluate indicating the presence of one or more target proteins or other target molecules in one or more biological samples exposed to a SOMAmer library, as previously described.
[0083] In step 254, a set of SOMAmer-specific capture probes or a SOMAmer-specific capture probe library produced by method 200 is combined with the SOMAmer-containing eluate following the assay. In one example, 25 μΐ of the SOMAmer-containing eluate is combined with 25 μΐ of the probe-containing solution to form a hybridization solution having a volume of 50 μΐ. To facilitate hybridization, the concentration of capture probes can be comparable to or higher than the concentration of SOMAmers in the eluate. For example, a suitable probe concentration can be in the range of 0.05 nM to 5.0 nM, such as 0.5 nM (where nM is nanomoles per liter).
[0084] In some cases, in optional step 253 of method 250, the SOMAmer-containing eluate can be separated and selectively diluted prior to hybridization, and then recombined prior to sequencing. More specifically, the post-assay eluate can be separated into two or more dilution groups (e.g., four dilution groups) according to the expected relative abundance of SOMAmers in each group. The sample with the lowest degree of concentration (highest degree of dilution) can contain the largest amount of SOMAmers in the eluate. Conversely, the sample with the highest degree of concentration (lowest degree of dilution) can contain the smallest amount of SOMAmers in the eluate. In this way, the SOMAmer count can be "leveled" to improve the accuracy and precision of detecting smaller abundance SOMAmers. Each dilution group can then be hybridized individually by exposure to a corresponding subset of probes. Further details regarding the use of dilution groups are given in subsequent sections of this disclosure.
[0085] In addition, leveling can be achieved by introducing a fixed proportion of H1 probes with or without capture tags for certain high abundance SOMAmers in the eluate. SOMAmers that form tri-molecular complexes with H1 probes lacking capture tags will be removed in the washing step in method 300, as described in detail later.
[0086] In step 256, the first probes and the second probes are hybridized to the SOMAmers to form tri-molecular complexes, each tri-molecular complex including (i) a SOMAmer, (ii) a first probe bound to the SOMAmer through a hybridization region H1, and (iii) a second probe bound to the SOMAmer through a hybridization region H2. Each second probe includes a SOMAmer ID sequence I Swhich can be sequenced to indicate the presence of the corresponding SOMAmer in the eluate, and thus the presence of the corresponding protein captured by that SOMAmer in the original biological sample. Hybridization of the capture probes to the SOMAmers can be accomplished using any suitable technique, such as appropriate thermal cycling, and can include additives to enhance hybridization kinetics.
[0087] Figure 3A is an exemplary method 300 of performing a next generation sequencing content assay using a trimolecular complex, such as Figure 1 as shown and produced by methods such as Figure 2B In step 302, a hybridized trimolecular complex corresponding to the desired post-content assay SOMAmers library (e.g., each SOMAmer has the structure of complex 100 and is produced by methods such as method 250) is provided for sequencing content assay.
[0088] In step 304, the trimolecular complex is captured on a magnetic bead. For example, the complex can be captured by binding the biotin attached to the hybridization region H1 of the first probe to streptavidin on the bead. Capture can be accomplished by any suitable technique. For example, in one example, a solution of 30 μΐ hybridization volume is combined containing the beads at a concentration of 20 mg / ml, and then mixed by a thermal mixer at 1200 rpm for 30 minutes at a temperature of 45°C.
[0089] In step 306, the solution containing the bead-captured probes is washed one or more times to remove unbound H2 probe reporters, i.e., probes that are not hybridized to a corresponding SOMAmer. For example, the wash can be performed with a suitable buffer solution, such as a 20 mM phosphate buffer solution containing 1 mM EDTA and 0.05% sodium dodecyl sulfate (SDS). The wash phase can be performed statically, dynamically using a thermal mixer, or a combination of the two types sequentially. In one example, there can be two 5 minute static wash phases and two 10 minute dynamic wash phases at 1200 rpm. In any case, the resulting solution after washing should contain the trimolecular complex bound to the beads, with at most a small amount of unbound H2 probe / reporter remaining.
[0090] In certain embodiments of SOMAmer eluate leveling or dynamic range compression, trimolecular complexes that lack the bead capture tag on H1 will be removed in step 306 along with unbound H2 probe reporters. These complexes will result in a reduced copy number of those SOMAmers in the final NGS sequencing, and thus a reduced number of those abundant SOMAmers.
[0091] In step 308, the bound ternary complex is eluted from its attached beads, for example, by exposure to a solvent, heat, or by any other suitable elution method. As part of this step, the components of the complex may also dissociate, resulting in separated SOMAmers and probes. In one example, the complex is eluted by adding 85 μl of 20 mM NaOH to the eluent containing the bound complex, followed by mixing with a hot mixer at 1200 rpm for 3 minutes, and then dissociating the complex for 5 minutes. The solution containing the eluted complex is then mixed with 20 μl of hydrochloric acid.
[0092] In some examples, in optional step 309 of method 300, the eluted solution (i.e., the eluent) produced in step 308 can be partitioned and / or diluted into two or more groups, such as four dilution groups or primer amplification groups. As described in the context of the preceding method 250, using multiple dilution groups or primer amplification groups (which may not be diluted in some cases) corresponding to subsets of SOMAmers with different expected abundances levels to level the relative abundance of the entire SOMAmers set or compress its wide distribution can result in greater accuracy and precision when detecting relatively scarce target molecules. The partitioning into these groups can be performed before hybridization (as in step 253 of method 20) or after hybridization (as in step 309 as described here), or both. Further details of possible dilution and recombination techniques will be described in subsequent sections of this disclosure.
[0093] In step 310, the solution produced by steps 308 and optionally 309 is prepared for next-generation sequencing (NGS). This may include solutions containing universal primers ( Figure 1 (P1 and P2 in the sequence) and the associated SOMAmer ID sequence I S PCR amplification of the reporter region. If the elution buffers from multiple samples are to be combined before sequencing, the NGS preparation process may also include ligating aptamer sequences and / or barcode sequences for demultiplexing. The generation and ligation of aptamer and barcode sequences in the reporter region can be performed in any suitable manner known in the art, which is common practice in preparing samples for next-generation sequencing. In some cases, barcode sequences may be added as part of a first preparation step, and NGS aptamers may be added as part of a second preparation step.
[0094] In optional step 311, the groups that remain separated after step 310 can be recombined to prepare for sequencing.
[0095] In step 312, the prepared sample is sequenced using a next generation sequencing technology. In some examples, the prepared sample can be sequenced using a next generation sequencing platform developed by Illumina, Inc. of San Diego, California. However, the methods of the present disclosure are also suitable for use with other NGS sequencing platforms.
[0096] Following NGS, the sequencing-derived data can be analyzed or otherwise processed in optional step 314 to determine the concentration of the analyte (e.g., the protein of interest) in the original biological sample. Generally, such analysis includes demultiplexing the sequencing data using the barcode corresponding to each original sample (if multiple samples were multiplexed), counting the reporter factor sequences, and scaling and / or normalizing the data to extract accurate results. For analysis, the sequencing data can be written to a data file in a standard format, such as the ADAT format developed by SomaLogic. Possible methods of quantitative analysis are discussed in more detail below.
[0097] Figure 3B is a flowchart of steps of an exemplary method 350 of performing a next generation sequencing content assay, which includes capturing a target protein with an aptamer, forming a ternary complex from the aptamer, and then using the ternary complex as a basis for identifying the captured target protein. It will be appreciated that any of the steps of method 350 can be similar to corresponding steps of the previously described methods (i.e., methods 200 and 300), and thus the same details are not described again.
[0098] In step 352 of method 350, target proteins are captured by exposing a biological sample to a plurality of aptamers, such as SOMAmers, each of which is configured to bind to a particular protein. By exposing the sample to a library containing many such SOMAmers, a large number of target protein species can be detected in a single content assay.
[0099] In step 354, the aptamer that captured one of the target proteins is isolated in an aptamer-containing eluate. For example, in the course of a SomaLogic-performed assay called SomaScan, an aptamer-containing eluate can be formed that includes the process of binding the aptamer to a content assay bead, capturing proteins with the aptamer, washing away unbound proteins, labeling the bound proteins with biotin, releasing the aptamer from the bead, capturing the labeled proteins to a new bead, removing unbound aptamer, denaturing the aptamer from the captured proteins, and then isolating the aptamer to an eluate.
[0100] In optional step 356, the aptamer-containing eluate can be divided into a plurality of groups, which can be dilution groups, more details of which are provided below in connection with method 400.Figures 4-6 and Figure 11 as shown.
[0101] In step 358, a plurality of trimolecular complexes is formed by exposing the aptamers in the eluate (or each eluate dilution set) to a plurality of capture probes, each capture probe configured to hybridize with a particular aptamer. For example, each complex can have a structure similar to complex 100 shown in Figure 1 Thus, each trimolecular complex includes (i) a particular one of the aptamers from the eluate; (ii) a first capture probe including a portion that hybridizes to a first portion of the aptamer; and, (iii) a second capture probe including a portion that hybridizes to a second portion of the aptamer, further including one or more DNA primer regions and an aptamer ID sequence corresponding to the particular aptamer. If separate dilution sets are formed, each dilution set will be exposed to a particular set of capture probes corresponding to a particular subset of aptamers. Optionally, some H1 probes can lack a bead capture tag for additional leveling of SOMAmer counts.
[0102] In step 360, the sets formed in step 356, if any, can be recombined.
[0103] In step 362, the trimolecular complexes are separated from capture probes that are not bound to aptamers. For example, the hybridized complexes can be captured to magnetic beads, and then unbound probes are removed by washing.
[0104] In step 364, the capture probes in the trimolecular complexes are dissociated from the corresponding aptamers. This can include eluting the complexes from the beads, but in any event the result of step 362 is that the capture probes are no longer bound to aptamers.
[0105] In step 366, the eluate containing unbound capture probes can optionally be separated and / or diluted (possibly a second dilution, as described below in Figure 6 to form a plurality of PCR sets.
[0106] In step 368, the aptamer ID sequences in the dissociated capture probes in the eluate are amplified, for example by PCR amplification of the DNA primer regions and associated ID sequences. Other preparations for NGS can also be performed at this stage, for example ligation of aptamer sequences and / or demultiplexing barcode sequences.
[0107] In step 370, the different PCR sets, if any, can be recombined.
[0108] In step 372, the aptamer ID sequences are sequenced using next generation sequencing technology. In some examples, the sequencing can be accomplished using a next generation sequencing platform developed by Illumina, Inc. of San Diego, CA.
[0109] In step 374, the data obtained by sequencing the aptamer ID sequences can be used to determine the abundance of the target protein in the original biological sample.
[0110] C. Exemplary dilution panels or dynamic range compression for next generation sequencing
[0111] This section describes possible approaches for achieving higher assay efficiency, reproducibility, performance, and / or production feasibility in next generation sequencing systems in accordance with aspects of the present specification, see Figures 4-6 .
[0112] First, it should be understood that NGS can be performed on SOMAmer-containing eluate without dividing the eluate into multiple dilution groups, i.e., on a single elution solution that has never been grouped, diluted, or reconstituted. Such an assay is within the scope of the present disclosure and has the advantage of requiring less eluate and only one hybridization plate per 96 samples. However, such an assay faces challenges in sensitivity and accuracy, e.g., because the maximum range of SOMAmer abundance in the original eluate can span multiple orders of magnitude. For example, the target protein in the sample can have a concentration in the fM- mM range (i.e., spanning about 9 orders of magnitude), resulting in eluted SOMAmer concentrations spanning 5 or more orders of magnitude. Assaying such an eluate can result in overcounting of more SOMAmers and undercounting of fewer SOMAmers. Therefore, it can be desirable to level or compress the range of SOMAmer abundances prior to sequencing and counting.
[0113] The systems and methods of the present disclosure address this problem by subdividing the SOMAmers into subgroups prior to counting by incorporating dilution, thereby enabling leveling or dynamic range compression of counts across subgroups when they are combined together for sequencing and counting. In some examples, the SOMAmer probe set is subdivided into multiple subgroups (first group of very few SOMAmers, second group of few SOMAmers, third group of abundant SOMAmers, etc.) according to SOMAmer eluate abundance. The dynamic range of each subgroup is smaller, and in some examples much smaller, than the dynamic range of the original (un-divided) eluate. As described below, the dilution groups can be formed prior to and / or after hybridization of the SOMAmers to the probes, i.e., prior to and / or after formation of the tri-molecular complex suitable for NGS.
[0114] In addition to leveling by dilution, high abundance SOMAmers can also be "leveled" by introducing H1 probes that lack the bead capture tag as well as probes that contain these tags. The ratio of H1 probes that do and do not contain the bead capture tag will cause the tri-molecular complexes captured in step 304 of method 300 and step 358 of method 350 to be reduced by an amount corresponding to the ratio. For example, if the ratio of untagged probes to tagged probes is 10: 1, then only 10% of the tri-molecular complexes are captured, and the counts in the NGS output are reduced by an order of magnitude compared to an assay that does not introduce untagged probes. The ratio of untagged probes to tagged probes can be different for different SOMAmers, depending on the expected counts for each SOMAmer.
[0115] 1. Four hybridization panels
[0116] Figure 4 The steps and byproducts of an exemplary NGS assay involving four hybridization panels (generally indicated at 400) are shown. In step 402, a SOMAmer-containing eluate 404 is provided. As previously described, eluate 404 contains SOMAmers that were produced by prior exposure to a biological sample and separation from target molecules, as previously described.
[0117] In step 406, eluate 404 is divided into four equal portions or samples 408, 410, 412, and 414. In this example, each of the four samples is diluted by a different amount. In step 408, sample 408 is diluted by a ratio of 1 : 16, i.e., one part eluate to sixteen parts buffer solution; in step 410, sample 410 is diluted by a ratio of 1 :4; and in step 412, sample 412 is diluted by a ratio of 1 :2. In step 414, sample 414 is not diluted. Figure 4
[0118] In step 416, the four samples are each combined with a set of hybridization capture probes, hybridization capture probe sets 418, 420, 422, and 424 are labeled "Panel 1," "Panel 2," "Panel 3," and "Panel 4," respectively. In this example, capture probe set 418 is combined with the eluate that was diluted the most, and therefore contains capture probes configured to bind to the most common SOMAmers in the eluate. Similarly, capture probe set 420 contains probes configured to bind to the second most common SOMAmers, and capture probe sets 422 and 424 both contain probes configured to bind to different subsets of relatively less common SOMAmers. The result of step 416 is therefore four different solutions, each configured to produce a set of tri-molecular compounds upon hybridization, each compound including a SOMAmer and the corresponding probe that hybridizes to the SOMAmer. Each compound can be roughly analogous to compound 100 in Figure 1
[0119] Any of the four sets of hybridization capture probes may contain a fixed proportion of H1 probes that include and exclude some subset of bead-capturing markers for SOMAmers within each set, for further count leveling.
[0120] In step 426, the four solutions generated in step 416 are respectively hybridized, captured onto beads, washed, and eluted. This can be done according to the previously described... Figure 3A Steps 304, 306, and 308 of method 300 shown are implemented.
[0121] In step 428, the different eluents generated in step 426 are recombine into a single eluent 430. In some cases, the different solutions may be recombinated in different volumes, thereby further diluting the relatively abundant capture groups (i.e., those corresponding to abundant SOMAmers, and thus to the abundant target molecule species in the original biological sample). This produces normalized combined solutions with smaller overall variations in the concentrations of different three-molecule compounds, which can be analyzed with relatively fewer sequencing “reads.” For example, the combination of dilution and normalization may reduce the number of reads required per sample from approximately 200 million to less than 5 million, allowing multiple samples to be reused in each sequencing cycle and reducing the cost per sample while still achieving acceptable accuracy (measured by the coefficient of variation (CV) in the results).
[0122] In step 432 (which can be considered a combination of steps 310, 312, and 314 of the aforementioned method 300), the solution 430 generated in step 428 is prepared for next-generation sequencing (NGS), sequenced, and the results are written to a data file and analyzed as needed. Preparation may include the use of universal primers ( Figure 1 P1 and P2) and related SOMAmer ID sequence I S PCR amplification is performed on the reporter region. As previously mentioned, preparation may also include ligating the aptamer sequence and / or the barcode sequence for demultiplexing to the trimolecular compound. The prepared solution is then sequenced using NGS technology, such as a next-generation sequencing platform developed by Immena, San Diego, California, or any other NGS sequencing platform. Following NGS, the sequencing data is analyzed or otherwise processed to determine the concentration of the analyte, such as the target protein, in the original biological sample. This may include demultiplexing the sequencing data using barcodes corresponding to each original sample (if multiple samples are pooled), counting reporter factor sequences, and scaling and / or normalizing the data to extract accurate results. Sequencing data may be written to a data file in a standard format, such as the ADAT format developed by SomaLogic.
[0123] 2. One hybridization panel and four PCR panels
[0124] Figure 5 The steps and byproducts of an exemplary NGS content assay (generally denoted 500) involving one hybridization panel and four PCR panels are shown. In step 502, an eluate 504 containing SOMAmers is provided. As described previously, the eluate 504 contains SOMAmers produced from a previous exposure to a biological sample and separation from the target molecules.
[0125] In step 506, the eluate 504 is combined with a full panel of hybridization capture probes 508 (i.e., a panel of probes configured to bind to all SOMAmers in the eluate). As described above, this panel of probes can also contain a fixed proportion of labeled and unlabeled Hl probes.
[0126] In step 510, the solution produced in step 506 is hybridized, captured to beads, washed, and eluted. This can be accomplished in accordance with steps 304, 306, and 308 of the method 300 described in Figure 3A However, in this case, four sets of universal primers can be used instead of one, each set of primers associated with a particular subset of SOMAmer IDs corresponding to a panel of SOMAmers falling within a particular expected concentration range. In other words, step 510 produces four different sets of trimeric compounds corresponding to different abundance groups of SOMAmers and thus to different abundance groups of target molecules in the original biological sample from which the SOMAmer eluate was produced, each set of trimeric compounds containing different PCR primers and thus can be amplified separately.
[0127] In step 512, the eluate produced by step 510 is divided into four equal parts or aliquots 514, 516, 518, and 520. At this stage, each aliquot can optionally be diluted to any desired degree to normalize the expected concentration of SOMAmer ID sequences to be amplified in the next step. However, Figure 5 No dilution in step 512 is described.
[0128] In step 522, the different eluates produced from step 512 are prepared for next generation sequencing (NGS), including PCR amplification of the reporter regions. However, in this case, a different set of primers and associated reporter regions are amplified in each separate eluate, resulting in only a known subset of SOMAmer ID sequences being amplified in each eluate.
[0129] In step 524, the different solutions containing amplified SOMAmer ID sequences in each sample are recombined into a single eluate 526. In some cases, the separated solutions can be recombined at different volumes, achieving the degree of dilution required for relatively abundant SOMAmer ID sequences, and resulting in a normalized combined solution with a generally smaller overall change in SOMAmer ID concentration, which can be analyzed with relatively fewer reads.
[0130] In step 528, solution 526 is further prepared for next generation sequencing (NGS), sequenced, and the results written to data files and analyzed as needed. Preparation of the amplified eluate can include ligation of adaptor sequences and / or barcode sequences for demultiplexing to the reporter region of the tri-molecular compounds. The prepared solution is then sequenced using NGS technology, after which the data obtained from sequencing can be analyzed or otherwise processed to determine the concentration of the target analyte in the original biological sample, as previously described.
[0131] 3. Four hybridization panels and four PCR panels
[0132] Figure 6 The steps and byproducts of an exemplary NGS content assay (generally referenced at 600) involving four hybridization panels and four PCR panels are shown, and thus incorporate aspects of content assays 400 and 500 of Figures 4-5 . In step 602, a SOMAmer-containing eluate 604 is provided, which includes SOMAmers produced from a previous exposure to a biological sample and separation from the target molecules.
[0133] In step 606, eluate 604 is divided into four equal portions or samples 608, 610, 612, and 614. These samples are optionally diluted to different degrees, or in some cases, the samples can not be diluted.
[0134] In step 616, the four samples are each combined with a set of hybridization capture probes, hybridization capture probe sets 618, 620, 622, and 624 are labeled “Panel 1,” “Panel 2,” “Panel 3,” and “Panel 4,” respectively. Each set of probes is configured to bind to a particular subset of SOMAmers in the eluate, and then the samples are hybridized separately. Thus, the result of step 616 is four different solutions, each containing a set of tri-molecular compounds, including a SOMAmer and the corresponding probe hybridized to the SOMAmer. Each compound can be generally similar to Figure 1 compound 100 of
[0135] Any of the four sets of hybridization capture probes can optionally include a fixed ratio of unlabeled and labeled H1 probes for further count leveling.
[0136] In step 626, the four hybridization solution sets produced in step 616 are combined into a single eluate 628, which is then captured onto beads, washed, and eluted in sequence. This can be accomplished according to the methods 300 shown in steps 304, 306, and 308, as previously described. As previously described, each hybridization solution can or can not be diluted prior to recombination, and the differential volumes of each set can also be used to compress the variation of the analyte prior to bead capture and washing. Figure 3A
[0137] In step 630, the eluate produced in step 626 is divided into four equal parts or samples 632, 634, 636, and 638. At this stage, each sample can optionally be diluted to any desired degree to normalize the expected concentration of the SOMAmer ID sequences to be amplified in the next step. However, Figure 6 No dilution in step 630 is described.
[0138] In step 640, the different eluates produced from step 630 are used for next generation sequencing (NGS), a process that includes PCR amplification of the reporter region. As Figure 5 As shown in content assay 500, different primer sets and associated reporter regions are amplified in each of the separate eluates, thereby amplifying a subset of the SOMAmer ID sequences in each eluate.
[0139] In step 642, the different solutions containing amplified SOMAmer ID sequences in each sample are recombined into a single eluate 644. In some cases, the different solutions can be recombined at different volumes, achieving the degree of dilution required for relatively abundant SOMAmer ID sequences, as well as resulting in a normalized combined solution with a relatively small overall variation in SOMAmer ID concentration, which can be analyzed with relatively fewer reads.
[0140] In step 646, solution 644 is further prepared for NGS and sequencing, and the results are written to a data file and analyzed as desired. Preparation of the amplified eluate can include ligation of adaptor sequences and / or barcode sequences for demultiplexing to the reporter region of the tri-molecular compound. The prepared solution is then sequenced using NGS technology, after which the data obtained from sequencing can be analyzed or otherwise processed to determine the concentration of the target analyte in the original biological sample.
[0141] D. PCR panel quantification spike normalization
[0142] In NGS-based systems, signals are measured as sequence read counts, with the read counts for all analytes measured in a given sample in the same sequencing run mixed together in a fixed or limited set of total read counts. As parts of the same mixture, the NGS read counts for all analytes measured in a given sample influence each other, so the signal counts observed from each analyte are the "net result" of all the increases and decreases of the analytes measured in a "zero-sum game" for each sample with a fixed total number of reads. More specifically, in NGS systems with a fixed total number of reads, any increase in the count of one analyte results in a corresponding decrease in the count of the other analytes, distributed according to the proportion of each analyte in the total number of reads.
[0143] Figure 7 This "zero-sum game" is graphically depicted by showing the results of a simplified two-analyte assay for two different samples in the form of a histogram, where the vertical axis represents the total number of reads for each analyte, and the total number of reads is fixed at 2 million. In Sample 1, the counts for analytes A and B are equal. In Sample 2, analyte A is increased by 0.5 million counts, and analyte B is correspondingly decreased by 0.5 million counts. However, since the total number of reads is limited, it is not possible to know from Figure 7 Sample 2 whether the difference in counts between analytes A and B in Sample 2 is due to the increase in analyte A, the decrease in analyte B, or a combination of both.
[0144] Figure 8 It is described how to normalize the compensated reads across samples by introducing a reference reporter, or quantitative spike control reporter ("qSpike"), allowing the identification of true signal changes. In the NGS assay shown in Figure 8 In the NGS assay shown in FIG. 1, a qSpike reporter is physically added (spiked) into a two-analyte system containing analytes A and B, with all samples having the same known concentration. In this case, the qSpike reporter is exactly the same concentration as analytes A and B in Sample 1. In Sample 2, as before, an increase in analyte A and a decrease in analyte B are observed. However, it is now observed that the qSpike is decreased in Sample 2 relative to Sample 1, while it is known that Sample 1 has the same concentration of qSpike added. As shown in the "Sample 2 qSpike adjustment" histogram in FIG. 2, proportionally adjusting / adjusting all analytes in Sample 2 to push the spike back to its expected concentration results in a clearly visible increase in analyte A relative to analyte B. Figure 8
[0145] In more practical NGS content assays, a mixture of multiple qSpike reference reporter factors can be used to correct for compensatory changes in the proportion of analyte counts. For example, a content assay according to the present specification can use a mixture of four unique H2 reporter factors, i.e., four unique amplifiable reporter factors form part of a second probe that is introduced after elution of the trimolecular complex in step 308 of the content assay 300 shown in Figure 3A Alternatively, specific qSpike SOMAmers can be introduced into the elution fluid and the SOMAmer-specific probe pool contains the appropriate qSpike reporter factors. The qSpike reporter factors or SOMAmers can be provided at different relative concentrations.
[0146] Figures 9-10 are histograms depicting the raw results and the qSpike-adjusted results, respectively, of such a content assay, with the legend reference having the following meanings:
[0147] • QSpike-H is a high concentration qSpike reporter factor
[0148] • QSpike-MH is a medium-high concentration qSpike reporter factor
[0149] • QSpike-ML is a medium-low concentration qSpike reporter factor
[0150] • QSpike-L is a low concentration qSpike reporter factor
[0151] • Apolipoprotein E2, transferrin, kininogen HMW are the SOMAmer analytes
[0152] In the content assay shown in Figures 9-10 Three SOMAmers, Apolipoprotein E2, transferrin, and kininogen HMW, were titrated in buffer at a concentration range of 50 pM - 50 aM and measured in an NGS HC-content assay. Each measurement point was a separate sample in which the three SOMAmer analytes were sequenced together, and qSpike was added to all samples at the same concentration.
[0153] In Figure 9 (A), the qSpike exhibits compensatory changes due to changes in the analyte dose response signal. The NGS content assay signal (reads) is compared to the known spike and a scaling factor is generated. In Figure 10 (B), the NGS counts have been scaled so that the spike is uniform for all samples, which restores the actual SOMAmer dose response to the three SOMAmers measured in the content assay.
[0154] The use of qSpike reporters to compensate for limited read counts can be incorporated into any of the next generation sequencing content assays described above. For example, Figure 11 Some steps and byproducts of an exemplary NGS content assay (generically represented as 1100) are described, in which four PCR panels are involved and qSpike reporters are added in each panel. Thus, the steps of content assay 1100 can be incorporated into any NGS content assay that uses multiple PCR panels, such as content assays 500 and 600 shown in Figures 5-6 and 600 shown in
[0155] In step 1102, an eluate 1104 containing SOMAmers that have been hybridized to probes, captured onto beads, washed, and eluted is provided. Thus, eluate 1104 should be considered substantially similar to the eluate produced by, for example, Figure 5 step 510 of content assay 500 shown in Figure 6 step 626 of content assay 600 shown in
[0156] In step 1106, eluate 1104 is divided into four equal parts or samples 1108, 1110, 1112, and 1114. At this stage, each sample can optionally be diluted to any desired degree to normalize the expected concentration of SOMAmer ID sequences to be amplified in the next step.
[0157] In step 1116, the different eluates produced from step 1106 are prepared for next generation sequencing (NGS), including PCR amplification of the reporter regions. As in content assays 500 and 600, different primer sets and associated reporter regions are amplified in each individual eluate, such that a subset of SOMAmer ID sequences is amplified in each eluate. However, in this case, different qSpike reporters are added to each eluate at known concentrations prior to PCR amplification.
[0158] In step 1118, the different solutions containing amplified SOMAmer ID sequences and qSpike reporters in each sample are recombined into a single eluate 1120. In some cases, the different solutions can be recombined in different volumes, such that the degree of dilution required for relatively abundant SOMAmer ID sequences is achieved, and a normalized combined solution is obtained that has a relatively small overall variation in SOMAmer ID concentration that can be analyzed with relatively fewer reads.
[0159] In step 1122, the solution 1120 is further prepared for NGS, sequenced, and the results written to a data file and analyzed as needed. Preparation of the amplified eluate can include ligation of adaptor sequences and / or barcode sequences for demultiplexing to the amplified reporter region of the tri-molecular compound. The prepared solution is then sequenced using NGS technology, after which the data obtained from sequencing can be analyzed or otherwise processed to determine the concentration of the target analyte in the original biological sample. Due to the use of qSpike reporters, the analysis can include scaling or renormalizing the data to restore the qSpike concentration to a known level, thereby compensating for counting errors that can arise due to limited sequencing reads.
[0160] E. Assaying SOMAmer-probe stability
[0161] As previously described, according to aspects of the present description, hybridization regions H1 and H2 are configured to bind to respective compensating portions of the SOMAmer, and can be designed to have similar or deliberately different melting temperatures (T m ), to enable simultaneous hybridization under a given set of assay conditions. As discussed in this section, in some cases, the melting curves of experimentally determined SOMAmer-probe pairs can be used to calculate estimated melting temperatures of the hybridization regions.
[0162] 1. BACKGROUND
[0163] The most widely used method to predict the stability of nucleic acid duplexes is known as the nearest-neighbor model. The nearest-neighbor model assumes that the thermodynamic properties of helix formation depend primarily on the identity of adjacent base pairs in the duplex. This model has been widely applied to predict the stability of duplex formation in the design of primers required in PCR, as well as other applications where oligonucleotide duplex formation is critical. For NGS assays according to the present description (i.e., involving SOMAmers), it is necessary to extend this method to accurately predict the stability of duplexes composed of one strand containing a modified DNA base and another strand containing a natural DNA base.
[0164] Traditionally, the absorbance as a function of temperature curve (melting curve) measured with a UV-Vis spectrophotometer has been used to study the stability of DNA secondary structures. Hybridization is typically performed in 1.0 M NaCl, 10 mM sodium carboxylate, and 0.5 mM Na2EDTA buffer at pH 7. The oligonucleotide concentration is varied over a 100-fold range, and the thermodynamic parameters are obtained from the curve of the inverse melting temperature (T M -1 ) versus the natural logarithm of the total DNA concentration, and fitted to
[0165]
[0166] In addition, AH° and AS° can also be obtained from the melting curves, respectively, and averaged between different concentrations. Both methods are essentially a van't Hoff analysis of the data. The thermodynamic data obtained from both methods are typically within 10% agreement. This section will specifically use the latter method - separate fitting of the melting curves - to extract the thermodynamic parameters needed to predict the stability of the duplexes formed in the mix. In addition, the buffer composition from which the thermal melting curves are obtained will match the composition of the typical SOMAmer-containing assay readout.
[0167] According to various aspects of the present specification, an extension of the nearest neighbor model was developed for SOMAmers containing the three most common modified bases: Nap-dU, 2-Nap-dU, and benzyl-dU. As described below, melting curves were experimentally obtained for over 400 SOMAmer-probe pairs by fluorescence measurements. These data were used to define the nearest neighbor parameters needed to predict the stability of SOMAmer-probes under the conditions of the assay readout.
[0168] 2. Modeling double-stranded formation of SOMAmer-probe binding
[0169] SOMAmer-probe duplex formation follows the following process:
[0170]
[0171] where S is the SOMAmer, p is the hybridization probe, and S:p is the duplex. C T defined as the total concentration of initial DNA:
[0172] C T = [S] + [p],
[0173] Assuming equal initial concentrations of the SOMAmer and the probe, the following equation can be derived from stoichiometry:
[0174]
[0175] where a is the molar fraction of duplex. The equilibrium constant for duplex formation is
[0176]
[0177] where AH and AS are the enthalpy and entropy of duplex formation, T is the absolute temperature (K), and R is the gas constant (1.9872 cal / K mol). Substituting the equation into the concentrations gives:
[0178]
[0179] By definition, T M corresponds to the temperature at which equal amounts of duplex and non-duplex exist, i.e. The expression for the melting temperature is given as follows:
[0180]
[0181] 3. SOMAmer melt model
[0182] SOMAmers with internal structure can also be viewed as a simple two-state model as follows:
[0183]
[0184] where S and U are structured and unstructured SOMAmer, respectively. T defined as the total concentration of initial DNA,
[0185] C T = [S] + [U]
[0186] From chemical thermodynamics the following equation is derived
[0187]
[0188] where β is the mole fraction of structured SOMAmer. The equilibrium constant for structured SOMAmer is
[0189]
[0190] where ΔΗ and Δ5 are the enthalpy and entropy of structure formation, T is the absolute temperature (K), and R is the gas constant (1.9872 cal / K mol). Substituting the concentration equations into the mole fraction of structured SOMAmer gives
[0191]
[0192] Similarly, by definition, T M corresponds to the temperature at which the amounts of structured and unstructured SOMAmer are equal, i.e. as follows:
[0193]
[0194] The concentration of SOMAmer does not contribute to the entropy because this is a single molecule reaction and, under conditions of proper dilution, all reactions are independent. The simplest model assumes that SOMAmer melting and subsequent primer melting are independent processes, and vice versa. Without additional experiments, such as a separate SOMAmer melt, it is not possible to know a priori which of the two transitions is due to SOMAmer structure melting or hybridization primer melting. Primer melting is most likely the higher free energy data because it corresponds to melting better than 16 base pairs.
[0195] 4. Experimental determination of melting thermodynamics
[0196] The fluorescence intensity as a function of temperature was measured (melting curve) using the fluorescent dye SYBR Green I. SYBR Green I increases its fluorescence intensity 100-fold upon binding to double-stranded DNA compared to single-stranded DNA, thus the fluorescence intensity decreases as the SOMAmer-probe double-stranded structure melts.
[0197] Thermal melting of the defined components in the elution buffer, i.e. 100 mM Tris-Hydroxymethylaminomethane, pH 8.0, 200 mM NaCl and 0.9 M perchlorate, was performed using SOMAscan content assays. Perchlorate is known to decrease the stability of DNA double strands. All thermal melting was achieved when the concentration of both the SOMAmer and the probe was 100 pM in 120 μL (8.3 x 10 -7 M). Thus, all thermodynamic parameters were obtained by single fitting of individual thermal melting curves. Melting curves that showed more complex behavior than expected for a bimodal model of double-strand formation were excluded from the analysis. Four plates were measured for each of the H1 and H2 probes. The 800 melting curves were evaluated to obtain data consistent with a hypothetical bimodal model of double-strand formation. Of the 800 curves, 408 melting curves of SOMAmers containing three different modified nucleotides, Nap-dU, 2-Nap-dU and benzyl-dU, were used in this analysis.
[0198] a. Single phase model fitting
[0199] Figure 12 Data of a typical thermal melting of a SOMAmer-probe double strand is shown, where the vertical axis is the fluorescence measured in RFU and the horizontal axis is the temperature, and is overlaid by a bimodal model fit as follows. First, the high and low temperature baselines are data fits using the first 15 data points and the last 15 data points. The low temperature baseline corresponds to the double-stranded material and the high temperature baseline corresponds to the single-stranded material, denoted as
[0200] bl ds = b ds + m ds T
[0201] bl ss = b ss + m ss T.
[0202] For a given value of ΔΗ and Δ5, the fraction of double strand as a function of temperature is obtained by first calculating K and then calculating a,
[0203]
[0204] where
[0205] b = 1 + 1 / (KC T ).
[0206] The curve of the hot melt, RFU(T), is calculated according to:
[0207] RFU(T) = a b ds + (1 - a) b ss .
[0208] The optimal values of the six free parameters b ds , m ds , b ss , m ss , AH, and AS are found using non-linear regression. The initial estimates of the single- and double-stranded baselines are as described above, and the initial values of AH and AS are -200 kcal / mol and -0.6 kcal / mol K, respectively. Figure 12 The model fit to the data in Figure 6 is shown as the red solid line in Figure 6. The subscript 'p' denotes the SOMAmer-probe thermodynamics. This model fits the data very well. Figure 12
[0209] b. Two phase model fitting
[0210] Typically, the data exhibit more complex melting behavior, which is likely due to the melting of the internal SOMAmer structure first, followed by the melting of the double-stranded SOMAmer-probe. Figure 13 Data representing a typical two-phase behavior are shown, and again the theoretical model fit is overlaid by the red solid line. Two distinct transitions are apparent in the data, the first likely being the melting of the internal SOMAmer structure, followed by the melting of the double-stranded SOMAmer-probe. It is assumed that these two transitions are independent. To fit a two-phase model, three baselines are required. The first corresponds to the temperature dependence on the internal SOMAmer structure, the second is the temperature dependence on the double-stranded SOMAmer-probe, and the third is the temperature dependence on the combined single-stranded material. The latter two are the same as described above. The former is denoted as
[0211] b int = b int + m int T
[0212] where the fluorescence is assumed to be additive, so that the fluorescence of the double-stranded SOMAmer-probe complex is the sum of the fluorescence of each individual structure. Let the molar fraction of the internal SOMAmer structure be β, and the molar fraction of the SOMAmer-probe double-stranded structure be α, then the thermal melting curve is
[0213] RFU(T) = β(bl int -bl ds )+ αbl ds +(1-α)bl ss .
[0214] At low temperatures, both α and β are 1, and the temperature dependence is given by bl int . As the SOMAmer structure melts, the fluorescence reaches the baseline for the duplex. As in the single phase case, the two-phase model fits the data very well.
[0215] 5. Nearest neighbor model
[0216] Once the experimental data is fit to an appropriate model (e.g., single phase or two phase as described above), and the various thermodynamic parameters are in agreement, the parameters for the nearest neighbor model can be obtained from the data. The change in free energy is approximated as:
[0217]
[0218] where the i subscript indicates the SOMAmer-probe double-stranded structure, ΔG j is the free energy of the nearest neighbor stacking interaction, n ij is the number of times the nearest neighbor j occurs in the double-stranded structure i, and ΔG(init) is the initial free energy due to entropic considerations. The nearest neighbor interactions include the ten standard Watson-Crick nearest neighbor stacking interactions (e.g., and the like). The notation (AC / TG) indicates that the 5'-AC-3' Watson-Crick base pairs with the 3'-TG-5' base pair. In addition, each modified base introduces an additional 7 modified nucleotide stacking interactions (e.g., and the like), where X indicates a modified T nucleotide that base pairs with a standard A nucleotide. ΔH and ΔS have similar expressions.
[0219] For a set of d double-stranded and n nearest neighbor interactions, the free energy of the double-stranded structure is given by: ijA "stacked matrix" matrix N of dimension d x n is constructed from the sequence data. The experimental observed thermodynamic values T total are represented as a column vector of length d. The unknown nearest neighbor stacking interactions I nn are represented as a column vector of length n and are solved by using ordinary least squares regression to solve the following overdetermined linear equations:
[0220] NI nn = T total .
[0221] The parameter I nn takes the minimum value of the Euclidean l 2 norm, ||NI nn -T total ||.
[0222] 6. Exemplary results
[0223] In the exemplary process according to the above description, a nearest neighbor model parameter was expanded using a double stranded thermal melt containing three different modified nucleotides: Nap-dU, 2-Nap-dU, and benzyl-dU. Python code has been developed to calculate the thermodynamics and melting temperature based on these expanded nearest neighbor parameters.
[0224] The following table summarizes the 31 parameters required for the nearest neighbor model, including the standard 4 base 10 parameters and 7 additional parameters for each modified base. Included are the AH and AS values for each NN pair, as well as the total number of occurrences of each nearest neighbor parameter in the data set. (AT / TA) and (TA / AT) occur relatively infrequently in the data, as they only need to occur in the fixed regions.
[0225]
[0226] Using the parameters in the above table, one can calculate the estimated melting temperature T m for a double stranded molecule containing the modified nucleotide bases Nap-dU, 2-Nap-dU, and benzyl-dU. A similar process can be used to calculate the estimated melting temperature for any other SOMAmer containing double stranded molecule. According to aspects of the present description, these melting temperatures can then be used to determine where to divide the hybridization complement into the SOMAmer, i.e., the trimolecular compound (e.g., compound 100 shown in Figure 1 between the first and second hybridization regions H1 and H2 (and thus between the first and second probes 104 and 106) of the SOMAmer.
[0227] F. Illustrative combinations and additional examples
[0228] This section describes other aspects and features of systems and methods of detection and quantification of target molecules in a biological sample in accordance with aspects of the present description, which are presented by way of a series of paragraphs, without limitation, as a matter of clarity and efficiency, some or all of which can be combined in any appropriate manner, and / or with disclosures elsewhere in the application, including material incorporated by cross-reference. Each of these paragraphs can stand alone as a separate disclosure that incorporates any or all of the features of the other paragraphs, and / or the disclosure elsewhere in the application, including material incorporated by cross-reference. Certain paragraphs below expressly incorporate and further limit the other paragraphs, and provide some examples of suitable combinations, without limitation.
[0229] A system for quantifying abundance of a target protein in a biological sample, comprising a plurality of aptamers, each aptamer configured to bind to a particular target protein when a biological sample containing the target protein is exposed to the aptamer, thereby forming an aptamer-containing eluate; a plurality of capture probes, each capture probe configured to hybridize to a particular aptamer, wherein the plurality of capture probes comprises a first capture probe having a portion that hybridizes to a first portion of a particular aptamer, and a second capture probe having a portion that hybridizes to a second portion of the particular aptamer, a DNA primer region, and an aptamer ID sequence corresponding to the aptamer; means for exposing the aptamer in the eluate to the capture probes to form a plurality of tri-molecular complexes; and means for sequencing the aptamer ID sequence, thereby determining the abundance of the target protein in the biological sample.
[0230] B A system for quantifying abundance of two or more target proteins in a biological sample, comprising a plurality of aptamers, each aptamer configured to capture a particular protein in the sample, thereby forming an aptamer-containing eluate upon isolation of the aptamer that captured a target protein in the biological sample; a plurality of first probes, each first probe hybridizing to a respective first portion of a particular aptamer; a plurality of second probes, each second probe hybridizing to a respective second portion of a particular aptamer, each second probe comprising at least one DNA primer region, and an aptamer ID sequence corresponding to the particular aptamer; means for sequencing the aptamer ID sequence; and means for quantifying the abundance of the target protein based on the sequenced aptamer ID sequence.
[0231] C A system for detecting a target protein in a biological sample, comprising a plurality of aptamers, each aptamer configured to capture a particular protein; a plurality of first probes, each first probe comprising a portion that hybridizes to a respective first portion of one of the aptamers that captured a target protein; a plurality of second probes, each second probe comprising a portion that hybridizes to a respective second portion of one of the aptamers that captured a target protein, at least one DNA primer region, and an aptamer ID sequence corresponding to the aptamer that captured the target protein; means for amplifying the aptamer ID sequences; and means for sequencing the aptamer ID sequences to identify the aptamer ID sequences, thereby identifying the aptamer that captured the target protein and the target protein.
[0232] D The system of any of the preceding paragraphs, further comprising means for normalizing the compensated read counts across samples.
[0233] E The system of any of the preceding paragraphs, further comprising means for dynamic range compression of the aptamer abundances and / or the aptamer ID sequences prior to sequencing and counting.
[0234] Advantages, features, and benefits
[0235] The different embodiments and examples of the methods and systems for detecting and quantifying the presence of a target molecule in a biological sample described herein have many advantages over previously known approaches. For example, the illustrative embodiments and examples described herein enable quantification of aptamer-based protein detection using next-generation sequencing by simplifying the sequencing target from the aptamer to the aptamer identification sequence.
[0236] In addition, the illustrative embodiments and examples described herein enable accurate detection of the abundance of a target molecule across multiple orders of magnitude by separating the content assay eluate into multiple dilution groups at one or more stages of the content assay and then recombining prior to next-generation sequencing, among other benefits.
[0237] In addition, the illustrative embodiments and examples described herein enable correction of errors caused by limited sequencing reads by adding a quantitative peak reporting factor to the content assay eluate to normalize the compensated reads across samples, among other benefits.
[0238] No existing system or device is capable of these functions. However, not all of the embodiments and examples described herein have the same advantages or the same degree of advantages.
[0239] CONCLUSION
[0240] The foregoing disclosure can include a variety of different embodiments with separate utility. Although each of these embodiments has been disclosed in its preferred form, this disclosure is not to be construed in a limiting sense as the specific embodiments disclosed can be modified in various ways. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. The subject matter of the present disclosure includes all novel and nonobvious combinations and subcombinations of the various elements, features, functions and / or properties disclosed herein. The following claims particularly point out certain combinations and subcombinations regarded as novel and nonobvious. Other combinations and subcombinations can be claimed in dependent claims. Such claims, whether broader, narrower, equal, or different, are regarded as being included in the subject matter of the present disclosure.
Claims
1. A method of quantifying abundance of a target protein in a biological sample, comprising: capturing the target protein by exposing the biological sample to a plurality of aptamers, each of the plurality of aptamers configured to bind to a particular protein; forming an aptamer-containing eluate by isolating the aptamer that captured one of the target proteins; forming a plurality of tri-molecular complexes by exposing the aptamer in the eluate to a plurality of capture probes, each of the plurality of capture probes configured to hybridize to a particular aptamer, each of the tri-molecular complexes comprising: one of the plurality of aptamers from the eluate, a first capture probe comprising a portion that hybridizes to a first portion of the aptamer, and a second capture probe comprising a portion that hybridizes to a second portion of the aptamer, a DNA primer region, and an aptamer ID sequence corresponding to the aptamer; adding a content assay bead, separating the tri-molecular complex from capture probes that did not bind to the aptamer, obtaining a tri-molecular complex; dissociating the capture probes in the tri-molecular complex from the corresponding aptamer; amplifying the aptamer ID sequence of the dissociated second capture probe; sequencing the amplified aptamer ID sequence of the dissociated second capture probe by next generation sequencing; and determining the abundance of the target protein in the biological sample from data obtained by sequencing the aptamer ID sequence. wherein the aptamer is a SOMAmer, the protein-aptamer dissociation rate t 1 / 2 between 30 and 240 minutes, the first capture probe comprises one or more elements suitable for binding to the assay bead, the one or more elements comprising biotin for binding to streptavidin-coated beads.
2. The method of claim 1, further comprising dividing the aptamer-containing eluate into two or more groups prior to forming the tri-molecular complex.
3. The method of claim 1, further comprising dividing the dissociated capture probes into a plurality of aliquots prior to amplifying the aptamer ID sequence of the dissociated second capture probe, wherein amplifying the aptamer ID sequence comprises amplifying a subset of aptamer ID sequences in each of the aliquots, respectively.
4. The method of claim 3, further comprising diluting at least one of the aliquots prior to amplifying the aptamer ID sequence of the dissociated second capture probe.
5. The method of claim 3, further comprising adding a quantification spike-in reporter to each of the aliquots prior to amplifying the aptamer ID sequence of the dissociated second capture probe.
6. The method of claim 5, further comprising recombining the aliquots prior to sequencing the amplified aptamer ID sequence of the dissociated second capture probe.
7. The method of claim 1, wherein, the aptamer has chemically modified nucleotides.
8. A method of quantifying abundance of two or more target proteins in a biological sample, comprising: capturing the target protein by exposing the biological sample to a plurality of aptamers, each of the plurality of aptamers configured to capture a particular protein; forming an aptamer-containing eluate by isolating the aptamer that captured one of the plurality of target proteins in the biological sample; forming a plurality of tri-molecular complexes, each of the tri-molecular complexes comprising: a specific aptamer present in the aptamer-containing eluate, a first probe hybridized to a corresponding first portion of the specific aptamer, and a second probe comprising a portion hybridized to a corresponding second portion of the specific aptamer, at least one DNA primer region, and an aptamer ID sequence corresponding to the specific aptamer; adding content measurement beads to separate the tri-molecular complex from probes not bound to the specific aptamer, resulting in a tri-molecular complex; dissociating the probe in the tri-molecular complex from the corresponding aptamer; amplifying the aptamer ID sequence of the second probe; and sequencing the amplified aptamer ID sequence of the second probe; and quantifying the abundance of the target protein from the sequenced aptamer ID sequence. wherein the aptamer is a SOMAmer, the protein-aptamer dissociation rate t 1 / 2 Between 30 and 240 minutes, the first probe comprises one or more elements suitable for binding to an assay bead, the one or more elements comprising biotin for binding to a streptavidin-coated bead.
9. The method of claim 8, further comprising dividing the aptamer-containing eluate into two or more dilution groups prior to forming the tri-molecular complex.
10. The method of claim 8, further comprising dividing a capture probe into a plurality of aliquots prior to amplifying the aptamer ID sequence of the second probe, wherein at least one aliquot of the plurality of aliquots is diluted, the amplifying the aptamer ID sequence of the second probe comprising amplifying a subset of aptamer ID sequences in each aliquot, respectively.
11. The method of claim 10, further comprising adding a quantification spike-in reporter to each aliquot prior to amplifying the aptamer ID sequence of the second probe.
12. The method of claim 11, further comprising recombining the aliquots prior to sequencing the amplified aptamer ID sequence of the second probe.
13. The method of claim 8, wherein, the sequencing the amplified aptamer ID sequence of the second probe is performed by next-generation sequencing.
14. The method of claim 8, wherein, the aptamer has chemically modified nucleotides.
15. A method of detecting a target protein in a biological sample, comprising: capturing the target protein by aptamer by binding the biological sample to a plurality of aptamers, each of the plurality of aptamers configured to bind to a specific protein; forming an aptamer-containing eluate by isolating aptamers that captured a target protein in the biological sample; forming a tri-molecular complex comprising: the aptamer that captured the target protein, a first probe comprising a portion hybridized to a corresponding first portion of the aptamer that captured the target protein, and a second probe comprising a portion hybridized to a corresponding second portion of the aptamer that captured the target protein, at least one DNA primer region, and an aptamer ID sequence corresponding to the aptamer that captured the target protein; adding content measurement beads to separate the tri-molecular complex from probes not bound to the aptamer, resulting in a tri-molecular complex; dissociating the probe in the tri-molecular complex from the corresponding aptamer; amplifying the aptamer ID sequence of the second probe; and sequencing the amplified aptamer ID sequence of the second probe to identify the aptamer ID sequence of the second probe, thereby identifying the aptamer that captured the target protein and identifying the target protein; wherein the aptamer is a SOMAmer, the protein-aptamer dissociation rate t 1 / 2 between 30 and 240 minutes, the first probe comprises one or more elements adapted to bind to an assay bead, the one or more elements comprising biotin for binding to a streptavidin-coated bead.
16. The method of claim 15, further comprising forming a first probe and a second probe by determining a target region on the aptamer, then determining a boundary for dividing a complementary structure of the target region into the first probe and the second probe.
17. The method of claim 16, wherein, determining the boundary based on a melting temperature required to reach a hybridized portion of the first probe and the second probe.
18. The method of claim 17, further comprising calculating a melting temperature of a complementary structure of the target region, then adjusting a boundary change between a hybridized portion of the first probe and the second probe step by step until a melting temperature required to reach the hybridized portion is reached.
19. The method of claim 17, wherein, the sequencing the amplified aptamer ID sequence of the second probe is done by next generation sequencing.
20. The method of claim 15, further comprising dissociating the second probe from the aptamer before amplifying the aptamer ID sequence of the second probe.
Citation Information
Patent Citations
Aptamer barcoding
CN112912512A
Synthetic nucleic acid spike-ins
US20170275691A1
Method for identification and analysis of certain molecules using the dual function of single strand nucleic acid
WO2005108609A1