Methods for detecting differentially abundant analytes - Patent Application 20070122997
By preparing aliquots based on analyte abundance and using internal controls with UMIs in PCR reactions, the method addresses signal interference in multiplexed detection, enabling accurate detection of proteins across varying concentrations.
Patent Information
- Application Number
- JP2022558354
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-03-27
- Filing Date
- 2021-03-26
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2041-03-26
AI Technical Summary
Existing multiplexed detection methods, such as PEA and PLA, struggle to accurately detect proteins over a wide concentration range due to signal interference from higher-abundance proteins masking lower-abundance proteins.
The method involves preparing multiple aliquots of a sample, performing separate multiplex assays on each aliquot based on analyte abundance, and using an internal control with unique molecular identifiers (UMIs) or reverse sequences in PCR reactions to enhance detection accuracy.
This approach allows reliable detection of analytes across a wide concentration range, improving the accuracy of multiplexed detection methods by minimizing signal interference and enabling precise quantification of multiple analytes.
Smart Images

Figure 0007815132000012 
Figure 0007815132000013 
Figure 0007815132000014
Abstract
Description
[Technical Field]
[0001] The present invention provides a method for detecting multiple analytes in a sample, the analytes being present at different levels of abundance in the sample. The method comprises preparing multiple aliquots of the sample, and detecting a subset of analytes in each aliquot, selected based on their predicted abundance in the sample. Also provided is a method for detecting analytes in a sample, wherein the analytes are detected by detecting a reporter nucleic acid molecule specific for the analyte. The method comprises performing a PCR reaction to amplify the reporter nucleic acid molecule, and the PCR employs an internal control. The method of the present invention is particularly useful in proximity extension assays (PEA).
[0002] background Modern proteomic techniques require the ability to detect many different proteins (or protein complexes) from small amounts of sample. This requires multiplexed analysis. Common methods that can achieve multiplexed detection of proteins in a sample include proximity extension assays (PEA) and proximity ligation assays (PLA). PEA and PLA are described in WO 01 / 61037, and PEA is further described in WO 03 / 044231, WO 2004 / 094456, WO 2005 / 123963, WO 2006 / 137932, and WO 2013 / 113699. However, when the proteins of interest are present over a wide concentration range, as is often the case, the signal from the higher-abundance proteins can drown out the signal from the lower-abundance proteins, resulting in the inability to detect the proteins present at low concentrations.
[0003] The present invention provides a detection method that can reliably detect analytes (e.g., proteins) present in a sample over a wide concentration range and improve the accuracy of multiplexed detection methods. The method of the present invention can be applied to PEA or PLA as described above, but may also be applied to other techniques used in multiplexed analyte detection.
[0004] PEA and PLA are proximity assays that rely on the principle of "proximity probing." In these methods, analytes are detected by binding to multiple (i.e., two or more, usually two or three) probes. These probes generate a signal when they bind to the analyte and are brought into proximity (hence the term "proximity probes"). Typically, at least one of the proximity probes contains a nucleic acid domain (or nucleic acid portion) linked to the analyte-binding domain (or analyte-binding portion) of the probe, and signal generation requires interactions between the nucleic acid portions and / or between the nucleic acid portion and additional functional portions carried by other probes. Thus, signal generation depends on interactions between the probes (more specifically, between the nucleic acid portions / nucleic acid domains or other functional portions / functional domains carried by them), and therefore, a signal is generated only when the required probe binds to the analyte. This results in improved specificity of the detection system.
[0005] In PEA, nucleic acid moieties linked to the analyte-binding domains of a probe pair hybridize to each other when the probes are brought into proximity (i.e., bound to the target) and are then extended by a nucleic acid polymerase. The extension product forms a reporter nucleic acid, the detection of which proves the presence of a specific analyte (the analyte bound by the associated probe pair) in a sample of interest. In PLA, binding of a probe of a probe pair to a target brings the nucleic acid moieties linked to the analyte-binding domains of the probe pair into proximity, allowing ligation between the nucleic acid moieties, or the nucleic acid moieties can serve together as templates for ligation of additionally added oligonucleotides that can hybridize to the nucleic acid domains upon proximity. The ligation product is then amplified and serves as the reporter nucleic acid. Multiplexed analyte detection using PEA or PLA may be achieved by including a unique barcode sequence in the nucleic acid moiety of each probe. Reporter nucleic acid molecules corresponding to specific analytes may be identified by the barcode sequence they contain. The methods of the present invention are particularly useful in multiplex PEA and multiplex PLA methods.
[0006] The methods of the present invention may be useful in at least any field in which proteomics is utilized, and may be particularly useful in diagnostics in terms of biomarker identification and quantification. Modern personalized medicine, for example in the field of oncology, requires the ability to evaluate large panels of biomarkers. As personalized medicine becomes ever more prevalent, it is becoming increasingly important to be able to accurately identify and quantitate many biomarkers (over a wide concentration range) in a sample. The present invention addresses this need.
[0007] Summary of the Invention To this end, in a first aspect, the present invention provides a method for detecting a plurality of analytes in a sample, said analytes being at different levels of abundance in said sample, said method comprising: (i) preparing a plurality of aliquots from said sample; (ii) detecting a different subset of the analytes in each aliquot by performing a separate multiplex assay on each aliquot, wherein the analytes in each subset are selected based on their expected abundance in the sample.
[0008] In a second aspect, the present invention provides a method for detecting an analyte in a sample, wherein the analyte is detected by detecting a reporter nucleic acid molecule specific for the analyte, the method comprising: performing a PCR reaction to generate a PCR product of the reporter nucleic acid molecule; and detecting the PCR product; An internal control is prepared for the PCR reaction, and the internal control comprises: (i) a separate component that is, contains, or produces a control nucleic acid molecule that is present in a predetermined amount and that is amplified with the same primers as the reporter nucleic acid molecule; and / or (ii) A unique molecular identifier (UMI) sequence present in each reporter nucleic acid molecule and / or each control nucleic acid molecule, which is unique to each molecule.
[0009] In a third aspect, the present invention provides a method for detecting an analyte in a sample, wherein the analyte is detected by detecting a reporter nucleic acid molecule for the analyte, the method comprising: performing a PCR reaction to generate a PCR product of the reporter nucleic acid molecule; and detecting the PCR product, wherein an internal control is included in the PCR reaction, the internal control being present in a predetermined amount and being, comprising, or generating a control nucleic acid molecule, the control nucleic acid molecule comprising a sequence which is the reverse sequence of the reporter nucleic acid molecule.
[0010] Detailed Description As detailed above, a first aspect of the present invention provides a method for detecting multiple analytes in a sample, the analytes being present at different levels of abundance in the sample, the method relying on performing a set of separate assays grouped according to the abundance of the analytes being analyzed.
[0011] Thus, in another aspect, the methods disclosed herein are methods for detecting a plurality of analytes in a sample, the analytes being at different levels of abundance in the sample, the method comprising: may be defined as a method comprising performing a separate block of assays on each of a plurality of separate aliquots obtained from the sample to detect a subset of analytes in each of the separate aliquots, the analytes in each subset being selected based on their expected abundance in the sample.
[0012] Thus, each assay block performed on an individual aliquot is a multiplex assay. Thus, a multiplex assay for detecting multiple analytes in an analyte subset (i.e., an analyte subset designated to be detected in any one particular aliquot) may be considered an "abundance-specific block." Thus, as used herein, the term "abundance-specific block" refers to a collection of assays (assay blocks) (or a set of assays (assay sets)) performed to detect a specific group or subset of analytes in a sample to be detected (i.e., analyzed), where the analytes are assigned to each assay block (or assay set) based on their abundance in the sample, i.e., their expected or predicted abundance or relative abundance in the sample. In other words, the assays are grouped or "blocked" based on abundance. Thus, different aliquots or different abundance-specific blocks may be designated for the detection of specific analyte subsets, for example, based on low abundance, high abundance, or varying degrees of intermediate abundance. This does not mean that the abundance of each analyte in an assay block or assay set is the same or approximately the same: abundance may vary between different analytes / assays in the block or set and / or may vary between different samples.
[0013] The term "analyte," as used herein (with respect to all aspects of the invention), refers to any substance (e.g., molecule) or entity that one desires to detect by the methods of the invention. Thus, an analyte is the "target" of the assay methods of the invention, i.e., the substance that is detected or screened for using the methods of the invention.
[0014] Thus, an analyte may be a biological molecule or compound desired to be detected, for example, a peptide or protein, or a nucleic acid molecule or a small molecule, and may include organic and inorganic molecules. An analyte may also be a cell, or a microorganism, including a virus, or a fragment or product thereof. Thus, it will be understood that an analyte may be any substance or object for which a specific binding partner (e.g., an affinity binding partner) can be developed. All that is required is that the analyte be capable of simultaneously binding to at least two binding partners (more specifically, the analyte-binding domains of at least two proximity-probes).
[0015] Proximity probe-based assays are particularly useful for the detection of proteins or polypeptides. Analytes of particular interest therefore include proteinaceous molecules such as peptides, polypeptides, proteins, or prions, or any molecule containing a protein or polypeptide component, or fragments thereof. In particularly preferred embodiments of the invention, the analyte is a fully proteinaceous molecule or a partially proteinaceous molecule, most particularly preferred is a protein. That is, the analyte preferably is or comprises a protein.
[0016] An analyte may be a single molecule or a complex containing two or more molecular subunits. The molecular subunits may or may not be covalently bound to each other. The molecular subunits may be the same or different. Thus, such complex analytes may be, in addition to cells or microorganisms, protein complexes or biomolecular complexes containing a protein and one or more other biomolecules. Thus, such complexes may be homomultimers or heteromultimers. Target analytes may also include aggregates of molecules such as proteins, e.g., aggregates of the same protein or aggregates of different proteins. An analyte may be a complex of a protein or peptide and a nucleic acid molecule such as DNA or RNA. Of particular interest may be an interaction between a protein and a nucleic acid, e.g., an interaction between a regulatory factor such as a transcription factor and DNA or RNA. Thus, in certain embodiments, the analyte is a protein-nucleic acid complex (e.g., a protein-DNA complex or a protein-RNA complex). In other embodiments, the analyte is a non-nucleic acid analyte, meaning an analyte that does not contain a nucleic acid molecule. Non-nucleic acid analytes include proteins and protein complexes, small molecules, and lipids, as described above.
[0017] The methods of the present invention involve detecting multiple analytes in a sample, which may be of the same type (e.g., all of the analytes may be proteins or protein complexes) or different types (e.g., some of the analytes may be proteins and others may be protein complexes, lipids, protein-DNA complexes, or protein-RNA complexes, or any combination of these types of analytes).
[0018] The term "multiple," as used in this disclosure, follows its standard definition: more than one (i.e., two or more). However, the method of the first aspect of the present invention requires that separate multiplex reactions be performed on multiple (i.e., at least two) aliquots of a sample. As used herein, the term "multiplex" refers to an assay in which multiple (i.e., at least two) different analytes are analyzed simultaneously, more specifically, in the same aliquot of sample or in the same reaction mixture. Thus, it is clear that the number of analytes sought to be detected according to the method of the first aspect of the present invention is a minimum of four (two analytes to be detected in each of two aliquots of sample). However, it is preferred that significantly more than four analytes be detected according to the present method. Preferably, at least 10, 20, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, or 1500 or more analytes are detected according to the present methods.
[0019] The terms "detecting" or "detected" are used broadly herein to encompass any means of determining the presence or absence of an analyte (i.e., determining whether a target analyte is present in a sample of interest). Thus, even if the methods of the invention are performed to attempt to detect a particular analyte of interest in a sample, and the analyte is not detected because it is not present in the sample, the step of "detecting the analyte" has still occurred, since the presence or absence of the analyte in the sample has been assessed. The step of "detecting" the analyte does not depend on the detection being successful, i.e., the analyte actually being detected.
[0020] Detecting the analyte may further include, in any form, measuring the concentration or abundance of the analyte in the sample. The absolute concentration of the target analyte may be determined, or the relative concentration of the analyte may be determined for purposes of comparing the concentration of the target analyte to the concentrations of other target analytes (or other target analytes) in the sample or other samples.
[0021] Thus, "detecting" may include determining, measuring, assessing, or analyzing the presence or absence or amount of an analyte in any manner. Quantitative and qualitative determinations, measurements, or assessments are included, including semi-quantitative determinations. Such determinations, measurements, or assessments may be relative or absolute, for example, when two or more different analytes are being detected in a sample. Thus, when used in the context of quantifying a target analyte in a sample, the term "quantifying" can refer to absolute or relative quantification. Absolute quantification may be achieved by including one or more control analytes of known concentration and / or by comparing the detected level of the target analyte to a known control analyte (e.g., by generating a standard curve). Alternatively, relative quantification may be achieved by comparing the detected levels or amounts of two or more different target analytes to provide a relative quantification of each of the two or more different target analytes, i.e., quantification relative to one another. Methods by which quantification can be achieved in the methods of the present invention are further described below.
[0022] The methods of the present invention are for detecting multiple analytes in a sample. Any sample of interest can be analyzed according to the present invention, i.e., any sample that contains or may contain an analyte of interest, and that one desires to analyze to determine whether it contains the analyte of interest and / or to determine the concentration of the analyte of interest in the sample.
[0023] Thus, any biological or clinical sample may be analyzed according to the present invention, for example, any cell or tissue sample from or derived from an organism, or any body fluid or preparation thereof, as well as samples such as cell cultures, cell preparations, cell lysates, etc. Environmental samples such as soil and water samples, or food samples may also be analyzed according to the present invention. Samples may be freshly prepared or may have been pre-treated in any convenient way, for example for storage.
[0024] Representative samples thus include any material that may contain biomolecules or other desired or target analytes, including, for example, food and related products, clinical samples, and environmental samples. The sample may be a biological sample and may contain viral or cellular material, including prokaryotic or eukaryotic cells, viruses, bacteriophage, mycoplasma, protoplasts, organelles, etc. Such biological material may thus include all types of mammalian and / or non-mammalian animal cells, plant cells, algae, including blue-green algae, fungi, bacteria, protozoa, etc.
[0025] The sample is preferably a clinical sample, such as whole blood, blood-derived products such as plasma, serum, buffy coat, and blood cells, urine, feces, cerebrospinal fluid or other bodily fluids (e.g., respiratory secretions, saliva, milk, etc.), tissue, biopsy, etc. The sample is particularly preferably a plasma or serum sample. Thus, the methods of the present invention may be used, for example, to detect biomarkers or to analyze samples for pathogen-derived analytes. The sample may be particularly derived from humans, although the methods of the present invention may be equally applicable to samples derived from non-human animals (i.e., veterinary samples). The sample may be pretreated and prepared for use in the methods of the present invention by any convenient or desired method, such as cell lysis or cell removal.
[0026] The method of the first aspect of the present invention is for detecting multiple analytes in a sample that are present at different levels of abundance in the sample. That is, the analytes are present in the sample at different concentrations or over a range of concentrations. Every analyte in a sample need not be present at a concentration that is substantially different from every other analyte, but rather not all analytes are present at substantially the same concentration. The analytes in a sample may be present at a range of concentrations, although certain analytes may be present at very similar concentrations.
[0027] Analytes may be present in a sample over a concentration range spanning several orders of magnitude. For example, the analyte present (or expected to be present) in the sample at the highest concentration may be present (or expected to be present) at a concentration that is about 1000 times greater than the concentration of the analyte present (or expected to be present) in the sample at the lowest concentration. For example, analytes in a sample may differ in concentration by about 10-fold, about 100-fold, about 1000-fold, or more compared to one another, and of course, the factor can be any value in between. In a clinical sample, analytes may be present over a range of several orders of magnitude, for example, three, four, five, or six or more orders of magnitude.
[0028] The abundance level or value used to block or group different analytes, or more specifically, assays for different analytes, together may not depend solely on the absolute level or concentration of the analyte present (or expected to be present) in the sample. Other factors may be taken into account, including the nature of the assay method, differences in analytical performance for different analytes, etc. For example, in the case of detection assays based on antibodies or other binding agents, the abundance may depend on the affinity or avidity of the antibody for the analyte. Such variations between assays for different analytes may be taken into account. For example, abundance may reflect the abundance of the analyte detected in the assay, converted into an assay output or measurement. Thus, the predicted abundance based on which analytes are selected for the subset may depend at least on the predicted level or predicted concentration of the analyte in the sample, but may also, or instead, depend on the predicted level or predicted value of abundance determined in a particular detection assay. In other words, the abundance of an analyte in a sample may be an apparent abundance or a theoretical abundance that depends on the detection assay. The apparent abundance of an analyte may vary depending on the assay used, and particularly on the sensitivity of the assay.
[0029] The method includes preparing multiple (i.e., at least two) aliquots of the sample. That is, separate portions of the sample are prepared. The sample may be divided into multiple aliquots (so that the entire sample is dispensed), or portions of the sample may be prepared as aliquots without using the entire sample. The aliquots may be the same size or volume, or may be different sizes or volumes, or some aliquots may be the same size and others may be different sizes.
[0030] At least some of the aliquots may be diluted. For example, the sample may be diluted 1:2, 1:4, 1:5, 1:10, etc. In particular, the aliquots may be diluted in 10-fold increments. That is, one or more aliquots may be diluted 10-fold (1:10), one or more aliquots may be diluted 100-fold (1:100), and one or more aliquots may be diluted 1000-fold (1:1000). If necessary, further dilutions (e.g., 1:10,000 or 1:100,000) may be performed, although it is generally expected that a dilution of up to 1:1000 will be sufficient. One or more aliquots may not be diluted (referred to herein as 1:1).
[0031] In certain embodiments, serial 10-fold dilutions may be performed to provide 1:1, 1:10, 1:100, and 1:1000 diluted aliquots. In this embodiment, the 1:10 dilution is performed by diluting the undiluted sample 10-fold. The 1:100 and 1:1000 dilutions may be performed by directly diluting the undiluted sample 100-fold and 1000-fold (respectively) or by serially diluting the 1:10 diluted aliquot 10-fold (i.e., diluting the 1:10 diluted aliquot 10-fold to obtain the 1:100 diluted aliquot, and diluting the 1:100 diluted aliquot 10-fold to obtain the 1:1000 diluted aliquot). Sample dilutions (and, indeed, all of the dispensing steps throughout the methods of the present invention) may be performed manually or using an automated dispensing robot (e.g., SPT LabTech Mosquito).
[0032] The sample may be diluted using any suitable diluent, depending on the type of sample being analyzed. For example, the diluent may be water, saline, or a buffer solution, particularly a buffer solution containing a biologically compatible buffer compound (i.e., a buffer compatible with the detection assay being used, e.g., a buffer compatible with PEA or PLA). Suitable buffer compounds include, for example, HEPES, Tris (i.e., tris(hydroxymethyl)aminomethane), disodium phosphate, etc. Buffers suitable for use as diluents include PBS (phosphate-buffered saline), TBS (Tris-buffered saline), HBS (HEPES-buffered saline), etc. The buffer (or other diluent) used must be made with a purified solvent (e.g., water) so as to be free of contaminant analytes. Therefore, the diluent should be sterilized. When water is used as a diluent or as a base for a diluent, the water used is preferably ultrapure water (e.g., Milli-Q water).
[0033] Any suitable number of aliquots may be provided from the sample. As noted above, at least two aliquots are provided, but in most embodiments, more than two will be provided. In certain embodiments, as detailed above, four aliquots may be provided: an aliquot of the undiluted sample and aliquots of the sample diluted 1:10, 1:100, and 1:1000. More or fewer aliquots may be provided if more or less dilution of the sample is desired. Furthermore, more than one aliquot at each dilution may be provided, depending on the needs / requirements of the particular assay being performed.
[0034] When multiple aliquots are provided from a sample, a separate multiplex assay is performed on each aliquot to detect a subset of target analytes in each aliquot. A separate multiplex assay is performed on each aliquot so that each aliquot is analyzed separately (i.e., the multiple aliquots are not mixed during the multiplex reaction). When the multiplex assay is performed across all prepared aliquots, all target analytes are detected. That is, assays are performed across all aliquots to determine the presence or absence of each target analyte in the sample of interest. However, an individual assay for detecting a particular analyte may be performed on only one aliquot. Thus, a different subset of analytes is detected in each aliquot. In other words, different analytes are detected in each aliquot. Preferably, the subsets detected in each aliquot are distinct. That is, each target analyte is detected in only one aliquot so that there is no overlap between the analyte subsets. However, in some embodiments, a particular analyte may be detected in multiple aliquots, if deemed appropriate. In this example, there will be some overlap of analytes between subsets, in that some analytes will be present in multiple analyte subsets, while other analytes will only be present in one subset.
[0035] The analytes in each subset are selected based on their expected abundance (i.e., concentration) in the sample. That is, analytes that are likely to be present in the sample at similar concentrations may be included in the same subset and analyzed in the same multiplex reaction. Conversely, analytes that are likely to be present in the sample at different concentrations may be included in different subsets and analyzed in different multiplex reactions. Each analyte is assigned to a subset of analytes that are likely to be present in the sample at similar concentrations (e.g., within a certain order of magnitude). Each analyte subset is then detected in a sample aliquot diluted by an appropriate factor, taking into account the likely concentration of the analyte. Thus, the analyte likely to be present at the lowest concentration may be detected in either the undiluted or the less diluted aliquot, the analyte likely to be present at the highest concentration may be detected in the most diluted aliquot, and the analyte likely to be present at a concentration between these extremes may be detected in the "intermediate" dilution aliquot.
[0036] As noted above, in some embodiments, a particular analyte may be included in more than one subset. This may be the case, for example, when the likely concentration of the analyte is essentially between the likely concentrations of two subsets and does not clearly "belong" to either of them. In this example, the analyte may be included in both subsets. If it is known that the analyte may be present in a sample over a significantly wide concentration range, the analyte may be included in two (or more) subsets.
[0037] It will be appreciated that, provided that the analytes in each subset are selected based on their expected abundance in the sample, the number of analytes in each subset may be different, or the number of analytes in each subset may be the same, as appropriate.
[0038] The abundance / concentration of each analyte in a sample may be predicted based on known facts about the normal levels of each analyte in the sample type being analyzed. For example, if the sample is a plasma or serum sample (or other bodily fluid sample), the concentration of the analyte therein may be predicted based on known concentrations of species in these bodily fluids. Normal plasma concentrations for a wide range of analytes of potential interest are available at https: / / www.olink.com / resources-support / document-download-center / . However, as noted above, the abundance values used to assign analytes to particular subsets (blocks) may depend on the assay and the results (e.g., measurements) obtained from that assay.
[0039] As detailed above, a multiplex reaction is performed on each aliquot to detect all analytes of the subset to be analyzed in the aliquot. As mentioned above, the term "multiplex" refers to an assay in which at least two different analytes are analyzed simultaneously. However, preferably, significantly more than two analytes are analyzed in each multiplex reaction. For example, each multiplex reaction may analyze at least 5, 10, 15, 20, 25, 30, 40, 50, 60, or more analytes. In certain multiplex reactions, more than this number of analytes may be analyzed, for example, at least 70, 80, 90, 100, 110, 120, 130, 140, 150, or more analytes.
[0040] In certain embodiments of this aspect of the invention, the analytes are detected in each aliquot by detecting a reporter nucleic acid molecule specific for each analyte. In this embodiment, the presence of a particular analyte in the sample results in the production of a nucleic acid molecule having a specific nucleotide sequence known to correspond to the particular analyte during the detection assay. Detection of a particular nucleotide sequence indicates the presence of the analyte to which that sequence corresponds in the sample. Thus, a "reporter nucleic acid molecule" is a nucleic acid molecule whose synthesis during the detection assay indicates the presence of a particular analyte in the sample. The reporter nucleic acid molecule may be an RNA molecule or a DNA molecule. Preferably, it is a DNA molecule.
[0041] Reporter nucleic acid molecules may be generated by any means known in the art for detection assays. For example, they may be generated by ligating two (or more) nucleic acids together to form a unique nucleotide sequence that indicates the presence of the analyte in a sample. Alternatively, reporter nucleic acid molecules may be generated by extending a provided nucleic acid molecule along a template nucleic acid molecule. A combination of extension and ligation may also be used.
[0042] Thus, reporter nucleic acid molecules are generated while performing a multiplex detection assay on each aliquot. Any detection assay that operates by generating such nucleic acid molecules may be used to generate reporter nucleic acid molecules. In a specific embodiment, reporter nucleic acid molecules are generated in a proximity extension assay (PEA). That is, multiple PEA may be performed to detect analytes in each aliquot, and thus the sample. In another embodiment, reporter nucleic acid molecules are generated in a proximity ligation assay (PLA). That is, multiple PLA may be performed to detect analytes in each aliquot. As mentioned above, methods for performing PEA and PLA are known in the art. It is particularly preferred that the detection assay performed is PEA.
[0043] After production, the reporter nucleic acid molecule is preferably amplified to facilitate detection. The reporter nucleic acid molecule is preferably amplified by PCR, but other nucleic acid amplification methods, such as loop-mediated isothermal amplification (LAMP), may also be used.
[0044] As mentioned above, each reporter nucleic acid molecule is specific to a particular analyte. Thus, a reporter nucleic acid molecule identifies a given analyte, and more specifically, may contain a sequence or domain that serves as an identification (ID) sequence or tag that can detect the analyte. The ID sequence may be detected, for example, by serving as a binding site for a probe or primer, as described in further detail below, or more directly by sequencing. Therefore, in other words, this specificity may be achieved by the presence of one or more barcode sequences in the reporter nucleic acid molecule. Generally, a barcode sequence can be defined as a nucleotide sequence within the reporter nucleic acid molecule that identifies the reporter and, therefore, the detected analyte. The entirety of each reporter nucleic acid molecule generated in a detection assay may be unique, in which case the entire reporter nucleic acid molecule may be considered a barcode sequence. More generally, one or more smaller portions of the reporter nucleic acid molecule serve as barcode sequences.
[0045] Analytes in a sample are detected by detecting specific barcode sequences in reporter nucleic acid molecules generated during a multiplex detection assay. This can be achieved in several ways. First, specific barcode sequences may be detected by sequencing all reporter nucleic acid molecules generated during a multiplex detection assay. By sequencing all of the reporter nucleic acid molecules generated, all of the different reporter nucleic acid molecules generated may be identified by their barcode sequences, thereby identifying all analytes present in the sample (this identification is based on whether or not a reporter nucleic acid molecule known to correspond to each target analyte is detected). Nucleic acid sequencing is a preferred method for detecting / analyzing reporter nucleic acids.
[0046] Other suitable methods for detecting reporter nucleic acid molecules include PCR-based methods. For example, quantitative PCR using "TaqMan" probes may be performed. In this example, reporter nucleic acid molecules (or at least a portion of each reporter nucleic acid molecule containing a barcode sequence) are amplified, and a probe complementary to each barcode sequence is provided, with each different probe being conjugated to a different distinguishable fluorescent substance. The presence or absence of each barcode (and thus the reporter nucleic acid molecule and, ultimately, the analyte) can then be determined based on whether a particular barcode is amplified. However, while PCR-based methods such as those described above are clearly only suitable for simultaneously analyzing a relatively small number of different sequences, combinatorial methods using probes to decode barcode sequences are known and may be used to expand multiplexing capacity to some extent. Because nucleic acid sequencing has virtually no limit on the number of sequences that can be identified in a single attempt and allows for a higher level of multiplexing than detection using PCR, sequencing is a preferred method for detecting reporter nucleic acid molecules.
[0047] Preferably, reporter nucleic acid molecules are detected using a form of high-throughput DNA sequencing. Sequencing by synthesis is a preferred DNA sequencing method. Sequencing by synthesis techniques include, for example, pyrosequencing, reversible dye terminator sequencing, and ion torrent sequencing, all of which can be used in the present method. Preferably, reporter nucleic acids are sequenced using massively parallel DNA sequencing. Massively parallel DNA sequencing is particularly applicable to sequencing by synthesis (for example, the above-mentioned reversible dye terminator sequencing, pyrosequencing, or ion torrent sequencing). Massively parallel DNA sequencing using reversible dye terminator sequencing is a preferred sequencing method. Massively parallel DNA sequencing using reversible dye terminator sequencing can be performed, for example, using an Illumina® NovaSeq™ system.
[0048] As known in the art, massively parallel DNA sequencing is a technique for sequencing multiple (for example, thousands, millions, or more) DNA strands in parallel, i.e., simultaneously.In massively parallel DNA sequencing, target DNA molecules need to be immobilized on a solid surface, for example, the surface of a flow cell or beads.Then, each immobilized DNA molecule is individually sequenced.Usually, the massively parallel DNA sequencing method that employs reversible dye terminator sequencing uses a flow cell as immobilization surface, while the massively parallel DNA sequencing method that employs pyrosequencing or ion torrent sequencing uses beads as immobilization surface.
[0049] As known to those skilled in the art, immobilization of DNA molecules on a surface in massively parallel sequencing methods is usually achieved by adding one or more sequencing adapters to the ends of the molecules. Thus, the method of the present invention may include adding one or more sequencing adapters to a reporter nucleic acid molecule.
[0050] Typically, the sequencing adapter is a nucleic acid molecule (particularly a DNA molecule). In this example, a short oligonucleotide complementary to the adapter sequence is attached to an immobilization surface (e.g., the surface of a bead or flow cell) to allow annealing of the target DNA molecule to the surface via the adapter sequence. Alternatively, the target DNA molecule may be attached to the immobilization surface using other binding partner pairs, such as biotin and avidin / streptavidin. In this case, biotin may be used as the sequencing adapter, and the biotin sequencing adapter may be bound using avidin or streptavidin attached to the immobilization surface, or vice versa.
[0051] Thus, the sequencing adapter may be a short oligonucleotide (preferably DNA), typically 10-30 nucleotides in length (e.g., 15-25 nucleotides, or 20-25 nucleotides in length). As detailed above, the purpose of the sequencing adapter is to enable annealing of the target DNA molecule to the immobilization surface, and therefore the nucleotide sequence of the nucleic acid adapter is determined by the sequence of the binding partner bound to the immobilization surface. Other than this, there are no particular restrictions on the nucleotide sequence of the nucleic acid sequencing adapter.
[0052] Sequencing adapters may be added to the reporter nucleic acid molecules of the present invention during PCR amplification. In the case of nucleic acid sequencing adapters, this can be achieved by including sequencing adapter nucleotides in one or both primers. Alternatively, if the sequencing adapter is a non-nucleic acid sequencing adapter (e.g., a protein / peptide or small molecule), the adapter may be attached to one or both PCR primers. Alternatively, the sequencing adapter may be added to the reporter nucleic acid molecule by directly ligating or attaching the sequencing adapter to the reporter nucleic acid molecule. Preferably, one or more sequencing adapters used in the present method are nucleic acid sequencing adapters.
[0053] One or more nucleic acid sequencing adaptors may be added to a reporter molecule in one or more ligation and / or amplification steps. Thus, for example, when two sequencing adaptors are added to a reporter nucleic acid molecule (one at each end), they may be added in one step (e.g., by PCR amplification using a pair of primers, both of which contain a sequencing adaptor) or in two steps. These two steps may be performed in the same way or different ways. For example, a first sequencing adaptor may be added to a reporter nucleic acid molecule by ligation, and a second sequencing adaptor may be added by PCR amplification, or vice versa. Alternatively, a first amplification reaction may be performed to add a first sequencing adaptor to a reporter nucleic acid molecule, and then a second amplification reaction may be performed to add a second sequencing adaptor to a reporter nucleic acid molecule.
[0054] As mentioned above, one or more sequencing adapters may be added to a reporter nucleic acid molecule. This means one or two sequencing adapters. Since sequencing adapters are added to the ends of DNA molecules, the maximum number of sequencing adapters that can be added to one DNA molecule (e.g., reporter nucleic acid) is two. Thus, one sequencing adapter may be added to one end of a reporter nucleic acid molecule, or two sequencing adapters may be added, one to each end of a reporter nucleic acid molecule. In certain embodiments, Illumina P5 adapters and Illumina P7 adapters are used. That is, a P5 adapter is added to one end of a reporter nucleic acid molecule, and a P7 adapter is added to the other end. The sequence of the P5 adapter is shown in SEQ ID NO: 1 (AAT GAT ACG GCG ACC ACC GA), and the sequence of the P7 adapter is shown in SEQ ID NO: 2 (CAA GCA GAA GAC GGC ATA CGA GAT).
[0055] Thus, in certain embodiments of the invention, the reporter nucleic acid molecule is subjected to at least a first (i.e., at least one) PCR amplification to add at least a first (i.e., at least one) sequencing adaptor to the reporter nucleic acid molecule. As noted above, the reporter nucleic acid molecule is produced during a detection reaction in response to the presence of the target analyte to which the reporter nucleic acid molecule corresponds (i.e., the analyte whose presence is indicated by the production of the reporter nucleic acid molecule). As further noted above, the reporter nucleic acid molecule is preferably amplified to enable or improve its detection.
[0056] Therefore, this amplification may be combined with the addition of one or more sequencing adaptors to the reporter nucleic acid molecule. This may be achieved by amplifying the reporter nucleic acid molecule using a primer pair containing at least one sequencing adaptor. In this example, at least one primer of the primer pair contains a sequencing adaptor upstream of the sequence that binds to the reporter nucleic acid molecule. Thus, the sequencing adaptor is usually located at the 5' end of either primer that contains it.
[0057] In certain embodiments, the amplification step is carried out using a primer pair in which one primer contains a sequencing adapter, such that one sequencing adapter is added to one end of the reporter nucleic acid molecule.
[0058] In another embodiment, the amplification step is carried out using a primer pair in which both primers contain a sequencing adapter, such that a sequencing adapter is added to each end of the reporter nucleic acid molecule in a single amplification step.
[0059] In another embodiment, sequencing adaptors are added to each end of the reporter nucleic acid molecule by performing two separate amplification reactions, with each amplification step adding a different sequencing adaptor to a different end of the molecule.
[0060] In another embodiment, an initial amplification step is performed using primers that do not contain sequencing adapters, and the amplified reporter nucleic acid molecule is then subjected to one or more additional amplification reactions, as described above, to add sequencing adapters to each end of the molecule.
[0061] As detailed above, each reporter nucleic acid molecule generated during a detection assay may contain a barcode sequence corresponding to a specific analyte. Thus, reporter nucleic acid molecules with different sequences are generated in response to the presence of different analytes in a sample. Nevertheless, to facilitate multiplexing, it is preferred that all reporter nucleic acid molecules generated during a detection assay share a common primer binding site so that the same primer pair can be used to amplify all different reporter nucleic acid molecules.
[0062] When a first PCR amplification is performed on a reporter nucleic acid molecule, in which only one sequencing adapter is added to the molecule, a second PCR amplification can be performed on the amplified reporter nucleic acid molecule (i.e., the product of the first PCR amplification) to add a second sequencing adapter. Thus, in this embodiment, the first PCR amplification is performed using a primer pair in which one primer contains a sequencing adapter, thereby adding a first sequencing adapter to one end of the reporter nucleic acid molecule. A second PCR amplification is then performed using a different primer pair. In the second primer pair, one primer contains a second sequencing adapter. The second sequencing adapter is different from the first sequencing adapter, i.e., has a different sequence. The primer containing the second sequencing adapter binds to the reporter nucleic acid molecule at the end opposite to the end containing the first sequencing adapter, so that the second sequencing adapter is added to the reporter nucleic acid molecule at the end opposite to the end containing the first sequencing adapter.
[0063] If necessary to amplify the product of the first PCR amplification, the second primer of the second primer pair may contain the sequence of the first sequencing adapter so that it can bind to the end of the reporter nucleic acid molecule to which the first sequencing adapter has been added during the first PCR amplification. In certain embodiments, the primer containing the first sequencing adapter used in the first PCR amplification to add the first sequencing adapter to the reporter nucleic acid molecule is also used in the second PCR amplification. That is, the same primer (containing the first sequencing adapter) may be used in the first PCR amplification and the second PCR amplification.
[0064] In embodiments in which two PCR amplifications are performed sequentially to add sequencing adapters to both ends of the reporter nucleic acid molecule, the product of the first PCR may be purified before performing the second PCR. Standard methods for purifying PCR products are known in the art.
[0065] As mentioned above, Illumina P5 sequencing adaptors and Illumina P7 sequencing adaptors are preferred sequencing adaptor pairs for use in the present invention. In a specific embodiment, the P5 sequencing adaptor is added to the reporter nucleic acid molecule in a first PCR amplification, and the P7 sequencing adaptor is added to the reporter nucleic acid molecule in a second PCR amplification. In another embodiment, the P7 sequencing adaptor is added to the reporter nucleic acid molecule in a first PCR amplification, and the P5 sequencing adaptor is added to the reporter nucleic acid molecule in a second PCR amplification.
[0066] At least one of the one or two PCR amplifications performed to add a sequencing adaptor to a reporter nucleic acid molecule is preferably performed to saturation. As is well known in the art, the amount of PCR amplification product versus the number of cycles follows an "S" curve. Initially, the amplicon concentration gradually increases, then reaches an exponential amplification phase, during which the amount of product (approximately) doubles with each amplification cycle. After the exponential phase, a linear phase is reached, in which the amount of product increases linearly, rather than exponentially. Finally, a plateau is reached, in which the amount of product reaches the maximum possible level, determined by the reaction settings, the concentrations of the components used, etc.
[0067] In the present invention, saturation PCR may generally be considered to be PCR that has passed the exponential phase, i.e., PCR that is in the linear phase or has reached a plateau. In certain embodiments, "saturation" as used herein means that the reaction is run until the maximum possible product is obtained (i.e., until the product amount reaches a plateau), so that further amplification cycles do not produce any more product. Saturation can be reached when reaction components are depleted, for example, when primers or dNTPs are depleted. When reaction components are depleted, the reaction rate slows and then enters a plateau state. Although uncommon, saturation can also be reached when polymerase is exhausted (i.e., when the polymerase loses its activity). Saturation can also be reached when the amplicon concentration reaches a high level such that the DNA polymerase concentration is no longer sufficient to maintain exponential amplification, i.e., when there are more amplicon molecules than polymerase molecules. In this example, amplification enters and remains in the linear phase as long as sufficient primers and dNTPs remain in the reaction mix.
[0068] In certain embodiments, two PCR amplifications are performed to add sequencing adaptors to reporter nucleic acid molecules, and both of these reactions are performed to saturation. In another embodiment, only the first of the two PCR amplifications is performed to saturation. Alternatively, only the second of the two PCR amplifications is performed to saturation. It is particularly preferred that only the first of the two PCR amplifications is performed to saturation.
[0069] PCR amplification may be carried out to saturation simply by performing many cycles of PCR amplification, so that saturation can be assumed. For example, PCR amplification carried out for at least 25, 30, 35, or more cycles can be assumed to have reached saturation by its endpoint, in that the exponential amplification phase has ceased by that stage. Alternatively, saturation can be measured by quantitative PCR (qPCR). For example, TaqMan PCR can be performed using a probe that binds to a sequence common to all reporter nucleic acid molecules, or qPCR can be performed using a dye, such as SYBR Green, that changes color upon binding to double-stranded DNA. In this way, the reaction can be tracked to determine the minimum number of amplification cycles required to reach saturation. In either case, if further processing of the amplified reporter nucleic acid molecules (before sequencing) is required, any such experimental qPCR will need to be performed on an aliquot separate from the aliquot used in the experiment to generate the reporter nucleic acid molecules for sequencing, in order to identify the saturation point. This is because TaqMan probes or intercalating dyes are likely to interfere with further steps of the method.
[0070] As detailed above, a separate multiplex reaction is performed for each aliquot of the sample of interest. Each aliquot is used to detect analytes present at different levels in the sample. Reporter nucleic acid molecules are initially generated in amounts corresponding to the amount of each analyte in the sample. Thus, for analytes present at high concentrations, high concentrations of reporter nucleic acid molecules can be expected to be generated, and for analytes present at low concentrations, low concentrations of reporter nucleic acid molecules can be expected. The amount of reporter nucleic acid molecules generated can be expected to be proportional to the amount of the corresponding analyte present in the sample. For example, for a first analyte present in the sample at a concentration 10 times higher than that of a second analyte, 10 times as many reporter nucleic acid molecules for the first analyte as for the second analyte can be expected to be generated. Thus, a much larger amount of reporter nucleic acid molecules will be generated in an aliquot used to detect an analyte expected to be present at a high concentration in the sample than in an aliquot used to detect an analyte expected to be present at a low concentration in the sample.
[0071] If this difference in reporter nucleic acid abundance is carried over to the analytical step (e.g., sequencing step) that identifies the reporter nucleic acid molecules, the most abundant reporter nucleic acid molecule may "sweep out" the signal of the reporter nucleic acid molecule present in lower amounts, resulting in insufficient detection of the analyte present in low amounts in the sample.
[0072] In PCR carried out until saturation, the reporter nucleic acid molecules from each multiplex reaction are amplified, eliminating differences in reporter nucleic acid concentration between aliquots. Once saturation is reached, each aliquot will contain essentially the same amount of reporter nucleic acid molecules. This means that for each analyte present in the sample, there will be a similar amount of reporter nucleic acid molecules, and therefore, when analyzing the reporter nucleic acid molecules, all of the reporter nucleic acid molecules (and therefore their corresponding analytes) should be detected.
[0073] As described above, the multiplex detection assays used in the present methods are performed separately on multiple aliquots of the sample of interest. The products of the multiplex detection assays are then used to identify which of the target analytes are present in the sample. As detailed above, this may be achieved using reporter nucleic acid molecules corresponding to different analytes and analyzed, for example, by sequencing, to determine which reporter nucleic acid molecules are present (and thus which analytes are present in the sample). It is possible to analyze each multiplex reaction performed on each sample aliquot separately. However, in a preferred embodiment of the present invention, the reaction products from each aliquot (i.e., the products of the multiplex detection assay) are pooled (i.e., mixed). In other words, such a pooling step can be considered to pool separate "abundance-specific blocks." This allows for more efficient analysis of the reaction products by enabling a single analytical reaction (e.g., a sequencing reaction) for all aliquots of the sample.
[0074] When the products of the multiplex detection assay are reporter nucleic acid molecules, it is preferable to first amplify the reporter nucleic acid molecules (e.g., by PCR) and pool the amplified products, optionally followed by a further amplification step. It is particularly preferred to perform a separate first PCR amplification, as described above, on the reporter nucleic acid molecules generated by each separate multiplex detection assay, in which a first sequencing adapter is added to the nucleic acid molecule, and then pool the products. In other words, a detection assay is performed on each separate aliquot to generate a reporter nucleic acid molecule, and a first PCR reaction is performed on the reporter nucleic acid molecule, which both amplifies the reporter nucleic acid molecule and adds a first sequencing adapter to one end of the reporter nucleic acid molecule. The products of this first amplification reaction are pooled. If necessary, the products of each separate first PCR reaction may be purified before pooling. Alternatively, the products of the separate first PCR reactions may be pooled, and then all PCR products in the pool may be purified together. However, purifying the products of the first PCR amplification before proceeding to the second PCR amplification is not a requirement.
[0075] After pooling, the pooled products of the first PCR amplification are subjected to a second PCR amplification. The second PCR is used to both amplify the products of the first PCR and add a second sequencing adaptor to the reporter nucleic acid molecule, as described above. When pooling the products of the first PCR, it is important to perform the first PCR to saturation so that the amplified reporter nucleic acid molecule is present in approximately the same amount in each aliquot when pooled. It is not important whether the second PCR performed on the pooled products of the first PCR amplification is also performed to saturation, but this may be done if necessary. In a preferred embodiment, both the first PCR amplification and the second PCR amplification are performed to saturation.
[0076] In another embodiment, each of the separate aliquots is subjected to a separate multiplex detection assay.The reporter nucleic acid molecules produced in each aliquot are then subjected to a single PCR reaction, which is carried out separately for each aliquot, until saturation, and a sequencing adaptor is added to each end of the reporter nucleic acid molecule (one sequencing adaptor is added to each end of each reporter nucleic acid molecule).The products of this PCR reaction are then pooled and sequenced.
[0077] In yet another embodiment, a separate multiplex detection assay is performed on each separate aliquot. The reporter nucleic acid molecules produced in each aliquot are then subjected to two PCR amplifications. Both of these are performed separately for each aliquot. A first PCR is used to add a first sequencing adapter to the reporter nucleic acid molecule, and a second PCR is used to add a second sequencing adapter to the reporter nucleic acid molecule (the end of the reporter nucleic acid molecule opposite the first sequencing adapter). The products of the second PCR are then pooled and sequenced. In this embodiment, it is important that at least one of the PCR amplifications is performed to saturation for each aliquot. As long as the same reaction is performed to saturation for each aliquot, either the first PCR or the second PCR, or both PCRs, can be performed to saturation.
[0078] When amplified reporter nucleic acid molecules from separate multiplex reactions are pooled, the amount of amplification product from each separate multiplex reaction added to the pool can be the same or different. The same amount of amplification product from each separate multiplex reaction can be added to the pool. This can be achieved by adding the complete amplification reaction mixture from each multiplex reaction to the pool, or by adding the same defined amount of each amplification reaction mixture to the pool. In this example, for example, if three aliquots are prepared from a sample, each of which is subjected to a separate multiplex detection assay, and the amplified reporter nucleic acid molecules from each aliquot are pooled, one-third of the pool will come from each aliquot. Similarly, if four aliquots are prepared from a sample, one-quarter of the pool will come from each aliquot.
[0079] Alternatively, different amounts of amplification product from each separate multiplex reaction may be added to the pool. "Different amounts of amplification product" simply means that the amount of amplification product added to the pool is not the same across all aliquots / multiplex detection assays. Thus, this may be the case where different amounts of amplification product from each multiplex detection assay are added to the pool, or the same amount of amplification product from some but not all aliquots may be added, such that different amounts of amplification product are added from some aliquots. For example, if three aliquots are prepared from a sample, each of which is subjected to a separate multiplex detection assay, and the amplified reporter nucleic acid molecules from each aliquot are pooled, different amounts of amplification product from all three aliquots may be added to the pool. Alternatively, the same amount of amplification product from two aliquots may be added to the pool, and a different amount of amplification product from the third aliquot may be added. Similarly, for example, if four aliquots are prepared from a sample, different amounts of amplification product from all four aliquots may be added to the pool. Alternatively, the same amount of amplification product from three aliquots may be added to the pool, and a different amount of amplification product from the fourth aliquot. If the same amount of amplification product from two aliquots is added to the pool, different amounts of amplification product from the other two aliquots may be added to the pool, or the same amount (first amount) of amplification product from two aliquots may be added to the pool, and the same amount (second amount) different from the first amount from the other two aliquots may be added to the pool.
[0080] When different amounts of amplification product are added to a pool from various aliquots, the amount added from each aliquot is preferably proportional to the number of analytes detected in each aliquot. Thus, for example, if twice as many analytes are detected in a first aliquot as in a second aliquot, twice as much of the first aliquot as the second aliquot is added to the pool. This can be considered as adding the same amount of amplification product to the pool for each analyte detected in the sample across all aliquots. For example, if 100 analytes are detected across three aliquots, 50 in the first aliquot, 30 in the second aliquot, and 20 in the third aliquot, the three aliquots will be added to the pool in a ratio of 5:3:2, such that 50% of the pool comes from the first aliquot, 30% from the second aliquot, and 20% from the third aliquot.
[0081] The method of the first aspect of the present invention can be used to analyze multiple samples in parallel.When analyzing multiple samples in parallel, the samples can be the same type or different types.Preferably, all samples are the same type, for example, all are plasma samples or all are saliva samples.The analyte set detected in each sample can also be the same or different.Preferably, the same analyte set is detected in each sample, and each specific analyte in all samples is identified using the same reporter nucleic acid molecule.Analyzing multiple samples in parallel means analyzing multiple samples simultaneously, while each step of the method is carried out for each sample essentially simultaneously.
[0082] When analyzing multiple samples in parallel, multiple aliquots are prepared from each sample, as described above, with each aliquot detecting a subset of analytes. Preferably, the same number of aliquots are prepared from each sample. For example, three aliquots may be prepared from each sample, or four aliquots may be prepared from each sample. However, this is not required, and different numbers of aliquots may be prepared from different samples, such as two aliquots from some samples, three aliquots from other samples, four aliquots from other samples, and five aliquots from still other samples.
[0083] As noted above, it is preferred that the same set of analytes be detected in each sample and that the same number of aliquots be prepared from each sample. It is even more preferred that the analytes be divided into aliquots in the same manner in each sample, so that the same subset of analytes is detected in each corresponding sample aliquot (i.e., aliquots from each sample at the same dilution).
[0084] When analyzing multiple samples in parallel using the method of the first aspect of the present invention, reporter nucleic acid molecules may be amplified as described above, and the amplification products for each specific sample may be pooled as described above to generate a first pool. Thus, a separate first pool may be generated for each sample, and each first pool contains the amplification products from all multiplex detection assays performed on that sample (i.e., the amplification products from all aliquots prepared for that sample).
[0085] In one embodiment, the separate first pools generated for each sample may be further pooled to facilitate subsequent analysis. In such an embodiment, after the first pooling step, a sample index is added to the amplification products of each first pool. The sample index is a nucleotide sequence that identifies the original sample from which the amplification product was derived. Thus, a different nucleotide sequence is used as the sample index sequence for the amplification products derived from each sample. When the amplification products are subsequently sequenced, the sample index indicates which sample each individual reporter nucleic acid molecule originated from. Any nucleotide sequence may be used as the sample index. The sample index sequence may be of any length, but is preferably relatively short, e.g., 3-12 nucleotides, 4-10 nucleotides, or 4-8 nucleotides.
[0086] Thus, a different sample index sequence is used to label the amplification products in each separate first pool. However, the sample index sequence is the same within each individual first pool. The sample index sequence may be added to the amplification product by any appropriate method; for example, the sample index may be added during the amplification reaction (e.g., by PCR) or during the ligation reaction. In particular, if the amplified reporter nucleic acid molecule is to be analyzed by massively parallel DNA sequencing and requires sequencing adapters at both ends, the sample index sequence cannot be added so that it is ultimately located at the end of the reporter nucleic acid molecule.
[0087] As described above, it is preferred to perform a first PCR amplification of the reporter nucleic acid molecules, including adding a first sequencing adaptor to the reporter molecules, and then pool the reporter nucleic acid molecules to create a first pool. This also applies when analyzing multiple samples in parallel. As described above, it is preferred to perform a first PCR amplification separately for each aliquot of each sample, and add a first sequencing adaptor to one end of the reporter nucleic acid molecule. As described above, it is preferred to pool the aliquots of each sample separately to obtain a separate first pool for each sample.
[0088] Once the distinct first pools are obtained, a sample index is added. As described above, this may be achieved by amplification or ligation. Regardless of how the sample index is added, it is added to the end of the reporter nucleic acid molecule opposite the end bearing the first sequencing adaptor. While a ligation step may be performed to add the sample index to the end of each reporter nucleic acid molecule, preferably, the addition of the sample index is achieved by amplification, typically by PCR. The sample index is added during amplification using a primer pair, one of the primers containing the sample index sequence, so that the sample index is incorporated into the amplification product.
[0089] The addition of the sample index may be performed in a dedicated amplification step solely for the purpose of adding the sample index to the reporter nucleic acid molecule. An additional amplification step may then be performed, if necessary, to add a second sequencing adaptor to the reporter nucleic acid molecule. In this example, the second sequencing adaptor is added to the reporter nucleic acid molecule at the same end where the sample index is located. This typically results in the sample index being located internally to, and immediately adjacent to, the second sequencing adaptor in the amplified and adaptor-labeled reporter nucleic acid molecule.
[0090] However, preferably, as detailed above, after pooling the products of the first PCR amplification to obtain a first pool, the first pool (i.e., the products of the first PCR amplification) is subjected to a second PCR amplification that adds both a sample index and a second sequencing adaptor to the reporter nucleic acid molecule. Thus, a separate second PCR amplification is performed for each first pool. That is, a separate second PCR is performed for each analyzed sample.
[0091] In this embodiment, the second PCR amplification is performed using a primer pair in which one primer contains both the sample index sequence and the second sequencing adaptor, so that both are simultaneously added to the reporter nucleic acid molecule. The primer containing the second sequencing adaptor and the sample index sequence has the second sequencing adaptor at its 5' end. The sample index sequence is located downstream, usually immediately downstream, of the second sequencing adaptor so that it is adjacent to the second sequencing adaptor, but it does not have to be adjacent. Thus, the product of the second PCR amplification contains two sequencing adaptors (one at each end) and a sample index that is internal to the second sequencing adaptor.
[0092] The second PCR may use a common first primer and a unique second primer that varies across the analyzed samples. In other words, one primer (the same primer) is used across all samples to bind to the end of the reporter nucleic acid molecule to which the first sequencing adapter was added in the first PCR amplification. A different second primer is used for each sample, where the second primer contains a sample index sequence that is unique to each sample.
[0093] After the second PCR amplification, the indexed first pools generated for each sample are pooled (i.e., added or mixed together) to create a second pool. This second pool is used for DNA sequencing. Thus, a single DNA sequencing reaction can identify the reporter nucleic acid molecules generated for each sample. The sample index attached to the reporter nucleic acid molecule allows the sample from which each nucleic acid molecule originates to be identified, thereby determining which analytes are present in each sample. Prior to DNA sequencing, the second PCR amplification product is preferably purified to remove excess primers and other residues remaining from the amplification reaction. This purification step may be performed regardless of whether one sample or multiple samples are analyzed in the method. When multiple samples are analyzed and the products of the second PCR amplification are pooled prior to sequencing, the products of the second PCR may be purified either before or after pooling. That is, a second PCR may be performed on each first pool, the products pooled to generate a second pool, and then the PCR products of the second pools purified together in a single purification reaction. Alternatively, a second PCR may be performed on each first pool, the products from each second PCR may be purified separately, and then the purified products of the second PCR amplifications may be pooled.
[0094] As described above, each reporter nucleic acid molecule contains at least one barcode sequence associated with a specific analyte. Thus, each specific reporter nucleic acid molecule is detected by detecting its barcode sequence, typically by sequencing. When analyzing a single sample using the method of the first aspect of the present invention, detecting all reporter nucleic acid molecules generated in a multiplex detection assay requires only detecting their barcodes. Detection of each specific barcode indicates the presence of its corresponding analyte in the sample. When analyzing multiple samples in parallel using this method, after amplification, each reporter nucleic acid molecule contains both a barcode sequence and a sample index. In this embodiment, detecting each reporter nucleic acid molecule involves detecting both the barcode sequence and the sample index. Detection of the sample index indicates which sample the reporter nucleic acid molecule originates from, and detection of the barcode indicates the presence of a specific analyte in that sample. Thus, detecting the reporter nucleic acid molecule allows for the identification of the analyte present in each analyzed sample.
[0095] As mentioned above, sequencing for this method is typically performed by massively parallel DNA sequencing. For this purpose, the purified product of the second PCR amplification (or an aliquot thereof) is denatured, for example, using sodium hydroxide to obtain single-stranded DNA molecules. The denatured (single-stranded) DNA may be diluted with an appropriate buffer, if necessary. An appropriate dilution buffer is generally provided with the DNA sequencing platform or by the manufacturer of the DNA sequencing platform. The denatured DNA is then placed on a solid support (e.g., a bead or a flow cell) by hybridizing its sequencing adapter to a complementary sequence protruding from the solid support. Once the DNA is placed on the solid support, DNA sequencing is performed using a selected method.
[0096] The above-described method allows for the detection of each analyte in a sample. The method also allows for the comparison of analyte levels within each subset for each sample. That is, it allows for the comparison of analyte levels in each specific sample aliquot analyzed. In each individual aliquot, the level of each different reporter nucleic acid molecule produced is proportional to the level of that analyte (e.g., if a first analyte is present at twice the level in a particular aliquot as in a second aliquot, then twice as many reporter nucleic acid molecules corresponding to the first analyte will be produced as reporter nucleic acid molecules corresponding to the second analyte). These differences in reporter levels are detected during reporter detection, e.g., sequencing, thereby allowing for the comparison of the relative amounts of analytes present in the samples, but only for analytes detected in the same aliquot.
[0097] It would be advantageous to be able to compare the relative amounts of all analytes present in a sample (i.e., to compare the analytes detected in different aliquots). It would be even more advantageous to be able to compare the relative amounts of analytes present in different samples. This can be achieved by including an internal control in each aliquot. The same internal control is included in each aliquot of each sample. The internal control is included in each aliquot of sample at a different concentration depending on the dilution factor of the aliquot. The concentration of the internal control is proportional to the dilution factor of the aliquot. Thus, for example, if an internal control is used at a particular given concentration in an aliquot of an undiluted sample, then an aliquot of a 1:10 diluted sample would use the internal control at one-tenth the concentration used in the undiluted sample, and so on. This allows for direct comparison of the relative concentrations of analytes between aliquots while ensuring that the signal of the internal control does not overwhelm or be obscured by the signal of the analyte detected in the aliquot. This is because an internal control is present in each aliquot at a concentration appropriate to the analyte detected in that aliquot.
[0098] An internal control is a control reporter nucleic acid molecule or one that results in the production of a control reporter nucleic acid molecule. By comparing the amount of each reporter nucleic acid molecule to the control reporter, the relative amounts of analyte analyzed in different aliquots and / or the relative amounts of analyte from different samples can be compared. This is possible because the relative difference between each reporter nucleic acid molecule and the control reporter can be compared.
[0099] For example, if two different reporter nucleic acid molecules from different samples are present at the same relative level to a control reporter (e.g., half or third, or two-fold or three-fold), this indicates that the analyte represented by the two reporter nucleic acid molecules is present at essentially the same concentration in the two samples. Similarly, if the ratio of a particular reporter nucleic acid molecule to a control reporter is twice the ratio of the same reporter nucleic acid molecule from a different sample to the control reporter (e.g., if the reporter molecule is present in a first sample at twice the level of the control reporter and in a second sample at essentially the same level as the control reporter), this indicates that the analyte represented by the particular reporter nucleic acid molecule is present in the first sample at approximately twice the level present in the second sample.
[0100] There are various options that can be used as an internal control. The appropriate control may depend on the detection technology used. In any detection assay, the internal control may be an added analyte, i.e., a control analyte added at a predetermined concentration to each aliquot to be analyzed. The control analyte is added to an aliquot before the multiplex detection assay and is detected in each aliquot, similar to other analytes in the sample. In particular, detection of the control analyte may lead to the generation of a control reporter nucleic acid molecule specific to the control analyte, as described above. When a control analyte is used, the control analyte is an analyte that may not be present in the sample of interest. For example, it may be an artificial analyte, or, if the sample is derived from an animal (e.g., human), the control analyte may be a biomolecule from a different species that is not present in the animal of interest. In particular, the control analyte may be a non-human protein. Examples of control analytes include fluorescent proteins such as green fluorescent protein (GFP), yellow fluorescent protein (YFP), and cyan fluorescent protein (CFP).
[0101] Another example of an internal control is a double-stranded DNA molecule that has the same overall structure as a reporter nucleic acid molecule generated in a multiplex detection assay. That is, the DNA molecule contains a barcode sequence that identifies the DNA molecule as a control reporter nucleic acid molecule and a common primer binding site that is shared by all other reporter nucleic acids generated in response to analyte detection and that allows binding of primers used in the amplification reaction. Notably, the control DNA molecule does not contain a sequencing adapter or a sample index. These sequencing adapters and sample indexes are added to the control DNA molecule at the same time as they are added to the reporter nucleic acid molecules generated in response to analyte detection (e.g., in PCR amplification), as described above.
[0102] The double-stranded DNA molecule used as a control in this manner is referred to herein as a detection control. This is because it is not only useful for assessing the analyte concentration (by comparing the concentration with that of a control, as described above), but also allows for confirmation that the reporter nucleic acid molecule produced during analyte detection is amplified, labeled, and detected (e.g., by sequencing), as described above. If the detection control is not detected when the reporter nucleic acid molecule is analyzed (e.g., sequenced), this indicates a failure of the detection method. For example, the amplification step may have failed, or the sequencing reaction may have failed. The detection control is preferably added to each aliquot before performing the multiplex detection assay.
[0103] In certain embodiments of the method, both a control analyte and a detection control are added to each aliquot. In this example, the barcode sequence for the control analyte is distinct from the barcode sequence for the detection control, such that the two internal controls can be individually identified.
[0104] As mentioned above, the multiplexed detection assay is preferably a multiplexed proximity extension assay or a multiplexed proximity ligation assay, most preferably a multiplexed proximity extension assay, which are briefly described above. As mentioned above, both of these techniques rely on the use of paired proximity probes.
[0105] A proximity probe is defined herein as an entity comprising an analyte-binding domain specific for the analyte and a nucleic acid domain. "Specific for the analyte" means that the analyte-binding domain specifically recognizes and binds to a particular target analyte, i.e., binds to the target analyte with higher affinity than it binds to other analytes or other moieties. The analyte-binding domain is preferably an antibody, particularly a monoclonal antibody. Antibody fragments or antibody derivatives containing the antigen-binding domain are also suitable for use as the analyte-binding domain. Such antibody fragments or derivatives include, for example, molecules such as Fab, Fab', F(ab')2, and scFv.
[0106] The Fab fragment consists of the antigen-binding domain of an antibody. An individual antibody may be considered to contain two Fab fragments, each consisting of a light chain and the N-terminal portion of the heavy chain to which it is attached. Thus, a Fab fragment contains an intact light chain and the V of the heavy chain to which it is attached. H Domain and C H 1 domain. Fab fragments may be obtained by digesting antibodies with papain.
[0107] An F(ab')2 fragment consists of two Fab fragments of an antibody and the hinge region of the heavy domain, containing a disulfide bond connecting the two heavy chains. In other words, an F(ab')2 fragment can be considered to be two Fab fragments covalently linked together. An F(ab')2 fragment can be obtained by digesting an antibody with pepsin. Reduction of the F(ab')2 fragment yields two Fab' fragments, which can be considered Fab fragments containing additional sulfhydryl groups that may be useful for conjugating the fragment to other molecules. An ScFv molecule is a synthetic construct produced by fusing the light and heavy chain variable domains of an antibody. Typically, this fusion is achieved recombinantly by engineering antibody genes to produce a fusion protein containing both the heavy and light chain variable domains.
[0108] The nucleic acid domain of a proximity probe may be a DNA domain or an RNA domain. Preferably, it is a DNA domain. The nucleic acid domains of the proximity probes in each proximity probe pair are typically designed to hybridize with each other or with one or more common oligonucleotide molecules (both nucleic acid domains of a pair of proximity probes may hybridize). Therefore, the nucleic acid domain must be at least partially single-stranded. In one embodiment, the nucleic acid domain of a proximity probe is entirely single-stranded. In another embodiment, the nucleic acid domain of a proximity probe is partially single-stranded, containing both single-stranded and double-stranded portions.
[0109] The proximity probes are typically provided as a proximity probe pair, each specific for a target analyte. As mentioned above, the target analyte may be a single entity, particularly a single protein. In this embodiment, the probes of the proximity pair both bind to the target analyte (e.g., a protein), but to different epitopes. The epitopes are non-overlapping, such that binding of one probe of the proximity probe pair to an epitope does not interfere with or inhibit binding of the other probe of the proximity probe pair to the epitope. Alternatively, as mentioned above, the target analyte may be a complex, such as a protein complex, where one probe of the proximity probe pair binds to one component of the complex and the other probe of the proximity probe pair binds to the other component of the complex. The probes bind to proteins contained in the complex at sites that are different from the protein interaction sites (i.e., the sites of the proteins that interact with each other).
[0110] As mentioned above, the proximity probes are provided as proximity probe pairs, each specific to a target analyte. This means that in each proximity probe pair, both probes contain analyte-binding domains specific to the same analyte. Since the detection assays used are multiplex assays, multiple different probe pairs are used in each detection assay, each probe pair being specific to a different analyte. That is, the analyte-binding domains of each of the different probe pairs are specific to different target analytes.
[0111] The nucleic acid domain of each proximity probe is designed according to the method in which the probe will be used. Representative examples of proximity extension assay formats are shown schematically in Figure 1, and these embodiments are described in detail below. Typically, in a proximity extension assay, when a pair of proximity probes binds to a target analyte, the nucleic acid domains of the two probes come into proximity with each other and interact (i.e., hybridize to each other directly or indirectly). The interaction between the two nucleic acid domains results in a nucleic acid duplex containing at least one free 3' end (i.e., at least one of the nucleic acid domains in the duplex has an extendable 3' end). Addition of a nucleic acid polymerase enzyme to the assay mix or activation of a nucleic acid polymerase enzyme in the assay mix extends the at least one free 3' end. Thus, at least one of the nucleic acid domains in the duplex is extended using its paired nucleic acid domain as a template. The resulting extension product, a reporter nucleic acid molecule as used herein, contains a barcode sequence that indicates the presence of the analyte bound by the proximity probe pair that produced the extension product.
[0112] Version 1 of Figure 1 shows a "traditional" proximity extension assay, in which the nucleic acid domain of each proximity probe (indicated by an arrow) is attached at its 5' end to an analyte-binding domain (indicated by an upside-down "Y"), leaving two free 3' ends. When the proximity probes bind to their respective analytes (analytes not shown in the figure), the nucleic acid domains of the probes that are complementary at their 3' ends can hybridize and interact, i.e., form a duplex. Addition of a nucleic acid polymerase enzyme to the assay mixture, or activation of a nucleic acid polymerase enzyme within the assay mixture, allows each nucleic acid domain to be extended using the nucleic acid domain of the other proximity probe as a template. As detailed above, the resulting extension product is a reporter nucleic acid molecule that is detected to identify the analyte to which the probe pair is bound.
[0113] Version 2 of Figure 1 shows an alternative proximity extension assay in which the nucleic acid domain of a first proximity probe is attached at its 5' end to an analyte-binding domain, and the nucleic acid domain of a second proximity probe is attached at its 3' end to an analyte-binding domain. This results in the nucleic acid domain of the second proximity probe having a free 5' end (indicated by a blunt arrow) that cannot be extended using typical nucleic acid polymerase enzymes (which only extend 3' ends). The 3' end of the second proximity probe is effectively "blocked" - that is, it is not "free" and cannot be extended because it is bound to and blocked by the analyte-binding domain. In this embodiment, when the proximity probes bind to their respective analyte-binding targets on the analyte, the nucleic acid domains of the probes that share a region of complementarity at their 3' ends can hybridize and interact, i.e., form a duplex. However, in contrast to version 1, only the nucleic acid domain of the first proximity probe (with a free 3' end) can be extended using the nucleic acid domain of the second proximity probe as a template to obtain an extension product (i.e., a reporter nucleic acid molecule).
[0114] In version 3 of Figure 1, as in version 2, the nucleic acid domain of the first proximity probe is attached at its 5' end to the analyte-binding domain, and the nucleic acid domain of the second proximity probe is attached at its 3' end to the analyte-binding domain. This results in the nucleic acid domain of the second proximity probe having a free 5' end (indicated by the blunt arrow) that cannot be extended. However, in this embodiment, the nucleic acid domain attached to the analyte-binding domain of each proximity probe does not have a region of complementarity and therefore cannot directly form a duplex. Instead, a third nucleic acid molecule is provided that has a region of homology to the nucleic acid domain of each proximity probe. This third nucleic acid molecule acts as a "molecular bridge" or "splint" between the nucleic acid domains. This "splint" oligonucleotide bridges the gap between the nucleic acid domains, allowing them to interact indirectly. That is, each nucleic acid domain forms a duplex with a splint oligonucleotide.
[0115] Thus, when a proximity probe binds to each analyte-binding target on the analyte, each of the probe's nucleic acid domains interacts with the splint oligonucleotide by hybridizing to it, i.e., forming a duplex with the splint oligonucleotide. Thus, the third nucleic acid molecule, or splint, can be considered to be the second strand of a partially double-stranded nucleic acid domain provided on one of the proximity probes. For example, one of the proximity probes may be provided with a partially double-stranded nucleic acid domain, which is attached to the analyte-binding domain via the 3'-end of one strand, and the other (unattached) strand has a free 3'-end. Thus, such a nucleic acid domain has a terminal single-stranded region that includes a free 3'-end. In this embodiment, the nucleic acid domain of the first proximity probe (with the free 3'-end) may be extended using the "splint oligonucleotide" (or the single-stranded 3'-terminal region of the other nucleic acid domain) as a template. Alternatively, or in addition, the free 3' end of the splint oligonucleotide (i.e. the unattached strand, or 3' single-stranded region) may be extended using the nucleic acid domain of the first proximity probe as a template.
[0116] As is clear from the above description, in one embodiment, the splint oligonucleotide may be provided as a separate component of the assay. In other words, the splint oligonucleotide may be added separately to the reaction mix (i.e., added to the sample containing the analyte separately from the proximity probe). Nevertheless, because it hybridizes to a nucleic acid molecule that is part of the proximity probe and will hybridize upon contact with such a nucleic acid molecule, it can still be considered a strand of a partially double-stranded nucleic acid domain, even if added separately. Alternatively, the splint may be pre-hybridized to one of the nucleic acid domains of the proximity probe, i.e., hybridized before the proximity probe is contacted with the sample. In this embodiment, the splint oligonucleotide can be considered directly to be part of the nucleic acid domain of the proximity probe. That is, a nucleic acid domain is a nucleic acid molecule that is partially double-stranded; for example, a proximity probe may be created by linking a double-stranded nucleic acid molecule to an analyte binding domain (preferably the nucleic acid domain is single-stranded and attached to the analyte binding domain) and modifying the nucleic acid molecule to generate a partially double-stranded nucleic acid domain (with single-stranded overhangs that can hybridize to the nucleic acid domain of another proximity probe).
[0117] Therefore, the extension of the nucleic acid domain of a proximity probe as defined herein also encompasses the extension of a "splint" oligonucleotide. Advantageously, when the extension product is generated by extension of a splint oligonucleotide, the resulting extended nucleic acid strand is linked to the proximity probe pair only by the interaction between the two strands of the nucleic acid molecule (by the hybridization of the two nucleic acid strands). Therefore, in these embodiments, the extension product can be dissociated from the proximity probe pair using denaturing conditions, such as increasing the temperature or decreasing the salt concentration.
[0118] Although the splint oligonucleotide depicted in Version 3 of Figure 1 is shown as being complementary to the entire length of the nucleic acid domain of the second proximity probe, this is by way of example only, and the splint may be any oligonucleotide that is capable of forming a duplex with (or near) the end of the nucleic acid domain of the proximity probe, i.e., that forms a bridge between the nucleic acid domains of the two probes.
[0119] In another embodiment, the splint oligonucleotide may be provided as the nucleic acid domain of a third proximity probe, as described in WO2007 / 107743, which is incorporated herein by reference, and it has been demonstrated that this can further improve the sensitivity and specificity of proximity probe assays.
[0120] Version 4 of Figure 1 is a variation of Version 1, in which the nucleic acid domain of a first proximity probe contains a sequence at its 3' end that is not perfectly complementary to the nucleic acid domain of a second proximity probe. Thus, when the proximity probes bind to their respective analytes, the nucleic acid domains of the probes are able to hybridize and interact, i.e., form a duplex, but the extreme 3' end of the nucleic acid domain of the first proximity probe (the portion of the nucleic acid molecule containing the free 3' hydroxyl group) is unable to hybridize to the nucleic acid domain of the second proximity probe and therefore exists as a single-stranded, unhybridized "flap." Upon addition or activation of a nucleic acid polymerase enzyme, only the nucleic acid domain of the second proximity probe can be extended using the nucleic acid domain of the first proximity probe as a template.
[0121] Version 5 of Figure 1 may be considered a variant of Version 3. However, in contrast to Version 3, the nucleic acid domains of both proximity probes are attached at their 5' ends to their respective analyte-binding domains. In this embodiment, the 3' ends of the nucleic acid domains are not complementary, and therefore the nucleic acid domains of the proximity probes cannot interact or directly form duplexes. Instead, a third nucleic acid molecule is provided that has regions homologous to the nucleic acid domains of each proximity probe. This third nucleic acid molecule acts as a "molecular bridge" or "splint" between the nucleic acid domains. This "splint" oligonucleotide bridges the gap between the nucleic acid domains, allowing them to indirectly interact. That is, each nucleic acid domain forms a duplex with a splint oligonucleotide. Thus, when the proximity probes bind to their respective analytes, each of the nucleic acid domains of the probes hybridizes and interacts with the splint oligonucleotide, i.e., forms a duplex with the splint oligonucleotide.
[0122] According to Version 3, the third nucleic acid molecule, or splint, can be considered to be the second strand of a partially double-stranded nucleic acid domain provided on one of the proximity probes. In a preferred example, one of the proximity probes may be provided with a partially double-stranded nucleic acid domain, which is attached to the analyte binding domain via the 5' end of one strand, and the other (unattached) strand has a free 3' end. Thus, such a nucleic acid domain has a terminal single-stranded region containing at least one free 3' end. In this embodiment, the nucleic acid domain of the second proximity probe (having a free 3' end) may be extended using the "splint oligonucleotide" as a template. Alternatively, or in addition, the free 3' end of the splint oligonucleotide (i.e., the unattached strand, or the 3' single-stranded region of the first proximity probe) may be extended using the nucleic acid domain of the second proximity probe as a template.
[0123] As described above in relation to Version 3, the splint oligonucleotide may be provided as a separate element of the assay. However, because it hybridizes to a nucleic acid molecule that is part of a proximity probe and will hybridize upon contact with such a nucleic acid molecule, even when added separately, it can still be considered to be a strand of a partially double-stranded nucleic acid domain. Alternatively, the splint may be pre-hybridized to one of the nucleic acid domains of the proximity probe, i.e., before contacting the proximity probe with the sample. In this embodiment, the splint oligonucleotide can be considered to be directly part of the nucleic acid domain of the proximity probe. That is, the nucleic acid domain is a partially double-stranded nucleic acid molecule; for example, a proximity probe may be created by linking a double-stranded nucleic acid molecule to an analyte-binding domain (preferably the nucleic acid domain is single-stranded and bound to the analyte-binding domain) and modifying the nucleic acid molecule to generate a partially double-stranded nucleic acid domain (with a single-stranded overhang that can hybridize to the nucleic acid domain of another proximity probe).
[0124] Therefore, the extension of the nucleic acid domain of a proximity probe as defined herein also encompasses the extension of a "splint" oligonucleotide. Advantageously, when the extension product is generated by extension of a splint oligonucleotide, the resulting extended nucleic acid strand is linked to the proximity probe pair only by the interaction between the two strands of the nucleic acid molecule (by the hybridization of the two nucleic acid strands). Therefore, in these embodiments, the extension product can be dissociated from the proximity probe pair using denaturing conditions, such as increasing the temperature or decreasing the salt concentration.
[0125] Although the splint oligonucleotide depicted in Version 5 of Figure 1 is shown as being complementary to the entire length of the nucleic acid domain of the first proximity probe, this is merely an example and the splint may be any oligonucleotide that is capable of forming a duplex with (or near) the end of the nucleic acid domain of the proximity probe, i.e., that forms a bridge between the nucleic acid domains of the proximity probe.
[0126] In another embodiment, the splint oligonucleotide may be provided as the nucleic acid domain of a third proximity probe, as described in WO2007 / 107743, which is incorporated herein by reference, and it has been demonstrated that this can further improve the sensitivity and specificity of proximity probe assays.
[0127] Version 6 of Figure 1 is the most preferred embodiment of the present invention. As shown, both probes of a probe pair are bound to a nucleic acid molecule that is partially single-stranded. A short nucleic acid strand is bound to an analyte-binding domain via its 5' end. The short nucleic acid strands bound to the analyte-binding domain are not hybridized to each other. Rather, each of the short nucleic acid strands hybridizes to a long nucleic acid strand, which has a single-stranded overhang at its 3' end (i.e., the 3' end of the long nucleic acid strand extends beyond the 5' end of the short strand attached to the analyte-binding domain). The overhangs of the two long nucleic acid strands hybridize to each other to form a duplex. When the 3' ends of the two long nucleic acid molecules are fully hybridized to each other, the duplex contains two free 3' ends, as shown; however, the 3' ends of the long nucleic acid molecules may be designed, as in version 4, such that the extreme 3' end of one of the long nucleic acid molecules is not complementary to the other and forms a flap, i.e., the duplex contains only one free 3' end. Two long nucleic acid molecules that interact with each other can be considered splint oligonucleotides in that they together form a bridge between the two short oligonucleotides that are directly attached to the analyte-binding domains.
[0128] Addition or activation of a nucleic acid polymerase extends the free 3' ends of one or both splint oligonucleotides. Specifically, either splint oligonucleotide extends using the other splint oligonucleotide as a template. Thus, as one splint oligonucleotide extends, the other "template" splint oligonucleotide displaces the short strand bound to the analyte-binding domain.
[0129] In a preferred embodiment, the short nucleic acid strand directly attached to the analyte-binding domain is a "common strand." That is, the same strand is directly attached to all proximity probes used in a multiplexed detection assay. Thus, each splint oligonucleotide contains a "common portion" consisting of a sequence that hybridizes to the common strand and a "unique portion" that contains a barcode sequence unique to the probe. Such proximity probes and methods for making them are described in WO2017 / 068116.
[0130] In all proximity detection assay technologies, the nucleic acid domain of each individual proximity probe preferably contains a unique barcode sequence that identifies that particular probe (as described above for PEA version 6). In this case, the reporter nucleic acid molecule (or, in a proximity extension assay, the extension product) contains the unique barcode sequence of each proximity probe. These two unique barcode sequences thus together form the barcode sequence of the reporter nucleic acid molecule. In other words, the barcode sequence of the reporter nucleic acid molecule is or comprises a combination of the two probe barcode sequences, and the barcode sequences of the proximity probes are combined to generate the reporter nucleic acid molecule. Thus, a specific reporter nucleic acid molecule is detected by detecting the specific combination of the two probe barcode sequences.
[0131] When using a multiplex proximity extension assay to detect analytes, it is preferable to use an additional internal control, an extension control. The extension control is a single probe containing an analyte-binding domain bound to a nucleic acid domain containing a duplex with an extendable free 3' end. Preferably, the extension control has a structure essentially equivalent to the duplex formed between two experimental probes upon binding to the target analyte, except that it contains only one analyte-binding domain. The analyte-binding domain used in the extension control does not recognize analytes that may be present in the sample of interest. Suitable analyte-binding domains are commercially available polyclonal isotype control antibodies, such as goat IgG, mouse IgG, or rabbit IgG.
[0132] Figure 2 shows examples of extension controls that can be used in the present invention. Parts A through F correspond to extension controls that can be used in versions 1 through 6 of the PEA assay in Figure 1, respectively. Extension controls are used to verify that the extension step is occurring as intended. Extension of the extension control generates a reporter nucleic acid molecule that contains a unique barcode so that it can be identified as an extension control reporter nucleic acid molecule. When multiplexed PEA is used in the method of the first aspect of the present invention, it is preferred that a control analyte, extension control, and detection control are all used in the assay (i.e., added to each aliquot). In other embodiments, only two internal controls are used, such as a control analyte and an extension control, a control analyte and a detection control, or an extension control and a detection control.
[0133] As detailed above, in a proximity extension assay, a reporter nucleic acid molecule is generated by extending the nucleic acid domain of one or both proximity probes using the nucleic acid domain of the other proximity probe as a template. In a preferred embodiment, the extension reaction is performed during PCR amplification. In other words, a single reaction including PCR amplification is performed to both extend the nucleic acid domain of the proximity probe to generate a reporter nucleic acid molecule and amplify the reporter molecule, including adding a first sequence adapter to the generated reporter nucleic acid molecule. In this embodiment, the reaction does not start with a denaturation step (as is common in PCR), but rather with an extension step to generate the reporter nucleic acid molecule. A conventional PCR is then performed, beginning with reporter molecule denaturation, to amplify the reporter nucleic acid molecule. As detailed above, PCR is performed using common primers that bind to common sequences at the ends of the reporter nucleic acid molecule, one of the primers containing a sequencing adapter. Alternatively, PCR may be performed using primers that all contain a sequencing adapter to add a sequencing adapter to each end of the reporter nucleic acid molecule in a single run, as detailed above.
[0134] It may be desirable to detect more analytes in a sample than there are different reporter nucleic acid barcode sequences available. In this case, multiple panels of proximity probes (i.e., at least two panels) may be used. Each panel contains a different set of proximity probe pairs; i.e., the proximity probe pairs in each panel bind to a different set of analytes. Typically, the proximity probe pairs in each panel bind to completely different sets of analytes; i.e., the analytes bound by the proximity probe pairs in different panels do not overlap. Thus, each panel of proximity probes is intended to detect a different group of analytes.
[0135] As described above, each panel of proximity probes contains a different set of proximity probe pairs. In each individual panel, each probe contains a different nucleic acid domain (i.e., each probe contains a nucleic acid domain with a different sequence). Thus, each probe pair contains a different pair of nucleic acid domains, and therefore a unique reporter nucleic acid molecule is generated for each probe pair within a panel. However, in each different panel, the same nucleic acid domains (and usually the same nucleic acid domain pairings) are used in the probe pairs. That is, in different panels, the probe pairs contain the same pair of nucleic acid domains. This means that the same reporter nucleic acid molecule is generated in each panel. However, because the reporter nucleic acid molecule is generated by each panel using different probe pairs, the same reporter nucleic acid molecule indicates the presence of different analytes in each panel of probes. Because each panel of probes generates the same reporter nucleic acid molecule, a separate sample aliquot must be prepared for a multiplex detection assay using each panel of probes. That is, a multiplex detection assay is performed using each panel of probes, and multiplex detection assays using different panels of probes are performed in different aliquots of the sample. As detailed above for a single probe panel, for each panel of probes, multiple sample aliquots are prepared at different dilutions, and a different analyte subset for each panel is detected in each aliquot. As detailed above, the analyte subset detected in each aliquot is determined based on the expected concentration in the sample.
[0136] The reporter nucleic acid molecules generated using each distinct probe panel are processed (i.e., amplified, optionally labeled with a sequencing adapter, etc.) and detected as described above. In certain embodiments, as described above, the reporter nucleic acid molecules are amplified by PCR, sequencing adapters are added to both ends of the reporter nucleic acid molecule, and a sample index is added to each reporter nucleic acid molecule. In this embodiment, as described above, a first PCR is preferably performed separately for each aliquot to amplify the reporter nucleic acid molecule and add a first sequencing adapter to one end of the reporter nucleic acid molecule. Then, as described above, the amplified reporter nucleic acid molecules from each sample generated using a particular probe panel are pooled to generate multiple separate first pools. Each separate first pool contains the products of the first PCR amplification performed on all aliquots of a particular sample analyzed using a particular probe panel.
[0137] Then, each of the first separate pools is subjected to a second PCR amplification, and a second sequencing adaptor and a sample index are added to each reporter nucleic acid molecule.After the second PCR, the PCR products generated from different samples using the same probe panel are pooled into a second pool, known as a panel pool.The entire first pools can be combined into a panel pool, or only a portion of each first pool can be combined.Therefore, each panel pool contains the reporter nucleic acid molecules generated from all analyzed samples using a specific probe panel.
[0138] The amplified reporter nucleic acid molecules containing the sequencing adaptors and sample indexes are then sequenced as described above. Each panel pool is sequenced separately. This is because, as described above, the same reporter nucleic acid molecules are generated in each probe panel, but these reporter nucleic acid molecules represent different analytes in each probe panel. It is impossible to distinguish identical reporter nucleic acid molecules generated using different probe panels and therefore representing different analytes at the sequence level. Therefore, in this embodiment, it is essential to sequence each panel pool separately.
[0139] In another embodiment of this method, a panel index sequence is added to the reporter nucleic acid molecule during one PCR amplification. The same panel index sequence is used to identify all reporter nucleic acid molecules (across all samples) generated using a particular proximity probe panel. The combination of the panel index and sample index allows for accurate identification of which analytes are present in each sample across all probe panels used in the detection assay. Thus, once both a panel index and a sample index are added to each reporter nucleic acid molecule, all PCR products generated in the detection assay across all samples and probe panels can be pooled and sequenced together.
[0140] Alternatively, reporter nucleic acid molecules generated using each probe panel may be labeled with a different sample index. Different sample index sequences are selected and used for each different sample, so that each sample index used is unique to that particular sample. However, for any given sample, reporter nucleic acid molecules generated using each different probe panel are labeled with different sample indexes. Thus, in this embodiment, the sample index serves the dual function of identifying both the sample and the probe panel for each reporter nucleic acid molecule. Thus, the specific sample index present in a reporter nucleic acid molecule links that reporter to a specific probe panel, and the combination of the sample index and barcode sequence of a reporter nucleic acid molecule serves to identify the analyte that led to the generation of that reporter nucleic acid molecule.
[0141] In a further embodiment of this method, as detailed above, the same nucleic acid domains are used in the probes in each probe panel. However, each panel contains different pairs of nucleic acid domains so that each panel generates a different reporter nucleic acid molecule. As described above, the nucleic acid domains of each probe contain a unique barcode sequence. By pairing the nucleic acid domains differently in each panel, different combinations of barcode sequences are paired in the reporter nucleic acid molecules generated in the detection assay. This means that different reporter nucleic acid molecules are generated for each panel. This method has the advantage that, because different reporter nucleic acid molecules are generated by each probe panel, they can be distinguished at the sequence level without the need for any panel index sequence. In this embodiment, as detailed above, all PCR products from each sample are pooled and assigned a sample index, and then all indexed PCR products from all samples and probe panels are combined into a single pool and sequenced.
[0142] As mentioned above, the advantage of this embodiment is that all reporter nucleic acid molecules from all samples and panels can be pooled and sequenced together without the need for a panel index to identify which reporter nucleic acid molecule comes from each panel. However, the advantage of using probe pairs with the same nucleic acid domain pairs for each panel so that each panel produces the same reporter nucleic acid molecule is that any nucleic acid molecules resulting from the hybridization of two unpaired nucleic acid molecules can be identified as non-specific background. If each probe panel produces a different reporter nucleic acid molecule, it is no longer possible to accurately determine which of the produced nucleic acid molecules is background.
[0143] As mentioned above, in a second aspect, the present invention provides a method for detecting an analyte in a sample, wherein the analyte is detected by detecting a reporter nucleic acid molecule specific for the analyte, the method comprising: performing a PCR reaction to generate a PCR product of the reporter nucleic acid molecule; and detecting the PCR product; An internal control was prepared for the PCR reaction, and the internal control was (i) a separate component that is, contains, or produces a control nucleic acid molecule that is present in a predetermined amount and is amplified with the same primers as the reporter nucleic acid molecule; and / or (ii) a unique molecular identifier (UMI) sequence present in each reporter nucleic acid molecule, which is unique to each molecule.
[0144] All details of this second aspect of the invention may be the same as those of the first aspect (e.g., analyte, sample, reporter nucleic acid molecule and techniques used to generate same, detection of reporter nucleic acid molecule, etc.).
[0145] In this second embodiment, the internal control is a component or sequence present in the PCR performed to generate a PCR product of the reporter nucleic acid molecule. As noted above, the internal control can be a control nucleic acid molecule that is present in a predetermined amount and is amplified with the same primers as the reporter nucleic acid molecule, or a separate component that contains or generates such a nucleic acid molecule.
[0146] When the internal control is a separate component present in the reaction in a predetermined amount, the internal control may be, among other things, a control analyte, an extension control, or a detection control, as described above. As detailed above, a control analyte is an analyte that is added to a sample and detected by detecting a control reporter nucleic acid molecule specific for the control analyte.
[0147] In the method of the second aspect of the invention, the analyte is preferably detected using proximity probes, for example in PEA or PLA as detailed above, most preferably PEA. Thus, if a control analyte is used as an internal control, a proximity probe for detecting the control analyte must be included. Binding of the control-specific proximity probe to the control analyte generates a control reporter nucleic acid molecule.
[0148] As described above, an extension control may be used. As detailed above, an extension control is a single control probe that generates a control reporter nucleic acid molecule during the extension step of the PEA.
[0149] Generally speaking, an internal control can be one or more molecules that are added to a sample to generate a control reporter nucleic acid molecule that is subsequently amplified in a PCR reaction.
[0150] Also, as described above, a detection control may be used. As detailed above, the detection control is a control reporter nucleic acid molecule that is added to the sample and amplified in a PCR reaction. The detection control is a double-stranded DNA molecule that has the same overall structure as the reporter nucleic acid molecule that is generated in response to the presence of the analyte. As in the first aspect of the present invention, it is preferred that a control analyte, an extension control, and a detection control are all used in this method. In certain embodiments, two types of internal controls may be used, the options for which are described above.
[0151] As detailed above, the control analyte, extension control, and detection control all generate or are control reporter nucleic acid molecules. In certain embodiments of the present invention, the control reporter nucleic acid molecule has a sequence that is the reverse sequence of the reporter nucleic acid molecule generated in response to the detection of an analyte. In particular, the control reporter nucleic acid molecule has the reverse sequence of the reporter nucleic acid molecule generated in response to the detection of an analyte, but does not have a reverse complementary sequence. Because the control reporter nucleic acid molecule simply has the reverse sequence of the reporter nucleic acid molecule generated in response to the detection of an analyte, the control reporter nucleic acid molecule cannot hybridize to the reporter nucleic acid molecule of interest. This allows for the maintenance of maximum similarity between the control reporter nucleic acid molecule and the reverse-sequence reporter nucleic acid molecule generated in response to the detection of an analyte, while preventing unwanted hybridization interactions between the control reporter nucleic acid molecule and the reporter nucleic acid molecule generated in response to the detection of an analyte, which is an advantage in PCR amplification. A control reporter nucleic acid molecule, the sequence of which is the reverse sequence of the reporter nucleic acid molecule produced in response to detection of the analyte, is preferably also used in the method of the first aspect of the invention.
[0152] As mentioned above, the method of this aspect of the present invention preferably uses a control analyte, an extension control, and a detection control as internal controls. It is clear that for these three controls to function together, the control reporter nucleic acid molecules generated / provided by the controls must be distinguishable from one another, i.e., must all have different sequences. Preferably, the control reporter nucleic acid molecules used in the method of the present invention each have the reverse sequence of a reporter nucleic acid molecule generated in response to detection of an analyte. In this case, it is clear that each control reporter nucleic acid molecule has the reverse sequence of a different reporter nucleic acid molecule generated in response to detection of an analyte.
[0153] Alternatively, the internal control may not be a separate component of the amplification reaction, but may be a unique molecular identifier (UMI) sequence present in each reporter nucleic acid molecule, which is unique to each molecule. This means that each individual reporter nucleic acid molecule generated upon detection of an analyte contains a UMI sequence. More specifically, it will be understood that each individual reporter nucleic acid molecule has a different UMI. The UMI is added to any sequence, such as a barcode, present in the reporter nucleic acid molecule as a means for detecting or identifying the analyte. As detailed above, the analyte is preferably detected by a proximity extension assay according to the method of the second aspect of the present invention. PEAs, including probes that can be used therein, are described above. As detailed above, the analyte is detected using a proximity probe pair, each of which binds to the analyte. Both probes of the probe pair contain nucleic acid domains that contain a barcode sequence specific to the analyte recognized by the probe.
[0154] Typically, when performing PEA, multiple identical probe pairs for each analyte to be detected are applied to a sample. By "identical" probe pair, we mean that all of the multiple probe pairs contain the same pair of analyte-binding molecules and the same pair of nucleic acid domains, so that each identical probe pair that binds to the target analyte generates the same reporter nucleic acid molecule, indicating that the analyte is present in the sample.
[0155] When a UMI sequence is used as an internal control, the probes used to detect each specific analyte are not identical. While a specific analyte-binding molecule pair is used, each individual probe, i.e., each individual probe containing at least one specific analyte-binding molecule of the pair, contains a different, unique nucleic acid domain. Each nucleic acid domain is unique due to the presence of a UMI sequence within it. This means that each specific probe pair that binds to a specific analyte molecule generates a unique reporter nucleic acid molecule. A unique reporter nucleic acid molecule is generated for each individual analyte molecule bound by a proximity probe pair. This allows absolute quantification of the amount of analyte present in a sample, because the exact number of detected analyte molecules can be counted based on the number of unique reporter nucleic acid molecules generated for a specific analyte.
[0156] UMIs can be advantageous because they not only enable quantification but also improve the resolution of measurements by allowing the number of reporter nucleic acid molecules generated in a detection assay to be calculated. UMIs allow for determining how many times a reporter nucleic acid molecule (e.g., the extension product of a PEA) has been amplified. Therefore, differences in UMI levels for reporter molecules for the same analyte can be detected. For example, individual reporter nucleic acid molecules for the same analyte may have the same barcode sequence but different UMIs. By detecting differences in different UMI levels, any biases that may occur in the PCR reaction can be detected and explained.
[0157] Improved resolution may also be useful or beneficial in control nucleic acid molecules. Accordingly, a UMI may alternatively or additionally be included in a control nucleic acid molecule. Thus, a UMI may be included in each individual control reporter nucleic acid molecule, e.g., a detection control molecule as described above, in the sense that it is added to each of the individual control nucleic acid molecules (it will be understood that each individual control nucleic acid molecule will have a different UMI). Alternatively, a UMI can be appropriately included in each different IC control format, e.g., extension control or control analyte, so that the UMI is included in the generated control reporter nucleic acid molecule. For example, the UMI may be included within the nucleic acid sequence of the nucleic acid domain of the extension control that serves as a template for the extension reaction, or within the sequence of a portion of the domain that serves as a primer for the extension reaction. Similarly, in the case of a control analyte, the UMI may be included in one or both of the nucleic acid domains of the proximity probe used to detect the control analyte, so that it is incorporated into the control reporter nucleic acid.
[0158] When UMIs are included in control nucleic acids, they can be used to improve the resolution of normalization. For example, UMIs can account for any PCR bias, as described above. This may allow very precise values to be used for normalization. Therefore, UMIs can be used as a tool for improving or ensuring data quality.
[0159] In an exemplary embodiment, the control reporter nucleic acid molecule comprises a sequence that is the reverse sequence of the reporter nucleic acid molecule that is generated in response to detection of the analyte, and a UMI.
[0160] UMI sequences may be used in proximity probes used in the method of the first aspect of the invention.
[0161] The method of the second aspect of the present invention may be applied to the detection of multiple analytes in the same sample (indeed, this is preferred). As detailed above, multiple analytes may be detected in a multiple detection assay. Each different analyte is detected based on the detection of an analyte-specific reporter nucleic acid molecule. As detailed above, the reporter nucleic acid for each different analyte has a unique barcode sequence to confer analyte specificity, but preferably all reporter nucleic acids contain a common primer binding site so that all reporter nucleic acid molecules can be amplified in a single PCR using the same primers. PCR amplification of the reporter nucleic acid molecules may include adding at least one (i.e., one or two) sequencing adapters to the ends of the reporter nucleic acid molecules, as detailed above.
[0162] As detailed above, when detecting multiple analytes in the same sample, different subsets of analytes may be detected in different aliquots of the sample based on the predicted abundance of the analytes in the sample, as detailed above. In this embodiment, a separate PCR is performed on each aliquot. The PCR products may then be pooled, as detailed above.
[0163] The method of the second aspect of the present invention may be used to detect one analyte in multiple samples or multiple analytes. In this embodiment, PCR is performed separately to amplify the reporter nucleic acid molecules generated from each sample. If different analyte subsets are detected in separate aliquots of each sample, PCR is performed separately for each separate aliquot of each sample. The same primers are used to amplify the reporter nucleic acid molecules generated for all analytes in all samples.
[0164] When multiple PCR amplifications are performed separately for multiple different samples and / or multiple different sample aliquots, if the internal control is a separate component present in the PCR mix, it will be present in each aliquot at a concentration proportional to the dilution of the aliquot, as described above. The concentration of the internal control will vary between aliquots with different dilutions, while the concentration of the internal control will be the same in aliquots from different samples with the same dilution (as in the first aspect of the invention). This allows for comparison of the relative amounts of each analyte present in each sample / aliquot, as detailed above.
[0165] In the method of the second aspect of the present invention, the PCR reaction is preferably carried out to saturation. Saturation of the PCR reaction has been described above. This is particularly advantageous when the method is used to detect multiple analytes present at different levels of abundance in one or more samples, with the detection assay being performed on multiple aliquots of each sample, detecting a subset of the analytes in each aliquot, as described above. Combining carrying out the PCR to saturation with the use of a separate component of the PCR mix as an internal control is a particularly preferred embodiment of the present invention. As detailed above, carrying out the PCR to saturation eliminates differences in the concentration of reporter nucleic acid molecules between different sample aliquots. Once saturation is reached, the overall concentration of reporter nucleic acid molecules present in each reaction will be essentially the same. Including an internal control in the reaction ensures that the relative levels of analyte detected in different aliquots or different samples can be compared.
[0166] As mentioned above, in the second aspect of the invention, it is particularly preferred to use analyte-specific probes to detect one or more analytes. When such probes are used to detect the analytes, an internal control (if a separate component of the PCR mixture) is typically added to the sample before or at the same time as adding the probes to the sample. Alternatively, as mentioned above, the internal control may comprise a UMI sequence present in each probe. Preferably, one or more analytes are detected by a proximity assay (e.g., PEA or PLA, particularly PEA) that generates a reporter nucleic acid molecule specific for each analyte. In this embodiment, it is preferred that at least an extension control is included. As noted above, it is most preferred that a control analyte, extension control, and detection control are all included.
[0167] In a preferred embodiment of the second aspect of the invention, the method is for detecting a plurality of analytes in a sample that are present at different levels in the sample, the method comprising: (i) preparing a plurality of aliquots from said sample; (ii) detecting a subset of analytes in each aliquot by performing a separate multiplex assay on each aliquot, wherein the analytes in each subset are selected based on their expected abundance in the sample; Each aliquot contains at least one internal control.
[0168] All parts of this embodiment may be as defined above in relation to the first aspect of the invention. The internal control may be any internal control defined above.
[0169] As mentioned above, when detecting analyte subsets in aliquots at different dilutions of the original sample, different amounts of internal control are added. The amount of internal control added to each aliquot is determined by the expected abundance of the analyte subset to be detected in that aliquot. As detailed above, this means that in practice, the amount of internal control used in each aliquot is proportional to the dilution of the aliquot.
[0170] The reporter nucleic acid molecules produced by the method of the second aspect of the present invention (i.e., more precisely, the PCR products resulting from the amplification of the reporter nucleic acid molecules) are preferably detected by DNA sequencing, most preferably using massively parallel DNA sequencing methods, as described above.
[0171] A third aspect of the present invention provides a method for detecting an analyte in a sample, wherein the analyte is detected by detecting a reporter nucleic acid molecule for the analyte, the method comprising performing a PCR reaction to generate a PCR product of the reporter nucleic acid molecule and detecting the PCR product, wherein an internal control is included in the PCR reaction, the internal control being present in a predetermined amount and being, comprising, or generating a control nucleic acid molecule, the control nucleic acid molecule comprising a sequence which is the reverse sequence of the reporter nucleic acid molecule.
[0172] All features of the third aspect of the invention may be as described in relation to the first and / or second aspect of the invention.
[0173] The present invention may be further understood with reference to the following non-limiting examples and drawings. [Brief explanation of the drawings]
[0174] [Figure 1]Figure 1 shows a conceptual diagram of the six different versions of the proximity extension assay detailed above. The upside-down "Y" represents an antibody as the analyte-binding domain of an exemplary proximity probe. [Figure 2] Figure 2 shows a schematic diagram of examples of extension controls that can be used in proximity extension assays. Parts A through F show suitable extension controls for use with versions 1 through 6 of Figure 1, respectively. Parts B through E show different possible extension controls, option (i) and option (ii), for use with versions 2 through 5 of Figure 1, respectively. The legend for Figure 1 also applies to Figure 2. [Figure 3] Figure 3 shows the counts (correctly paired barcodes) obtained from 367 assays performed on one plasma sample on a Log10 scale. A comparison was made between exposing the sample to a probe pool containing all 367 assays and exposing the sample to the same probe sets divided into four abundance blocks. The counts in assays in Blocks A and B were significantly increased compared to assays with lower counts without the abundance blocks, enabling high detection for the corresponding assays. The counts in Block D were similarly decreased compared to assays with higher counts without the abundance blocks, mitigating the loss of flow cell real estate. [Figure 4]Figure 4 shows the counts (with correctly paired barcodes) obtained from 367 assays performed on a single plasma sample, plotted on a linear scale. A comparison was made between exposing the sample to a probe pool containing all 367 assays and exposing the sample to the same probe set divided into four abundance blocks. The counts for assays in Blocks A and B were significantly increased compared to assays with lower counts without the abundance blocks, enabling high detection for the corresponding assays. The counts for Block D were similarly decreased compared to assays with higher counts without the abundance blocks, mitigating the loss of space on the flow cell. [Figure 5] Figure 5 shows box plots of 54 plasma samples exposed to the 372-assay probe pool, divided into four abundance blocks and sorted by median count within the block. Abundance blocks allow for detection of a wide range of protein abundance across samples without sacrificing detection or risking that the low end of an assay with high sample-to-sample variability falls below robust count detection. The dashed line indicates 100 counts as the threshold for sufficient count detection. [Example]
[0175] Example 1 - Exemplary Experimental Protocol Step 1 - Sample preparation and incubation Sixteen aliquots from each of the 48 to 96 plasma samples are incubated with each of up to 16 proximity probe pools (four abundance-specific blocks for each of four 384-probe pair panels) in 96- or 384-well incubation plates. For probe pools containing assays requiring pre-dilution, samples may be pre-diluted at 1:10, 1:100, 1:1000, and 1:2000. The dilution and dispensing of the plasma sample into the incubation solution can be done manually or by a dispensing robot, such as the LaboTec Mosquito® HTS. The incubation solution is dispensed into the wells of the plate. Add 1 μl of sample to 3 μl of incubation mix in the bottom of each well, seal the plate with adhesive film, spin at 400 x g for 1 minute at room temperature, and incubate at 4°C overnight. If using the dispensing robot mentioned above, the sample volume may be reduced to 0.2 μl and the incubation mix volume to 0.6 μl (5 times smaller).
[0176] The following table shows an exemplary reagent formulation: The probe solution may also contain other components, such as other blocking agents. [Table 1] [Table 2] [Table 3] [Table 4] [Table 5]
[0177] Step 2 - Proximity extension and PCR1 amplification Extension and amplification are performed using Pwo DNA polymerase. PCR1 is performed using a common primer for amplification of all extension products. The incubation plate (from step 1) is brought to room temperature and centrifuged at 400 x g for 1 minute. The extension mix (containing ultrapure water, DMSO, Pwo DNA polymerase, and PCR1 solution) is added to the plate, which is then sealed, vortexed briefly, and centrifuged at 400 x g for 1 minute. The plate is then placed in a thermal cycler for PEA reaction and preamplification (50°C for 20 minutes, 95°C for 5 minutes, (95°C for 30 seconds, 54°C for 1 minute, 60°C for 1 minute) x 25 cycles, 10°C hold). Preferably, the extension mix can be dispensed into the plate using a dispenser robot, such as a Thermo Scientific™ Multidrop™ Combi Reagent Dispenser. The forward common primer contains the Illumina P5 sequencing adapter sequence (SEQ ID NO: 1). [Table 6] [Table 7]
[0178] Step 3 – Pooling Blocks by Abundance The PCR1 products from each of the four abundance-specific blocks from the 384 probe pair panel are pooled together, resulting in a maximum of four PCR1 pools per sample for each 384 probe pair panel. Different amounts can be drawn from each block to balance the relative levels of the assay between blocks. PCR1 products can be pooled manually or by robotic pipetting.
[0179] Step 4 - PCR2 indexing Prepare a primer plate containing 48 to 96 reverse primers (typically one primer for each well of a 96-well plate). Each reverse primer contains an "Illumina P7" sequencing adapter sequence (SEQ ID NO: 2) and a sample index barcode. A unique barcode sequence is used for PCR1 products from each different sample. Preferably, up to four PCR1 pools containing the same plasma sample (384 probe pairs per panel) are each labeled with the same sample index for ease of identification and data processing. Add a forward common primer (the same forward primer used in PCR1) containing an "Illumina P5" sequencing adapter sequence to the PCR2 solution. Each PCR1 pool is contacted with PCR2 solution containing a forward common primer, a single reverse (sample index) primer from the primer plate, and DNA polymerase (Taq DNA polymerase or Pwo DNA polymerase). Amplification is carried out by PCR (95°C for 3 minutes, (95°C for 30 seconds, 68°C for 1 minute) x 10 cycles, hold at 10°C) until the primers are exhausted. The theoretical final concentration of the pooled PCR1 product is 1 μM (all primers used). For PCR2, the PCR1 amplicon is diluted 1:20, resulting in a starting concentration of 50 nM in each PCR2 reaction. The concentration of each PCR2 primer is 500 nM. Therefore, the PCR2 primers should be depleted after 3.3 cycles (10-fold amplification). [Table 8] [Table 9] [Table 10]
[0180] Step 5 - Final Pool All 48 to 96 indexed sample pools that belong to the same 384 probe pair panel are pooled together, adding equal amounts from each sample, resulting in up to four pools (i.e., libraries) per 384 probe pair panel.
[0181] Step 6 - Purification and Quantitation (Optional) The libraries are purified separately using magnetic beads, and the total DNA concentration of the purified libraries is determined by qPCR using a DNA standard curve. AMPureXP beads (Beckman Coulter, USA), which preferentially bind longer DNA fragments, can also be used according to the manufacturer's protocol. AMPureXP beads bind long PCR products but not short primers, allowing the PCR products to be purified from remaining primers. The depletion of PCR2 primers means that this purification step may not be necessary.
[0182] Step 7 - Quality Control (Optional) A small aliquot of each (purified) library is analyzed in an Agilent Bioanalyzer (Agilent, USA) according to the manufacturer's instructions to confirm successful DNA amplification.
[0183] Step 8 - Sequencing The libraries are sequenced using an Illumina platform (e.g., the NoveSeq platform). Up to four libraries (obtained from each of the 384 probe pair panels) are run in separate "lanes" of the flow cell. Depending on the size and model of the flow cell and sequencer used, up to four libraries may be sequenced in parallel or sequentially (one after the other) in different flow cells.
[0184] Step 9 - Data Output The sequences of the barcodes (from each reporter nucleic acid molecule) and sample indexes (from the sample index primer) are identified in the data, counted, summed, and aligned / labeled according to a known barcode-assay-sample key. "Matching barcodes" represent interactions between two paired PEA probes. The count correlates to the number of interactions in the PEA. Counts for each assay and sample must be normalized using an internal reference control so that comparisons can be made between samples. Each of the four abundance blocks has its own internal reference control. Each of the 384 probe pair panels is separated based on the lane they are read in. Each panel contains the same 96 sample indexes, the same 384 barcode combinations, and an internal reference control.
[0185] Example 2 - PEA with and without abundance blocks Multiplex PEA (using probes containing antibodies bound to nucleic acid domains with the structure described in version 6 above) was performed to detect 367 proteins in plasma samples. Each probe contained a unique barcode sequence. A proximity-probe pool containing all 367 assays was incubated with the sample, and as a comparison, four aliquots from each plasma sample were incubated with each of the four proximity-probe pools (four abundance-specific blocks containing 367 assays) in 96- or 384-well incubation plates. PEA was performed as described above, except that step 3 was omitted for proximity probe pools without abundance-sorted blocks. During amplification of the extension products, P5 and P7 sequencing adapters were added to each end of the products, along with unique sample indices for the reporter nucleic acid molecules from each different sample. All extension products were sequenced by massively parallel DNA sequencing employing reversible dye terminator sequencing technology on the Illumina NovaSec platform. Extension products from the 367-assay probe pool and the pooled abundance-sorted blocks (total of 367 assays) were sequenced separately on separate flow cells at separate times.
[0186] The results for one of the plasma samples can be seen in Figures 3 and 4. The table below shows the ratio between the highest and lowest assays (counts) for the same plasma sample with and without abundance blocks. The ratios for the abundance blocks are significantly lower than the ratio for the pool of all 367 assays, meaning that measurements for these assays are obtained in a more optimal manner using the space on the flow cell (higher counts for low abundance assays and lower counts for high abundance assays). [Table 11]
[0187] Example 3 - PEA of samples with assays of different abundances using abundance blocks Multiplex PEA (using probes containing antibodies bound to nucleic acid domains with the structure described in version 6 above) was performed to detect 372 proteins in 54 plasma samples. Each probe contained a unique barcode sequence. Four aliquots from each plasma sample were incubated with each of four proximity probe pools (four abundance-specific blocks containing 372 assays) in 96- or 384-well incubation plates. PEA was performed as described above. During amplification of the extension products, P5 and P7 sequencing adapters were added to each end of the products, along with unique sample indices for the reporter nucleic acid molecules from each different sample. All extension products were sequenced by massively parallel DNA sequencing using reversible dye terminator sequencing technology on the Illumina NovaSec platform. The results in Figure 5 demonstrate that protein targets with a wide range of abundances can be detected in samples without sacrificing the low range of proteins with high variability between samples or assays with relatively low abundance across all 54 samples for signal loss (robust abundance, e.g., counts below 100 counts).
Claims
1. 1. A method for detecting a plurality of analytes in a sample, the analytes being present at different levels in the sample, the method comprising: (i) preparing a plurality of aliquots from said sample; and (ii) performing a separate block of detection assays for each of the plurality of aliquots; In the above (ii), each detection assay block detects a subset of analytes that is different from the subset detected by another detection assay block; A method wherein each analyte in each of said analyte subsets is assigned to a respective block of detection assays based on the predicted abundance of that analyte.
2. The method of claim 1 , wherein the analyte is a non-nucleic acid analyte.
3. The method of claim 1 or 2, wherein the analyte is or comprises a protein.
4. 4. The method of claim 1, wherein the analytes are detected in each aliquot by detecting a reporter nucleic acid molecule specific for each analyte.
5. The method of claim 4 , wherein the reporter nucleic acid molecule is generated in the detection assay performed on each aliquot.
6. The method of claim 4 or 5, wherein the reporter nucleic acid molecule is amplified by PCR and preferably detected by nucleic acid sequencing.
7. 7. The method of claim 6, wherein one or more sequencing adaptors are added to the reporter nucleic acid molecule in one or more amplification and / or ligation steps.
8. 8. The method of claim 6 or 7, wherein at least a first PCR reaction is performed on the reporter nucleic acid molecule to add at least a first nucleic acid sequencing adaptor.
9. 9. The method of claim 8, wherein a second PCR reaction is performed on the PCR product of the first PCR reaction to add a second nucleic acid sequencing adaptor.
10. 10. The method of claim 6, wherein at least one PCR reaction is carried out to saturation.
11. 11. The method of any one of claims 1 to 10, wherein the reaction products of the separate detection assays, or, if the reaction products are nucleic acid molecules, their amplification products, are pooled to form a first pool, in which the reaction products or amplification products are amplified.
12. The reaction product of the detection assay is a reporter nucleic acid molecule, and the method comprises:
12. The method of claim 11, comprising amplifying the reporter nucleic acid molecule in a first PCR reaction performed separately on each individual aliquot to generate a first PCR product, pooling the first PCR products from the individual aliquots to create a first pool, and performing a second PCR reaction on the first pool.
13. 13. The method of claim 11 or 12, wherein the reaction products or amplification products thereof are added to the first pool in different amounts.
14. 14. The method of any one of claims 11 to 13, wherein the method is performed separately and in parallel on a plurality of different samples to generate reaction products or amplification products thereof for each sample, a separate first pool is created for each sample, and a sample index is added to the products of the first pool by an amplification reaction and / or a ligation reaction.
15. 15. The method of claim 14, wherein the separate first pool created for each sample comprises a first PCR product, and a sample index is added to the first PCR product in the second PCR reaction performed on the first pool of each sample.
16. 16. The method of claim 14 or 15, wherein the indexed first pools generated for each sample are pooled together to create a second pool for nucleic acid sequencing.
17. 17. The method of claim 6, wherein the PCR reaction includes an internal control in each aliquot.
18. 18. The method of any one of claims 4 to 17, wherein the reporter nucleic acid molecule is generated in a proximity probe detection assay, in particular a proximity extension assay (PEA).
19. the reporter nucleic acid molecule comprises at least one barcode sequence, and detecting the reporter nucleic acid molecule comprises detecting the at least one barcode sequence, optionally together with a sample index; 17. The method of any one of claims 4 to 16, wherein the reporter nucleic acid molecule comprises a combination of barcode sequences derived from the nucleic acid domains of a pair of proximity probes, and wherein detecting the reporter nucleic acid molecule comprises detecting the combination of barcode sequences.
20. 20. The method of any one of claims 1 to 19, wherein the sample is a plasma or serum sample.
21. The analyte is detected using a proximity probe pair, each proximity probe comprising: (i) an analyte-binding domain specific for the analyte; (ii) a nucleic acid domain; both probes of each probe pair contain analyte-binding domains specific for the same analyte, and each probe pair is specific for a different analyte, and each probe pair is designed such that when the pair of proximity probes bind to their respective analytes in proximity, the nucleic acid domains of the proximity probes interact to generate a reporter nucleic acid molecule; at least two panels of proximity probe pairs are used, each panel for detecting a different group of analytes, and a separate aliquot of the sample is prepared for each panel to detect a different subset of the analytes within the group; (a) in each panel, each probe pair comprises a different pair of nucleic acid domains; and (b) in different panels, the probe pairs comprise the same pair of nucleic acid domains.
21. The method of any one of claims 18 to 20.
22. A method for detecting an analyte from different samples, comprising adding a sample index to the PCR products generated by amplification of the reporter nucleic acid molecules generated for each sample; The PCR products generated from each of the different samples using the same proximity probe pair panel are pooled into a nucleic acid sequencing panel pool, and the PCR products generated using each panel are pooled into separate panel pools; 22. The method of claim 21, wherein each panel pool is sequenced separately.
23. 23. The method of any one of claims 7 to 22, wherein the nucleic acid sequencing method is a massively parallel DNA sequencing method.
Citation Information
Patent Citations
Proximity probing methods and kits
JP2003524419A
immune amplification
JP2007525174A
Method for creating a proximity probe
JP2018533944A
Multiplexed digital assay for variant and normal forms of a gene of interest
US20140274799A1
Immunoassays with enhanced selectivity
WO2006081651A1