Methods and kits for whole genome amplification and analysis of target molecules in a biological sample

The whole-genome amplification technology, which combines labeled oligonucleotides with antibody conjugates, solves the problem of simultaneous analysis of genome and protein expression in single cells, and enables comprehensive analysis of genotype and phenotype.

CN114867867BActive Publication Date: 2026-03-24MENARINI SILICON BIOSYSTEMS SPA
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-16
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing technologies cannot simultaneously analyze genome sequences and protein expression at the single-cell level, nor can they reanalyze additional targeted genome information from single cells.

Method used

By using labeled oligonucleotides and antibody conjugates to bind to target molecules, genomic DNA and protein markers are simultaneously amplified using whole-genome amplification (WGA) technology to generate a sequenceable library, which is then analyzed using NGS technology.

Benefits of technology

It enables simultaneous analysis of genome copy number maps and protein expression at the single-cell level, providing complete single-cell information, including genotype and phenotype data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114867867B_ABST
    Figure CN114867867B_ABST
Patent Text Reader

Abstract

Disclosed is a method for whole genome amplification and analysis of a plurality of target molecules in a biological sample comprising genomic DNA and target molecules, comprising the steps of: contacting the biological sample with at least one binding agent for at least one target molecule conjugated to a labeled oligonucleotide comprising a binding agent barcode sequence (BAB) and a unique molecular identifier sequence (UMI); performing a separation step to selectively remove unbound binding agent, thereby obtaining a labeled biological sample; simultaneously performing whole genome amplification and amplification of the labeled oligonucleotide on the labeled biological sample; preparing a massively parallel sequencing library from the amplified labeled oligonucleotide; sequencing the massively parallel sequencing library; retrieving the sequences of the BAB and UMI from each sequencing read; calculating the number of different UMIs for each binding agent.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This patent application claims priority to Italian Patent Application No. 102019000024159, filed on December 16, 2019, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This invention relates to whole-genome amplification and analysis of target molecules in biological samples, particularly single-cell samples, and particularly to methods and kits for protein quantification. Existing technology

[0004] Single-cell analysis methods allow for the acquisition of information about cell state without the complications arising from heterogeneity in batch samples. Analyzing the proteome and genome within the same single cell can provide correlations between cell phenotype and genotype, thus offering unique insights into different biological and pathological processes. This is particularly true for tumors, where somatic cell-acquired genetic heterogeneity and its impact on transcription and protein translation are key factors in cancer initiation and evolution.

[0005] Whole-genome amplification (WGA) can be used to analyze genomes from single cells to obtain more DNA and simplify and / or allow for different types of genetic analysis, including sequencing, SNP detection, etc. WGA based on deterministic restriction sites (DRS-WGA) with LM-PCR is known from EP1109938.

[0006] WO 2017 / 178655 and WO 2019 / 016401 teach a method from DRS-WGA (e.g., Amplil) TM A simplified method for preparing massively parallel sequencing libraries (WGA) or MALBAC for low-pass whole-genome sequencing and copy number analysis.

[0007] Recently, methods for simultaneously analyzing the genome and transcriptome of single cells have been developed. A paper by Dey S. et al., 2015, "Integrated genome and transcriptome sequencing of the same cell," Nature Biotechnology, 33(3), 285-289, http: / / doi.org / 10.1038 / nbt.3129, teaches a method in which messenger RNA from hand-isolated single cells is first transcribed into single-stranded cDNA, which is then amplified along with genomic DNA via pseudo-linear whole-genome amplification. Two separate libraries, one from cDNA and one from genomic DNA, are then prepared and sequenced. In another approach, Macaulay, IC et al., 2015, G&T-seq: Parallel sequencing of single-cell genomes and transcriptomes, Nature Methods, 12(6), 519-522. http: / / doi.org / 10.1038 / nmeth.3370, oligo-dT coated beads were used to physically separate mRNA from gDNA in order to capture and isolate polyadenylated mRNA molecules from completely lysed single cells. Then, a modified Smart-seq2 protocol was used to amplify mRNA (Picelli, S. et al., 2013, Smart-seq2 for sensitive full-length transcriptome profiling in single cells, Nature Methods, 10(11), 1096-1100. http: / / doi.org / 10.1038 / nmeth.2639), while gDNA could be amplified and sequenced using available whole-genome amplification methods. While these methods help link genotype to messenger transcription, they do not directly detect proteins that are actively regulated in the cell, their translation, and flipping / degradation.

[0008] Currently, the most widely used single-cell protein detection methods rely on targeting specific proteins with labeled antibodies. Fluorescence-based detection and quantification of proteins via fluorescence-activated cell sorting (FACS) or fluorescence microscopy allows for low-level multiplexing of proteins in single cells using fluorescently labeled antibodies that recognize specific cellular proteins. However, this method is typically limited to 10-15 simultaneous measurements because highly multiplexed fluorophore-based assays are challenged by spectral overlap between the emission spectra of multiple dyes. Furthermore, complex algorithms are required to deconvolve overlapping spectra.

[0009] Fluidigm mass cytometer (CyTOF™) detects proteins using antibodies labeled with metal-containing polymer tags (MAXPAR™). The instrument is based on a non-optical physical detection principle and the diverse chemical properties of the labels. Fluorescent labels are replaced by specially designed polyatomic element labels, and detection leverages the high resolution, sensitivity, and speed of time-of-flight mass spectrometry (TOF-MS) analysis. Because many available stable isotopes can be used as labels, it is potentially possible to simultaneously detect many proteins in a single cell [Ornatsky, O. et al., 2010, High-multiparametric analysis by mass cytometry, Journal of Immunological Methods, 361(1-2), 1-20. http: / / doi.org / 10.1016 / j.jim.2010.07.002]. Frei et al., 2016, "Highly multiplexed simultaneous detection of RNAs and proteins in single cells," Nature Methods, 13(3), 269-275, http: / / doi.org / 10.1038 / nmeth.3742, taught a method for the simultaneous detection of RNA and proteins in single cells using RNA-based neighbor-linkage analysis (PLAYR). PLAYR enables highly multiplexed quantification of transcripts in single cells via mass flow cytometry, allowing for the simultaneous quantification of more than 40 different mRNAs and proteins. Finally, mass cytometry allows for the study of multiple cellular processes and phenotypic features, as well as protein and messenger RNA transcription, such as protein phosphorylation (Bendall, SC et al., 2011, Single-Cell Mass Cytometry of Differential Immune and Drug Responses Across a Human Hematopoietic Continuum, Science, 332(6030), 687-696).http: / / doi.org / 10.1126 / science.1198704) and cell proliferation (Behbehani, GK et al., 2012, Single-cell mass cytometry adapted to measurements of the cell cycle, Cytometry Part A, 81A(7), 552-566, http: / / doi.org / 10.1002 / cyto.a.22075).

[0010] The limitations of these methods are:

[0011] - Due to the dynamics of ion flight in mass spectrometers, the throughput of mass spectrometers lags behind that of fluorescence-based instruments. In addition, the sensitivity of mass reporter genes is lower than that of a few fluorophores with higher quantum efficiency (such as phycoerythrin), which makes it more difficult to measure molecular features expressed at very low levels using mass flow cytometry (Spitzer, MH et al., 2016, Mass Cytometry: Single Cells, Many Features, Cell, 165(4), 780-791, http: / / doi.org / 10.1016 / j.cell.2016.04.019).

[0012] Importantly, because the cells are atomized and ionized, they cannot be recovered after analysis, thus making it impossible to analyze genomic DNA.

[0013] Fredriksson et al. described a method for detecting proteins using oligonucleotide-labeled antibodies in a 2002 paper, *Protein Detection Using Proximity-Dependent DNA Ligation Assays* (Nature Biotechnology, 20(5), 473-477, http: / / doi.org / 10.1038 / nbt0502-473), which taught a proximity ligation technique (PLA) in which the coordinated and proximal binding of two DNA aptamers to a target protein facilitates the ligation of oligonucleotides to each aptamer affinity probe. The ligation of two such neighboring probes produces an amplifiable DNA sequence that reflects the identity and quantity of the target protein. Method 3PLA (Schallmeiner, E. et al., 2007, Sensitive protein detection via Triple-binder proximity ligation assays, Nature Methods, 4(2), 135-137, http: / / doi.org / 10.1038 / nmeth974) extends the sensitivity and specificity of proximity ligation methods by using three recognition events and allows the detection of as few as one hundred target molecules. In 3PLA, a set of three oligonucleotide-modified antibody reagents binds to a single target protein to generate a detectable signal via proximity ligation. The 3' and 5' ends of the oligonucleotides on two neighboring probes are able to hybridize with an oligonucleotide present on a third neighboring probe to form a complex containing the three probes and the target protein. This allows the two oligonucleotides to be ligated via an intermediate fragment through two ligation reactions, using the third neighboring probe as a template, to form a specific, amplifiable DNA strand that can be detected by qPCR.Proximity extension assay (PEA) is a variant of PLA in which an antibody labeled with two oligonucleotides binds to a single protein, the oligonucleotides partially anneal at their 3' ends, and extension by polymerase produces an amplifiable DNA sequence detectable by qPCR (Lundberg, M. et al., 2011, Homogeneous antibody-based proximity extension assays provide sensitive and specific detection of low-abundant proteins in human blood, Nucleic Acids Research, 39(15), http: / / doi.org / 10.1093 / nar / gkr424). Although the methods described above are not specifically designed for single-cell protein detection, Fluidigm Cl. TM A single-cell automated preparation system was used to automatically prepare an amplifiable target set of 92 proteins, using PEA analysis to detect up to 96 single cells per run (Egidio C. et al., 2014, using Cl...). TM A Method for Detecting Protein Expression in Single Cells Using the Automated Single-Cell Preparation System TM Single-Cell AutoPrep System (TECH2P.874), J Immunol, 192(1 Supplement) 135.5). Fluidigm Cl microfluidic systems support a range of single-cell biology methods to analyze transcriptome or genomic DNA sequences through whole-exome sequencing and targeted DNA sequencing, but these methods cannot be easily combined to obtain genotype and phenotype information from the same single cell.

[0014] Therefore, PLA and PEA assays based on qPCR are sensitive and highly specific, but are limited by their throughput and the fact that they can only detect proteins.

[0015] NanoString Technologies, Inc.'s US 9,714,937 teaches a method for detecting proteins using a capture antibody conjugated to a first region of a target protein (e.g., biotin) and a detection antibody specific to a second region of the target protein, along with a nanoreporter molecule comprising multiple separable tags linked to the detection antibody via hybridization with a linker oligonucleotide. The two antibodies form a complex with the target protein, which can form a bond with a matrix or beads that has a high affinity for that region. The target is detected and quantified by counting the number of nanoreporter molecules. Nanostring, Inc.'s commercially available assay, based on nCounter® digital molecular barcoding technology, detects proteins using antibodies labeled with unique oligonucleotides that target specific protein epitopes. Unique single-stranded DNA tags are detected using a combination of a biotinylated capture probe and a reporter probe made from single-stranded DNA molecules annealed with a series of fluorescently labeled RNA fragments. The linear sequence of these tags creates a unique barcode for each target of interest. The complex is then immobilized to an imaging surface via a non-covalent bond between biotin and immobilized streptavidin molecules, and the fluorescent barcode is imaged and counted. The count value of each protein-specific barcode is a numerical measurement directly related to the number of molecules present in the sample. Protein detection can be combined with messenger RNA detection by using capture probe-reporter probe pairs designed for specific target RNAs. A single analysis can analyze approximately 30 protein targets and 770 mRNA targets.

[0016] The drawback of this method is that it requires a large number of cells to analyze RNA (equivalent to 2500 cells) and / or to analyze proteins (equivalent to 100,000 cells), and it is not suitable for analyzing single cells. This method could potentially be used to detect other analytes, such as genomic DNA, but it does not provide a direct readout of the genomic sequence, only a signal indicating the presence / absence of a known sequence. Based on hybridization, it has partial inclusion for sequence variants, but may not be able to distinguish between different sequence variants.

[0017] Stoeckius et al., 2017, Simultaneousepitope and transcriptome measurement in single cells, Nature Methods, Vol. 14, pp. 865-868, and US 2018 / 0251825 first described an NGS-based method for integrating the analysis of multiple proteins and RNA transcripts in single cells, termed Cell Indexing of Transcriptome and Epitopes via Sequencing by Imaging (CITE-seq). This method relies on oligonucleotide-labeled antibodies that integrate cellular protein and transcriptome measurements into single-cell reads via a 3'-polyadenylate tail on the antibody label, similar to that present on messenger RNA. This method is compatible with droplet-based methods for sample dispensing in single-cell and single-cell library preparation, such as those offered by 10X Genomics. More specifically, in the CITE-seq method, cells stained with oligonucleotide-labeled antibodies targeting cell surface epitopes are dispensed via a microfluidic device into droplets containing lysin and barcode beads. Barcoded antibodies and mRNA from each single cell / droplet are captured by beads with unique cell barcodes. The mRNA is then reverse transcribed and amplified with oligonucleotides from the barcoded antibodies to generate an NGS library ready for sequencing. Finally, sequence counting is used to quantify the barcoded antibodies. Similarly, Peterson et al., 2017, Multiplexed quantification of proteins and transcripts in single cells, Nature biotech., (35)10:936-939, teach a method for RNA expression and protein quantification based on DNA-labeled antibodies and droplet microfluidics, which allows for the quantification of proteins using 82 barcoded antibodies and the analysis of >20,000 transcripts in a single cell. Both of these methods utilize the DNA polymerase activity of reverse transcriptase while simultaneously labeling antibodies with oligonucleotides from poly(dT) cell barcode extension primers and synthesizing complementary DNA from mRNA in the same reaction. On the other hand, other droplet-based methods can be used to analyze whole-genome copy number maps or analyze genomic sequences in single cells. For example, 10X Genomics' commercial solution, the Chromium Singlecell CNV Solution, allows copy number analysis of hundreds to thousands of single cells, and MissionBio's Tapestri® platform provides targeted single-cell DNA sequencing for sequencing and CNV analysis of multiple gene sets.

[0018] The disadvantages of these methods are:

[0019] Neither of the two methods described above for simultaneous transcriptomics / proteomics analysis allows for the simultaneous analysis of genomic sequences and proteins or transcripts because genomic DNA does not have the polyadenylate tail necessary for its amplification.

[0020] - Copy number and / or targeted sequencing methods based on Dropseq are only suitable for analyzing genomic DNA, but do not provide any information about single-cell phenotypes, such as single-cell transcriptomic profiles or quantification of surface markers or other proteins.

[0021] - In droplet-based single-cell partitioning methods, the single cell and all its information content are essentially "destroyed" during the process, and it is impossible to recover the single cell after the procedure is completed for further analysis. Summary of the Invention

[0022] Therefore, the object of this invention is to provide a method for whole-genome amplification and analysis of multiple target molecules in biological samples, which simultaneously allows for the analysis of whole-genome copy number maps / genome sequences and the analysis of protein expression on the same single cell, and in particular overcomes one or more of the following disadvantages of the prior art:

[0023] - Unable to detect and quantify proteins in the same sample and analyze the genome down to single-cell resolution.

[0024] - Unable to reanalyze additional targeted genomic information from single cells.

[0025] This objective is achieved by the method defined in claim 1.

[0026] Another object of the present invention is to provide the kit as described in claim 17.

[0027] Brief description of the attached figures

[0028] Figure 1 It shows the absence of ( ) according to the invention Figure 1 A) or having ( Figure 1 B) General structures of two possible implementations of 3'-labeled oligonucleotide sequences. PL = payload sequence; 5-TOS = first-labeled oligonucleotide amplification sequence; 3-TOS = second-labeled oligonucleotide amplification sequence; UMI = unique molecular identifier sequence; BAB = binding barcode sequence.

[0029] Figure 2The graph shows labeled oligonucleotide libraries representing oligonucleotides from two different labels, one of which readily forms intramolecular hairpins between 5-TOS and 3-TOS at different temperatures (“with hairpins”: Tm = 67℃; “without hairpins”: Tm = 45℃). The x-axis represents the expected UMI count, and the y-axis represents the UMI count after sequencing.

[0030] Figure 3 The graph shows the amplification of different amounts of oligonucleotide mixtures in 27 PCR cycles. Each dilution was performed from three independent dilution replicates. The x-axis represents the expected number of different UMIs, and the y-axis represents the experimentally observed number of UMIs.

[0031] Figure 4 Three extended plots are shown for four oligonucleotide mixtures with different amounts of different BABs. There are four data points for each dilution and one for each oligonucleotide. Figure 4 A: Perform 23 PCR cycles. Figure 4 B: Perform 27 PCR cycles. Figure 4 C: Perform 35 PCR cycles. The x-axis represents the expected number of different UMIs, and the y-axis represents the experimentally observed number of UMIs.

[0032] Figure 5 The structure of an oligonucleotide labeled with a primer amplification according to an embodiment of the present invention is shown. Additional notes: 5-WGAH = 5'WGA handle sequence; 3-WGAH = 3'WGA handle sequence; 1AH = first amplification handle sequence; 2AH = second amplification handle sequence.

[0033] Figure 6 The structure of an oligonucleotide with a label amplified by at least one second primer is shown according to another embodiment of the present invention. Figure 6 A: The labeled oligonucleotides and annealing sites correspond to the structure of the relatively extended oligonucleotides of BAB. Figure 6 B: The labeled oligonucleotide and annealing site do not correspond to the structure of the relative extended oligonucleotide of BAB. Additional notes: Ep = 5' extended oligonucleotide sequence; SS = spacer sequence; AS = annealing sequence; AS-RC = reverse complementary sequence of annealing sequence.

[0034] Figure 7 This shows the melting temperature ([Na+]) of the hairpin induced by the 15 nt complementary sequence located at the end of the ssDNA molecule. + ] = 150mM; [Mg +The figure shows the computer prediction of the change of molecular length (4mM) as a function of molecular length (in silico prediction).

[0035] Figure 8 The structure of a labeled oligonucleotide with at least three primers amplified is shown according to another embodiment of the present invention.

[0036] Figure 9 This demonstrates a general method for generating libraries. Figure 9 A: According to Figure 5 The implementation method generates a library from labeled oligonucleotides. Figure 9 B: According to Figure 6 Implementation of A involves generating a library from labeled oligonucleotides. Figure 9 C: According to Figure 8 The implementation method generates a library from labeled oligonucleotides. Additional note: 2AH-RC = reverse complementary sequence to the second amplification handle sequence.

[0037] Figure 10 The design of the P5-Synth oligonucleotide and corresponding library primers disclosed in Example 1 is shown.

[0038] Figure 11 The method for generating an NGS library of oligonucleotide P5-Synths using library primers according to Example 1 is shown.

[0039] Figure 12 Scatter plots of PBMC and SK-BR-3 cells stained with Ab oligonucleotides and secondary fluorescent antibodies are shown. On the x-axis, fluorescence levels in the APC channels are proportional to the amount of Ab oligonucleotide markers 1, 2, and 4. On the y-axis, fluorescence levels in the PE channels are proportional to the amount of Ab oligonucleotide marker 3.

[0040] Figure 13 Electrophoretic diagram of a library generated from P5-Synth-labeled oligonucleotides from a single cell according to Example 1 is shown.

[0041] Figure 14 Showing according to Figure 8 The implementation method describes the protein quantification results of labeled oligonucleotide amplification after WGA in single cells. Cytokeratin (CK) was quantified separately. Figure 14 A), Her2 ( Figure 14 B), CD45 Figure 14 C) and IgG1 isotype control ( Figure 14 D) UMI count. The y-axis shows the number of UMIs, and the x-axis shows the UMI count based on... Figure 9 The isolated cell types.

[0042] Figure 15Showing according to Figure 8 The implementation method describes the protein quantification results of labeled oligonucleotide amplification of single cells during WGA. Cytokeratin (CK) was quantified separately. Figure 15 A), Her2 ( Figure 15 B), CD45 Figure 15 C) and IgG1 isotype control ( Figure 15 D) UMI count. The y-axis shows the number of UMIs, and the x-axis shows the UMI count based on... Figure 9 The isolated cell types.

[0043] Figure 16 The design of the P5-Lib1 oligonucleotide and corresponding library primers disclosed in Example 3 is shown.

[0044] Figure 17 The protocol for generating an NGS library of P5-Lib1 oligonucleotides using library primers is shown.

[0045] Figure 18 An example of a low-pass map of CNA analysis obtained from a single cell is shown. Figure 18 A: Possesses a P5-Lib1 oligonucleotide peak and according to Figure 5 The single-cell CNA atlas was processed using the implementation method. Figure 18 B: Possesses a P5-Synth oligonucleotide peak and according to Figure 8 The single-cell CNA atlas was processed using the implementation method. Figure 18 C: Single-cell CNA maps of unlabeled oligonucleotide peaks. All maps correspond to SK-BR-3 cells with typical gain and loss. Minor variations are due to single-cell genomic heterogeneity.

[0046] Figure 19 The diagram shows a labeled oligonucleotide library obtained from a single cell with P5-Synth and P5-Lib1 peaks. The x-axis represents the expected UMI count, and the y-axis represents the UMI count after sequencing.

[0047] definition

[0048] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. While many methods and materials similar to or equivalent to those described herein may be used in the practice or testing of this invention, preferred methods and materials are described below. Unless otherwise stated, the techniques used in this invention described herein are standard methods well known to one of ordinary skill in the art.

[0049] "ab-oligo mix" refers to a solution containing all ab-oligonucleotides and / or isotype control ab-oligonucleotides that target intracellular (i.e., "internal ab-oligonucleotide mix") and / or external (i.e., "external ab-oligonucleotide mix") epitopes.

[0050] "Antibody-oligonucleotide conjugate," "ab-oligonucleotide conjugate," or "ab-oligonucleotide" refers to a synthetic molecule derived through the chemical conjugation of an antibody molecule with an ssDNA oligonucleotide molecule. Chemical conjugation typically involves a specific chemical reaction that allows the two molecules to be covalently linked. The antibody:oligonucleotide stoichiometry can be controlled to have a specified ratio. In the initial steps of WGA, the antibody portion is usually digested, leaving only the oligonucleotide portion. For simplicity, these molecules will still be referred to as ab-oligonucleotide molecules or ab-oligonucleotide amplicon in this description.

[0051] The acronym "APC" refers to the fluorophore phycocyanin.

[0052] "Binding agent barcode sequence (BAB)" refers to a unique DNA oligonucleotide sequence that identifies the binding agent.

[0053] "Balanced PCR amplification" refers to the characteristic of PCR that it performs amplification of multiple targets, so that in each PCR cycle, essentially every target molecule is amplified.

[0054] "Binding agent" refers to a molecule that can specifically bind to a specified target molecule (e.g., protein or glycosylated protein or phosphorylated protein) (e.g., antibodies, affinity molecules, ligands, aptamers, synthetic binding proteins, small molecules, as non-limiting examples).

[0055] “CITE-Seq” or “Cellular Indexing of Transcriptomes and Epitopes by Sequencing” refers to a method developed by Stoeckius et al. for simultaneous protein quantification and mRNA sequencing in single cells.

[0056] "CyTOF," or "cytometry by time of flight," refers to a device that performs mass spectrometry flow cytometry, enabling the quantification of proteins in single cells using a combination of mass spectrometry and flow cytometry. Cells are stained using a binding agent conjugated with heavy metal isotopes.

[0057] The term "conjugate" refers to a molecule obtained from the covalent conjugation of a binding agent and a labeled oligonucleotide.

[0058] "Copy number alteration (CNA)" refers to somatic variations in the copy number of a genomic region, usually defined relative to the genome of the same individual.

[0059] "DNA library purification" refers to the process of separating DNA library material from unwanted reaction components (such as enzymes, dNTPs, salts, and / or other molecules that are not part of the desired DNA library). Examples of DNA library purification processes include purification using Agencourt AMPure or Merck Millipore Amicon centrifuge columns or solid-phase reversible immobilization (SPRI) beads (such as those from Beckman Coulter).

[0060] "DNA library quantification" refers to the process of quantifying DNA library materials. Examples of DNA library quantification processes include using QuBit technology, electrophoresis analysis (Agilent Bioanalyzer 2100, Perkin Elmer LabChip technology), or the RT-PCR PicoGreen system (Kapa Biosystems).

[0061] "Dynamic range" refers to the ratio between the maximum and minimum values ​​of a quantity that can be assumed.

[0062] "Library primers" refer to ssDNA molecules used as primers to generate large-scale parallel sequenceable libraries from labeled oligonucleotides.

[0063] "Low-pass whole genome sequencing" or "low-pass sequencing" refers to whole genome sequencing with an average sequencing depth of less than 1.

[0064] "Massively parallel sequencing" or "next-generation sequencing (NGS)" refers to a method for sequencing DNA, including creating spatially and / or temporally isolated, cloned, and sequenced DNA molecular libraries (with or without prior clonal amplification). Examples include the Illumina platform (Illumina Inc.), the IonTorrent platform (ThermoFisher Scientific Inc.), the Pacific Biosciences platform, and Minion (Oxford Nanopore Technologies Ltd.).

[0065] "Multiple Annealing and Looping Based Amplification Cycle (MALBAC)" refers to a quasi-linear whole-genome amplification method (Zong et al., 2012, Genome-wide detection of single-nucleotide and copy-number variation of a single human cell, Science, Dec 21; 338(6114):1622-6, doi: 10.1126 / science.1229164.). MALBAC primers have an 8-nucleotide 3' random sequence that hybridizes to the template, and a 27-nucleotide 5' common sequence (GTG AGT GAT GGT TGAGGT AGT GTGGAG). After the first extension, the half-amplifier is used as a template for another extension, producing a complete amplicon with complementary 5' and 3' ends. After several cycles of quasi-linear amplification, the complete amplicon can be amplified exponentially in subsequent PCR cycles.

[0066] The term “oligonucleotide” or “oligo” refers to an oligomeric molecule that contains a nucleotide sequence, such as non-limiting examples of deoxyribonucleic acid (DNA), ribonucleic acid (RNA), locked nucleic acid (LNA), and peptide nucleic acid (PNA).

[0067] “Tagged oligonucleotide” or “tagged oligonucleotide” refers to an oligonucleotide molecule (e.g., an ssDNA molecule) that is directly conjugated to a binding agent (e.g., a primary antibody). Tagged oligonucleotides are used to indirectly quantify target molecules (e.g., proteins) that act as ligands for binding agents.

[0068] The acronym "PE" refers to fluorophora erythrin.

[0069] "2% PFA" refers to a phosphate buffer solution containing 2% w / v paraformaldehyde.

[0070] "Primary WGA DNA library (pWGAlib)" refers to a DNA library obtained from a WGA reaction.

[0071] The term "re-amplification" or "re-amp" refers to a PCR reaction in which all or most of the primary WGA DNA library is further amplified.

[0072] "Residue" refers to the amino acid residues that exist in the polypeptide chain of a protein.

[0073] A "sequencing barcode" refers to a polynucleotide sequence that, when sequenced within a sequencer read, allows that read to be assigned to a specific sample associated with that barcode.

[0074] “UMI” or “Unique Molecular Identifier sequence” refers to a degenerate or partially degenerate (i.e., random or semi-random) oligonucleotide sequence that is actually unique for each ssDNA or dsDNA molecule.

[0075] "Universal WGA primers" or "WGA PCR primers" refer to additional oligonucleotides linked to each fragment produced by restriction enzyme action. Universal WGA primers are used for DRS-WGA, such as Amplil. TM WGA. Detailed Implementation

[0076] The method according to the present invention for whole-genome amplification and analysis of multiple target molecules in biological samples (including genomic DNA and target molecules) includes the following steps.

[0077] In step a), a biological sample is provided. The biological sample is preferably a single cell, but it can also be a sample containing several cells.

[0078] In step b), the biological sample is contacted with at least one binding agent that targets at least one target molecule and is conjugated with a labeled oligonucleotide so that - when at least one target molecule is present in the biological sample - at least one binding agent binds to at least one target molecule.

[0079] The conjugate is preferably selected from the group consisting of free antibodies or fragments thereof, aptamers, small molecules, peptides, and proteins. The target molecule is preferably selected from the group consisting of free proteins, peptides, glycoproteins, carbohydrates, lipids, and combinations thereof. More preferably, the conjugate is an antibody. The conjugate binds to a target molecule with a specific stoichiometry, such as a monoclonal antibody or enzyme substrate, or to a target molecule with an unspecified stoichiometry, such as a polyclonal antibody or small molecule. The former allows for better quantification of the target compared to the latter. The conjugate is chemically conjugated to a labeled oligonucleotide through covalent or non-covalent interactions. In the former case, both the oligonucleotide and the conjugate have reactive moieties capable of binding to each other. The stoichiometry of the oligonucleotide can be controlled during the conjugation process.

[0080] A non-limiting list of examples of binders / target molecules is reported in Table 1 below.

[0081] Table 1

[0082]

[0083] The oligonucleotide used as the marker is preferably an ssDNA or dsDNA molecule with a chemically modified 5' or 3' end. This modification is used for covalent conjugation with the associated binding agent.

[0084] Conjugates formed from labeled oligonucleotides bound to a binding agent can target both extracellular and intracellular epitopes. “External” and “internal” conjugates can be applied to biological samples as two separate mixtures containing the final staining concentration of the conjugate. First, the external mixture is applied to label the external epitope. Second, cells are permeabilized using a detergent or similar method, and the internal mixture is applied to the sample to label the internal epitope. Alternatively, the external and internal conjugates can be mixed together for a single-step staining process. The final staining concentration varies for each binding agent and must be determined experimentally.

[0085] The labeled oligonucleotide sequence is preferably shorter than 300 bases, more preferably shorter than 120 bases, to facilitate conjugation with the binding agent and reduce costs. In a preferred embodiment, the labeled oligonucleotide sequence is 60-80 nucleotides.

[0086] refer to Figure 1 A, the labeled oligonucleotides contain:

[0087] i) The nucleic acid payload sequence (PL), which includes a binding barcode sequence (BAB) and a unique molecular identifier sequence (UMI), and

[0088] ii) At least one first-labeled oligonucleotide amplification sequence (5-TOS) of nucleic acid.

[0089] The payload sequence contains the necessary information for target counting.

[0090] The Unique Molecular Identifier (UMI) sequence is preferably a degenerate or semi-degenerate sequence in the range of 10 to 30 nucleotides. Preferably, the UMI has a length of at least 10 bases, corresponding to a theoretical 4^10 = 1,048,576 different combinations, which is sufficient for most target molecules. For highly abundant target molecules, longer UMIs can be used to increase the possible combinations, such as 12 bases. Using semi-degenerate bases reduces the possible combinations, and the UMI length is preferably increased, for example, up to 20 or 30 bases. Semi-degenerate UMIs can be advantageously used to introduce reference points that can be used for readouts to realign sequences and prevent overestimation of the presence of different UMIs. The UMI sequence may be located at the 5' or 3' of the BAB. The UMI sequence is preferably located after the annealing site of the Read 1 sequencing primer to increase the complexity of the first sequencing base. This is advantageous for the Illumina sequencing platform because the initial sequencing step requires high complexity for cluster differentiation. The BAB is a fixed sequence for each conjugated molecule. BABs are designed to avoid features that could interfere with primary PCR amplification and sequencing steps, such as homopolymers, hairpins, and / or heteroduplex formation [Frank, DN, 2009, BARCRAWL and BARTAB: software tools for the design and implementation of barcoded primers for highly multiplexed DNA sequencing, BMC Bioinformatics, 10, 362. http: / / doi.org / 10.1186 / 1471-2105-10-362], and are selected from a pool of all possible BAB sequences of a defined length to maximize their relative Hamming distance, thus minimizing the likelihood of any PCR or sequencing errors leading to incorrect sequencing read allocations. The BAB length must be chosen based on the number of target molecules to be detected. Preferably, the BAB has a length of at least 10 nucleotides, corresponding to a theoretical 4^10 = 1,048,576 different combinations, reduced to approximately 2,000 possible BAB sequences after applying a filter to the GC content (e.g., [30%...70%]), and is free of homopolymers, hairpins, and minimum Hamming distance (preferably ≥3 nt).

[0091] The first labeled oligonucleotide amplification sequence (5-TOS) is located at the 5' end of the labeled oligonucleotide. This sequence is essential for labeled oligonucleotide amplification and subsequent library generation. Labeled oligonucleotide amplification is necessary to avoid any bias due to molecule loss when processing samples that may interfere with accurate UMI counting.

[0092] refer to Figure 1 B, the target oligonucleotide preferably also comprises at least one second labeled oligonucleotide amplification sequence (3-TOS). This sequence is located at the 3' end of the labeled oligonucleotide. This sequence is essential for labeled oligonucleotide amplification and subsequent library generation. Labeled oligonucleotide amplification is necessary to avoid any bias due to molecule loss when processing samples that may interfere with correct UMI counting.

[0093] In a preferred embodiment, the 5-TOS and 3-TOS sequences are designed to avoid the formation of hairpins and other intramolecularly stable secondary structures within the amplification temperature range, as these could hinder the amplification of the labeled oligonucleotides. Figure 2 The diagram shows a labeled oligonucleotide library representing oligonucleotides from two different labeled oligonucleotides, one of which readily forms intramolecular hairpins between 5-TOS and 3-TOS at different temperatures (“hairpin-equipped”: Tm = 67℃, SEQ ID NO: 50). =-11.15kcal / mol; "No hairpin": Tm=45℃, SEQ ID NO: 51, =-1.52kcal / mol). The x-axis represents the expected UMI count, and the y-axis represents the UMI count after sequencing.

[0094] The labeled oligonucleotides and their amplification primers were optimized to achieve high sensitivity, a wide dynamic range, balanced PCR amplification, and reproducibility.

[0095] In a preferred embodiment of the present invention, the UMI sequence length (n = 10) is selected to quantize 0-10. 6 Targets within the molecular range.

[0096] The dynamic range is characterized by the amplification of labeled oligonucleotides at concentrations spanning four orders of magnitude (from 10^6 to 10^6). 2 Up to 10 6 (molecules). Figure 3 This graph shows the amplification of different amounts of oligonucleotide mixtures across 27 PCR cycles. Each dilution was performed as a triplet of three independent dilutions. The x-axis represents the expected number of different UMIs, and the y-axis represents the experimentally observed UMIs. A range of 10° between the observed UMIs and the expected UMIs before amplification can be observed. 2 Up to 10 6High linear correlation within the molecular range.

[0097] Balanced PCR amplification is characterized by performing different amplification cycles on the same starting sample. For example... Figure 4 As shown in A and 4B, amplification of the same labeled oligonucleotide pools with different total PCR cycle numbers (23 and 27 PCR cycles, respectively) did not lead to the observed differences in UMI, indicating that the PCR cycle number does not affect UMI counting.

[0098] The sensitivity is characterized by amplifying labeled oligonucleotide pools in varying numbers (down to 40 molecules). For example... Figure 4 As shown in C, quantization down to 10 is possible. 2 The labeled oligonucleotide molecules were analyzed. It should be noted that serially diluted solutions are prone to sampling bias due to the highly non-uniform distribution of molecules within the volume, especially with very diluted solutions. Therefore, the observed quantitation limitations may be an underestimation related to the experimental setup rather than an analytical limitation.

[0099] In step c) of the method according to the invention, a separation step is performed to selectively remove unbound binders, thereby obtaining a labeled biological sample. The separation step is typically performed by washing in a suitable buffer solution and collecting the labeled biological sample by centrifugation.

[0100] In step d), the labeled biological sample undergoes simultaneous whole-genome amplification of the genomic DNA and amplification of labeled oligonucleotides conjugated with at least one binding agent. Whole-genome amplification of the genomic DNA is performed via deterministic restriction-site whole genome amplification (DRS-WGA) or via multiple annealing and looping based amplification cycles (MALBAC).

[0101] In step e), a massively parallel sequencing library is prepared from the amplified labeled oligonucleotides.

[0102] In step f), the massively parallel sequencing library is sequenced.

[0103] In step g), the binding barcode sequence (BAB) and unique molecular identifier sequence (UMI) are retrieved from each sequencing read.

[0104] In step h), the number of distinct unique molecular identifier (UMI) sequences for each binder is calculated.

[0105] Steps e), f), g), and h) will be described in more detail in the following description with reference to the specific implementation.

[0106] The above method preferably further includes the step of isolating single cells from the biological sample. This separation can be achieved by sorting the cells, particularly using a cell sorter such as the DEPArray® NxT (Menarini Silicon Biosystems SpA), or alternatively by separating the cells into droplets. The separation step is preferably performed after step c) and before step d).

[0107] The above method preferably includes a step of purifying the massively parallel sequencing library before step f).

[0108] In more specific terminology, but not intended to limit the scope of this specification, the methods disclosed above are also referred to as Ampli l protein (A1-P), allowing for the quantification of proteins and whole-genome genetic characterization in single cells. The quantification of single or multiple proteins in a single cell is achieved using a set of binding agents (specifically antibodies (Abs)) conjugated to labeled oligonucleotides. These oligonucleotides are designed to explicitly identify the conjugated antibody via a DNA barcode sequence and to quantify the abundance of the epitope of interest via a random or partially degenerate sequence (i.e., a unique molecular identifier (UMI) for epitope quantification). Biological samples are labeled with one or more Ab-oligonucleotide conjugates, each carrying a unique DNA barcode sequence. Subsequently, the protein can be quantified in various ways (i.e., DEPArray). TM The NxT system is used to isolate single cells or cell pools, and then amplifies the entire genome (i.e., Amplil) to achieve this. TM The whole genome amplification kit amplifies its genomic contents. Immediately after the subsequent steps, labeled oligonucleotides are pre-amplified to avoid any downsampling during NGS library preparation. Specific primers, i.e., "library primers," are used to generate the NGS (Illumina) library ready for sequencing. The labeled oligonucleotides are designed with Amplil... TM WGA (Al-WGA) workflow compatible, allowing single-cell genetic analysis (e.g., Amplil) to be performed simultaneously with protein quantification using Al-P. TM LowPass).

[0109] The following discloses three specific embodiments of the present invention, which respectively use different numbers of primers for labeled oligonucleotide amplification and whole genome amplification.

[0110] In the first preferred embodiment, reference Figure 5The labeled oligonucleotides from 5' to 3' contain at least:

[0111] a) The first marker oligonucleotide amplification sequence of the nucleic acid (5-TOS) contains, in sequence, the 5' whole genome amplification handle sequence (5-WGAH) and the first amplification handle sequence (1AH);

[0112] b) Payload sequence (PL);

[0113] c) The second-labeled oligonucleotide amplification sequence of the nucleic acid (3-TOS), which contains the second amplification handle sequence of the nucleic acid (2AH) and the 3' whole genome amplification handle sequence (3-WGAH).

[0114] 3-WGAH is the reverse complementary sequence of 5-WGAH, enabling simultaneous amplification of gDNA and labeled oligonucleotides during whole-genome amplification. 1AH and 2AH are located at the 5' and 3' ends of the payload sequence, respectively, for subsequent library generation. 1AH and 2AH are preferably designed to avoid stable intramolecular secondary structures, such as hairpin structures, which could inhibit the amplification of labeled oligonucleotides. A fixed sequence may be present between each of the aforementioned sequences.

[0115] Whole genome amplification and amplification of labeled oligonucleotides are preferably performed using a single primer.

[0116] In the second preferred embodiment, refer to Figure 6 A and 6B, the labeled oligonucleotides from 5' to 3' contain at least:

[0117] a) The first marker oligonucleotide amplification sequence of the nucleic acid (5-TOS), which contains, in sequence, the 5' whole genome amplification handle sequence (5-WGAH) and the first amplification handle sequence (1AH);

[0118] b) Payload sequence (PL);

[0119] c) Optional, annealing sequence (AS).

[0120] At least one primer is used for whole-genome amplification and amplification of labeled oligonucleotides, and at least one oligonucleotide (Ep) is used for the extension of labeled oligonucleotides, said at least one oligonucleotide (Ep) comprising at least the following from 5' to 3':

[0121] d) 5' whole genome amplified handle sequence (5-WGAH);

[0122] e) Interval sequences (SS);

[0123] f) The second amplified handle sequence (2AH); and

[0124] g) A sequence that is inversely complementary to the annealing sequence (AS-RC) or a sequence that is inversely complementary to the binder barcode sequence (BAB-RC).

[0125] In other words, the amplification of the labeled oligonucleotide is achieved by annealing Ep to the AS located at the 3' of the labeled oligonucleotide. Figure 6 A) This occurs via the annealing sequence-reverse complementary sequence (AS-RC) located at the 3' end of Ep, thus causing the labeled oligonucleotide and the 3' extension of Ep in the reaction, resulting in the WGA primer amplifiable molecule. Alternatively, AS can be identical to the BAB sequence, and Ep can anneal to the BAB sequence via the BAB reverse complementary sequence (BAB-RC), as shown below. Figure 6 As shown in B. The advantage of the first option (annealing to AS) is that a single Ep can be used with any BAB, thereby reducing manufacturing costs and protocol complexity. The second option (annealing to BAB) can be advantageously used to normalize signals from targets with large abundance differences. As a non-limiting example, this can be achieved by using a limited number of primers for potentially high-abundance targets, or by reducing the amplification of high-abundance labeled oligonucleotides by using different BAB annealing temperatures. The annealing temperature can be adjusted by the BAB length and / or composition.

[0126] Following the extension of the labeled oligonucleotide and Ep, WGA primers amplify the labeled oligonucleotide within the resulting larger molecule. The spacer sequence (SS) increases the length of the amplicons generated by the labeled oligonucleotide. This increased fragment length disrupts the stability of intramolecular secondary structures, such as hairpins induced by complementary ends of the fragments, thereby lowering their melting temperature. Figure 7 This facilitates the amplification of labeled oligonucleotides along with other WGA fragments (M. Zuker. Mfold web server for nucleic acid folding and hybridization prediction, Nucleic Acids Res., 31(13), 3406-15, (2003)). The length of the extended oligonucleotide (Ep) sequence is preferably in the range of 60-300 bases. More preferably, the length of the extended oligonucleotide (Ep) sequence is in the range of 120-200 bases.

[0127] In the third preferred embodiment, reference Figure 8 The labeled oligonucleotides from 5' to 3' contain at least:

[0128] a) The first labeled oxo oligonucleotide amplification sequence (5-TOS) of the nucleic acid corresponding to the first amplification handle sequence (1AH);

[0129] b) Payload sequence (PL);

[0130] c) The second-labeled oligonucleotide amplification sequence (3-TOS) of the nucleic acid corresponding to the second amplification handle sequence (2AH).

[0131] At least one first primer is used for whole-genome amplification, and at least one second primer and at least one third primer are used for amplification of labeled oligonucleotides, wherein the at least one second primer has the same sequence as the first amplification handle sequence (1AH), and the at least one third primer has a sequence that is inversely complementary to the second amplification handle sequence (2AH-RC).

[0132] The labeled oligonucleotide amplification primers were designed with melting temperatures compatible with at least the first 10-15 cycles of the WGA PCR thermal profile.

[0133] Preferably, at least one second primer and at least one third primer are added in step d).

[0134] Step e) of preparing a massively parallel sequencing library from amplified labeled oligonucleotides is preferably performed by PCR using at least one first library primer and at least one second library primer, wherein the at least one first library primer contains a 3' sequence corresponding to a first amplification handle sequence (1AH) and the at least one second library primer contains a 3' sequence corresponding to a reverse complementary sequence (2AH-RC) to the second amplification handle sequence.

[0135] Therefore, library primers are advantageously used to generate NGS libraries for binding agent quantification analysis in a single PCR step. Specific examples of library primers described in this specification are for generating libraries compatible with the Illumina sequencing platform and are not intended to limit the scope of the invention.

[0136] In a preferred embodiment, the forward and reverse primers are designed based on Illumina adaptors and comprise 5' to 3' of:

[0137] 1) Illumina adaptor sequence (IA): Required for Illumina sequencing;

[0138] - Index sequencing primer / flow cell binding sequence: This region is essential for flow cell binding, as well as the annealing sequence for i5 / i7 index sequencing primers;

[0139] -i5 / i7 index: Index used for NGS multiplexing reactions;

[0140] - Read 1 / Read 2 sequencing primers: primers for Illumina sequencing, and annealing sequences for amplifying libraries from labeled oligonucleotides.

[0141] 2) The first amplification handle sequence (1AH) or the sequence that is the reverse complement of the second amplification handle sequence (2AH-RC): These sequences are annealed to the reverse complement of the first amplification handle sequence and the second amplification handle sequence on labeled oligonucleotides, respectively, and these sequences are double-stranded after amplification of the labeled oligonucleotides.

[0142] Figure 9 Shown in accordance with Figure 5 ( Figure 9 A) Figure 6 A ( Figure 9 B) and Figure 8 ( Figure 9 In the implementation of C), the structures of library primers used to generate massively parallel sequencing libraries from amplified labeled oligonucleotides are described.

[0143] After library generation, at least one purification step is preferably performed, followed by library quantization and pooling, which are necessary for subsequent sequencing steps. Sequencing is preferably performed as paired-end sequencing, generating two reads, each from one strand of the library DNA molecule.

[0144] The same method can also be used to generate NGS libraries for other sequencing platforms, such as Ion Torrent.

[0145] The analysis of paired end-read sequences generated from the NGS library is based on the following steps:

[0146] 1. Subsequence extraction. Extract subsequences corresponding to the UMI, BAB, and amplification handle sequences (1AH and / or 2AH) from two sequencing reads of each tagged oligonucleotide molecule.

[0147] 2. Read Realignment. If a subsequence of the BAB and / or amplification handle sequence does not match the reference sequence (tolerance of 0.5-2 mismatches per 5 bases), the subsequence position is shifted by a variable number, ranging from -n to +n, where n is the maximum allowed shift (e.g., n=8), and the subsequence is re-extracted. For each iteration, the Hamming distance to the reference sequence is calculated, and the shift that returns the minimum distance is selected. All subsequences (UMI, BAB, amplification handle sequence) are then extracted.

[0148] 3. Read filtering. Read the BAB and / or handle subsequences; reads that differ from the reference sequence by more than a defined number of bases are discarded as low-quality reads.

[0149] 4. UMI Determination. In the absence of sequencing errors, the UMI sequences from the read pair are expected to be perfectly complementary. Any differences between the first-strand UMI and the second-strand complementary UMI sequences will be considered if:

[0150] a. Reads may be discarded due to low confidence levels or

[0151] b. Consistency between two sequences can be calculated by selecting bases for each position of the UMI in the sequences from the two sequencing reads, where the bases have the highest base call score reported by the sequencer base caller.

[0152] It should be noted that the first method (a) is less likely to overestimate the target molecule due to sequencing bias, but may lose the true binding event between the binder and the target molecule, which can be recovered using the second method (b).

[0153] 5. Target molecule quantification. Target molecules are counted by determining the number of distinct UMI sequences representing each BAB sequence of a specific binder in the analytical sample.

[0154] The kit according to the present invention comprises:

[0155] a) At least one binding agent for at least one target molecule in a biological sample conjugated to a labeled oligonucleotide, said labeled oligonucleotide comprising:

[0156] i) The nucleic acid payload sequence (PL), which includes a binding barcode sequence (BAB) and a unique molecular identifier sequence (UMI).

[0157] ii) At least one first-labeled oligonucleotide amplification sequence (5-TOS);

[0158] (b) at least one primer for whole-genome amplification and at least one primer for amplifying labeled oligonucleotides, wherein the at least one primer for whole-genome amplification and the at least one primer for amplifying labeled oligonucleotides have the same sequence.

[0159] In a preferred embodiment, the kit contains oligonucleotides for extension labeling.

[0160] In a preferred embodiment, at least one labeled oligonucleotide has a sequence corresponding to SEQ ID NO:1, and two primers are used for amplifying the labeled oligonucleotide, having sequences SEQ ID NO:2 and SEQ ID NO:3, respectively. The kit further preferably comprises one or more first library primers and one or more second library primers. More preferably, the first library primer has a sequence selected from the group consisting of SEQ ID NO:8 to SEQ ID NO:15, and the second library primer has a sequence selected from the group consisting of SEQ ID NO:16 to SEQ ID NO:27.

[0161] In another preferred embodiment, at least one labeled oligonucleotide has a sequence corresponding to SEQ ID NO:28, and the oligonucleotide used for extending the labeled oligonucleotide is one and has a sequence corresponding to SEQ ID NO:29. The kit preferably further comprises one or more first library primers and one or more second library primers. More preferably, the first library primer has a sequence selected from the group consisting of SEQ ID NO:30 to SEQ ID NO:37, and the second library primer has a sequence selected from the group consisting of SEQ ID NO:38 to SEQ ID NO:49.

[0162] Example

[0163] Example 1

[0164] In this embodiment, the labeled oligonucleotides are designed to be used after WGA according to Figure 8 The amplification was performed using the specified settings. The labeled oligonucleotide was named "P5-Synth" (SEQ ID NO:1). Figure 10 ,NNNNNNNNNN: UMI sequence), designed to be similar to Amplil TM WGA kit (Menarini Silicon Biosystems) compatible. The first amplification handle sequence (1AH) is identical to the last 19 bases of the index 2 (i5) adaptor from the Illumina TruSeq DNA and RNA CD index. The second amplification handle sequence (2AH) was generated in a computer (in silico) to avoid any intramolecular secondary structures and possible matches on the human genome. The melting temperatures of both amplification handle sequences were designed to be similar to the melting temperatures of the WGA primers. Labeled oligonucleotide amplification primers (SEQ ID NO: 2 and SEQ ID NO: 3) were designed based on the first and second amplification handle sequences.

[0165] like Figure 10 As shown, the forward library primers are identical to the Illumina linker in index 2 (i5), while the reverse library primers are identical to the Illumina linker in index 1 (i7), with the addition of the reverse complementary sequence of the second amplification handle sequence. More detailed... Figure 10 The design of P5-Synth oligonucleotides and corresponding library primers is shown.

[0166] Oligonucleotide P5-Synth (SEQ ID NO:1): The white box indicates the internal structural domain with UMI and binding barcode. The gray box with a black border indicates the first and second amplification handle sequences.

[0167] P5 library primer (SEQ ID NO:8): Forward primer used for NGS library generation. The gray box shows the annealing site with the oligonucleotide P5-Synth. The short dashed box shows the sequencing primer site; the dotted-line box shows the i5 index used in multiplexed sequencing reactions; the long dashed box shows the index sequencing primer / flow-through-cell adaptor sequence.

[0168] Synthetic library primers (SEQ ID NO:16): Reverse primers used for NGS library generation. The gray box shows the annealing site with the oligonucleotide P5-Synth. The short dashed box shows the sequencing primer site; the dotted-dash box shows the i5 index used in multiplexed sequencing reactions; the long dashed box shows the index sequencing primer / flow-through-cell adaptor sequence.

[0169] like Figure 11 As shown, during library generation, Illumina index 1 and 2 adaptors are added to the 5' and 3' of the labeled oligonucleotides, respectively.

[0170] In this embodiment, the labeled oligonucleotide was conjugated via a 5' amino modifier. A C6 or C12 spacer region exists between the amine moiety and the 5' of the oligonucleotide to avoid any steric hindrance affecting subsequent PCR reactions. Using amines typically found in antibodies derived from lysine, glutamine, arginine, and asparagine residues, the antibody was covalently bound to the labeled oligonucleotide using an amino-reactive reagent. Four antibody-oligonucleotides were generated using the labeled oligonucleotides (Table 2), and antibody-oligonucleotide conjugation was performed by Expedeon Ltd (25 Norman Way, Over, Cambridge CB24 5QE, United Kingdom) with an antibody:labeled oligonucleotide stoichiometric ratio of 1:2. Epitope localization: Indicates the position relative to the cell membrane.

[0171] Table 2

[0172]

[0173] Ab-oligonucleotides were used to stain two different cell lines. The first cell type was SK-BR-3 cells, a cell line derived from breast tumors, which overexpress cytokeratin and Her2 protein. The second cell type was peripheral blood mononuclear cells (PBMCs), leukocytes extracted from whole blood, which express CD45 and negligible levels of cytokeratin and Her2.

[0174] SK-BR-3 cells (ATCC) HTB-30™ (ATCC) cells were grown in culture medium according to the manufacturer's procedures. PBMCs were extracted from human blood samples. Both cell types were fixed using 2% PFA according to a custom protocol.

[0175] Ab-oligonucleotide staining was performed on 100,000–50,000 previously fixed and infiltrated cells. Cells were collected by centrifugation at 1000 x g for 5 min at RT. Cells were washed with at least 1 mL of running buffer (autoMACS running buffer, reference 130-091-221, Miltenyi Biotec) and collected by centrifugation. This last step was repeated twice. The external Ab-oligonucleotide and its isotype Ab-oligonucleotide control were diluted to their working concentrations with 100 µl of running buffer. The external Ab-oligonucleotide mixture (Ab-oligonucleotide labeled 3) was added to the cells and incubated at RT for 15 min. The sample was then washed twice with 1 mL of running buffer and collected by centrifugation. 500 µl of running buffer was added to the goat anti-mouse IgG2a-PE antibody and incubated at +4 °C for 30 min. This step enabled staining of PBMC cells in PE. The sample was washed twice with running buffer. Dilute the internal Ab-oligonucleotides and their isotype Ab-oligonucleotide controls to their working concentrations using 200 µl of internal Perm buffer (Inside Stain Kit, reference 130-090-477, Miltenyi Biotec). Add the internal Ab-oligonucleotide mixture (Ab-oligonucleotide labels 1, 2, and 4) to the cells and incubate at RT for 10 min. Wash the samples twice with 1 mL of internal Perm buffer and collect by centrifugation. Add 500 µl of internal Perm buffer to the Hoechst and goat anti-mouse IgG1-APC antibodies and incubate at +4 °C for 30 min. This step stains PE and SK-BR-3 cells in all nuclei. Wash the samples twice with run buffer.

[0176] Adding a secondary antibody conjugated to a fluorophore enabled fluorescence recognition of SK-BR-3 cells (APC channel) and PBMCs (PE channel). Furthermore, the fluorescence level reflected the relative abundance of the Ab-oligonucleotide. This was achieved using DEPArray. TM The NxT system (MenariniSilicon Biosystems) is based on its immunofluorescence labeling for the purification of single cells. Figure 12 ).

[0177] Specifically, Figure 12 Scatter plots of PBMCs and SK-BR-3 cells stained with Ab-oligonucleotides and secondary fluorescent antibodies are shown. On the x-axis, fluorescence levels in the APC channels are proportional to the number of Ab-oligonucleotide markers 1, 2, and 4. On the y-axis, fluorescence levels in the PE channels are proportional to the number of Ab-oligonucleotide marker 3. The scatter plots are divided into four quadrants based on their immunofluorescence levels, each quadrant containing a specific cell type: 1) PBMCs (high levels of CD45 and low levels of CK and Her2); 2) Double-positive cells (high levels of CD45, CK, and Her2); 3) Double-negative cells (low levels of CD45, CK, and Her2); 4) SK-BR-3 cells (low levels of CD45 and high levels of CK and Her2). Single cells highlighted with empty / filled squares / circles were isolated and used for library generation.

[0178] Alternatively, to perform labeled oligonucleotide amplification after WGA, prepare forward / reverse primers and Amplil according to Table 3 (left insert). TM Custom reaction mixture for the PCR kit reagents. Add 15 µl of the reaction mixture to each tube containing the WGA product. Incubate each sample according to the heat distribution shown in Table 3 (right insert).

[0179] Table 3

[0180]

[0181] Left insert: Composition of the reaction mixture for labeled oligonucleotide amplification and translation. Right insert: Specific cyclic program for labeled oligonucleotide amplification.

[0182] Library preparation was performed using 1 µL aliquots of WGA containing Ab-oligonucleotide amplicon amplified using the Amplil™ PCR kit and with P5 and Lib1 library primers at a final concentration of 0.5 µM. PCR thermal cycling curves are shown in Table 5. Each sample had a different combination of NGS library primers for dual indexing to allow for data multiplexing during bioinformatics analysis. A list of library primers used is provided in Table 4. The P5 library primers are forward primers used to label the oligonucleotide P5-Synth.

[0183] Table 4

[0184]

[0185]

[0186]

[0187]

[0188] Table 5

[0189]

[0190] Thermal cycling is used for NGS library generation. The number of cycles in step 3 depends on the number of recovered cells and the effective amount of total Ab-oligonucleotides in the cells. Typically, 27 amplification cycles in stage 3 result in a sufficient number of amplicones per cell.

[0191] Library samples were purified using Agencourt Ampre XP beads (Beckman-Coulter). KAPA SYBR was used. NGS DNA quantification was performed using a FAST qPCR kit (KAPA Biosystems). Each NGS library was examined using an Agilent Bioanalyzer 2100 (Agilent), and the library electrophoresis pattern is shown, which typically consists of a single peak at 185 bp. Figure 13 ).

[0192] Samples were pooled and sequenced on a MiSeq system (Illumina) using the MiSeq Reagent kit v3 150-cycle (reference MS-102-3001, Illumina). Data analysis was performed using custom software developed in Python. Figure 14The report describes the quantification of protein targets based on UMI counts. As expected, SK-BR-3 cells exhibited high expression of cytokeratin and Her2, and low levels of CD45, while PBMCs showed the opposite behavior. Higher protein expression levels were observed in double-positive cells, particularly the isotype control group, suggesting that these cells are more prone to nonspecific staining. Conversely, double-negative cells showed lower levels of all four targets.

[0193] Example 2

[0194] In this embodiment, the labeled oligonucleotides are designed to react during WGA according to Figure 8 Amplification was performed using the same settings. Except for the following, the experimental procedure was the same as in Example 1. The labeled oligonucleotide amplification primers were added directly to the primary PCR reaction mixture to a final concentration of 0.02 µM. Library generation and data analysis were performed as shown in Example 1.

[0195] Figure 15 The report describes the quantification of protein targets based on UMI counts. As expected, SK-BR-3 cells exhibited high expression of cytokeratin and Her2, and very low CD45 levels, while PBMCs showed the opposite behavior. Higher protein expression levels were observed in double-positive cells, particularly the isotype control group, suggesting that these cells are more prone to nonspecific staining. Conversely, double-negative cells showed lower levels of all four targets. Based on the results of Example 1, it can be inferred that amplification of labeled oligonucleotides is feasible during or after WGA. However, it must be noted that there was a significant difference in absolute UMI counts between the two procedures. When labeled oligonucleotide amplification was performed during WGA, the difference between the two cell types was more consistent with expectations regarding the CD45 target (which showed lower expression compared to CK).

[0196] Example 3

[0197] In this embodiment, the labeled oligonucleotide was added directly to a single cell. The labeled oligonucleotide, “P5-Lib1” (SEQ ID NO: 28), was designed for amplification using the Ampli1™ WGA kit (Menarini Silicon Biosystems). The labeled oligonucleotide amplification primers have the sequence SEQ ID NO: 29 (the forward and reverse primers are identical, sharing the sequence of the Ampli1 WGA Lib1 primer). Specifically, the 5'-WGA handle sequence is identical to the Lib1 WGA primer, while the 3'-WGA handle sequence is the reverse complementary sequence of the Lib1 WGA primer. The first amplification handle sequence is identical to the sequence described in Example 1. The second amplification handle sequence consists of the 3'-WGA handle sequence and an additional 5 bp sequence located at its 5' end (…). Figure 16 ).

[0198] To be more detailed, Figure 16 The design of P5-Synth oligonucleotides and corresponding library primers is shown.

[0199] Oligonucleotide P5-Lib1: The white-filled box with a thick border represents the internal structural domain with UMI and binding agent barcodes. The gray-filled box with a thick border represents the annealing positions of the two library primers. The gray box with a thin border is the WGA handle sequence (Lib1).

[0200] P5 library primers: The forward primer is used to generate the NGS library. In the gray-filled box, the annealing position is marked with the oligonucleotide P5-Lib1. The short dashed box indicates the sequencing primer site; the dotted-dash box indicates the i5 index used for multiplexing sequencing reactions; the long dashed box indicates the index sequencing primer / flow-through-cell adaptor sequence.

[0201] Lib1 library primers: Reverse primers are used to generate the NGS library. In the gray-filled box, the annealing site is marked with the oligonucleotide P5-Lib1: this sequence consists of a portion of the Lib1 reverse complementary sequence and a small tail (ACCAC) that allows annealing to proceed only to the 3' end of the oligonucleotide P5-Lib1. The short dashed box indicates the sequencing primer site; the dotted-dash box indicates the i5 index used for multiplexing sequencing reactions; the long dashed box indicates the indexed sequencing primer / flow-through-cell adaptor sequence.

[0202] The forward library primers are identical to the Illumina linker for index 2 (i5), while the reverse library primers are identical to the Illumina linker for index 1 (i7), with the addition of the reverse complementary sequence of the second amplification handle sequence. Figure 16 Therefore, during library generation, Illumina index 1 and 2 adaptors are added to the 5' and 3' of the labeled oligonucleotides, respectively. Figure 17 ).

[0203] SK-BR-3 cells (ATCC®HTB-30™, ATCC) were grown in culture medium according to the manufacturer's program and fixed with 2% PFA according to a custom protocol. Single cells were purified based on their morphology using the DEPArray™ NxT system (Menarini Silicon Biosystems). P5-Lib1 and P5-Synth-labeled oligonucleotides were added directly to tubes containing single cells. Different amounts of each oligonucleotide were added to each single cell, and Ampli1™ WGA was performed. Samples containing P5-Synth-labeled oligonucleotides were amplified as shown in Example 1.

[0204] The labeled oligonucleotide libraries were generated by taking 1 µL aliquots of WGA containing Ab-oligonucleotide amplicons amplified using the Amplil™ PCR kit and primers for the P5 and Lib1 libraries at a final concentration of 0.5 µM. PCR thermal cycling curves are shown in Table 5. Different NGS library primer combinations were used for dual indexing of each sample to allow for data multiplexing during bioinformatics analysis. Table 6 reports a list of the library primers used.

[0205] Table 6

[0206]

[0207]

[0208]

[0209] 10 µl aliquots of WGA samples were purified using SPRI beads (Beckmann Coulter) and then treated with the Ampli1™ LowPass kit to generate an NGS library for CNA analysis. Prior to the WGA procedure, the peak values ​​of labeled oligonucleotides in single cells did not affect downstream genetic analysis. Figure 18 To adapt Figure 5 and Figure 8 The labeled oligonucleotides designed according to the workflow shown do not affect or interfere with the WGA procedure. Furthermore, under both conditions, NGS libraries can still be obtained from the labeled oligonucleotides. Both labeled oligonucleotides can be correctly quantified, demonstrating the robustness of both methods and the design of the labeled oligonucleotides. Figure 19 ).

[0210] Advantages

[0211] The method of the present invention for whole-genome amplification and analysis of multiple target molecules in biological samples allows for simultaneous whole-genome copy number analysis / genome sequence and protein expression analysis on the same single cell.

[0212] The method of this invention enables whole-genome amplification of genomic DNA, facilitating further analysis such as whole-genome copy number analysis of gene arrays of interest via low-pass sequencing or targeted sequencing. It also allows for the detection and digital quantification of a wide range of proteins to single-cell resolution with very few samples, in cases where only a small number (or even a single) of circulating tumor cells (CTCs) are present. This is particularly advantageous for measuring each molecular type in different cells within genetically heterogeneous samples, where differences in genotype, phenotype, and environment can confound and completely prevent the correlation between genotype (copy number of sequence alterations) and phenotype (protein expression).

[0213] The method according to the invention surprisingly outperforms prior art in one or more of the following dimensions (given by way of non-limiting examples), performance previously considered unattainable by those skilled in the art:

[0214] • Proteins in a single cell are digitally quantified to hundreds of copies per cell.

[0215] • Because the process uses the inherent WGA, additional genetic material can be obtained in addition to the points mentioned above, for studying other characteristics of the single cell, and the single cell can be reliably reanalyzed for verification, which is impossible with the droplet-based method proposed by 10X Genomics.

[0216] The primary application of this method is in oncology, but it can also be applied to other fields, such as mosaic disorders, dermatology, or overgrowth phenotypes.

Claims

1. A method for whole-genome amplification and analysis of multiple target molecules in a biological sample, said biological sample comprising genomic DNA and target molecules, said method comprising the following steps: a) Provide biological samples; b) Contacting a biological sample with at least one binding agent, said at least one binding agent targeting at least one target molecule conjugated to a labeled oligonucleotide, said labeled oligonucleotide comprising: i) The nucleic acid payload sequence (PL), which includes a binding barcode sequence (BAB) and a unique molecular identifier sequence (UMI), and ii) At least one first-labeled oligonucleotide amplification sequence (5-TOS) of nucleic acid; This ensures that when at least one target molecule is present in a biological sample, at least one binder binds to at least one target molecule. c) Perform a separation step to selectively remove unbound binders, thereby obtaining labeled biological samples; d) Perform the following on the labeled biological samples: - The whole-genome amplification of the genomic DNA is achieved through: i) Deterministic restriction site whole-genome amplification (DRS-WGA), or ii) Multi-annealing and circularization-based cyclic genome amplification (MALBAC), and -Amplification of labeled oligonucleotides conjugated with at least one binding agent, The whole genome amplification and the amplification of labeled oligonucleotides are performed simultaneously; e) Prepare massively parallel sequencing libraries from amplified labeled oligonucleotides; f) Sequencing of massively parallel sequencing libraries; g) Retrieve the binding barcode sequence (BAB) and unique molecular identifier sequence (UMI) from each sequencing read; h) Calculate the number of distinct unique molecular identifier (UMI) sequences for each binder.

2. The method according to claim 1, wherein, The labeled oligonucleotide also contains at least one second labeled oligonucleotide amplification sequence (3-TOS).

3. The method according to claim 1 or 2, wherein, The unique molecular identifier sequence (UMI) is a degenerate or semi-degenerate sequence in the range of 10 to 30 nucleotides.

4. The method according to claim 1, wherein, The method also includes the step of isolating single cells from the biological sample.

5. The method according to claim 4, wherein, The separation step is performed by sorting cells.

6. The method according to claim 4, wherein, The separation step is performed by dividing the cells into droplets.

7. The method according to any one of claims 4 to 6, wherein, The separation step is performed after step c) and before step d).

8. The method of claim 1, further comprising the step of purifying the massively parallel sequencing library prior to step f).

9. The method according to claim 1, wherein, The at least one binding agent is an antibody or a fragment thereof.

10. The method according to claim 1, wherein, The at least one binder is an aptamer.

11. The method according to claim 1, wherein, The at least one binder is a small molecule.

12. The method according to claim 1, wherein, The at least one binding agent is a peptide.

13. The method according to claim 1, wherein, The at least one binding agent is a protein.

14. The method according to claim 1, wherein, The target molecule is a protein.

15. The method according to claim 1, wherein, The target molecule is a peptide.

16. The method according to claim 1, wherein, The target molecule is a glycoprotein.

17. The method according to claim 1, wherein, The target molecule is a carbohydrate.

18. The method according to claim 1, wherein, The target molecules are proteins and lipids.

19. The method according to claim 2, wherein, The labeled oligonucleotides from 5' to 3' contain at least: a) The first marker oligonucleotide amplification sequence of the nucleic acid (5-TOS), which contains, in sequence, the 5' whole genome amplification handle sequence (5-WGAH) and the first amplification handle sequence (1AH); b) Payload sequence (PL); c) The second-labeled oligonucleotide amplification sequence of the nucleic acid (3-TOS), which contains the second amplification handle sequence of the nucleic acid (2AH) and the 3' whole genome amplification handle sequence (3-WGAH).

20. The method according to claim 1, wherein, The whole genome amplification and the amplification of the labeled oligonucleotides were performed using a single primer.

21. The method according to claim 1, wherein, The labeled oligonucleotides from 5' to 3' contain at least: a) The first marker oligonucleotide amplification sequence of the nucleic acid (5-TOS), which contains, in sequence, the 5' whole genome amplification handle sequence (5-WGAH) and the first amplification handle sequence (1AH); b) Payload sequence (PL); c) Optional, annealing sequence (AS); At least one primer is used for whole-genome amplification and amplification of labeled oligonucleotides, and at least one oligonucleotide (Ep) is used for the extension of labeled oligonucleotides, said at least one oligonucleotide (Ep) comprising at least the following from 5' to 3': d) 5' whole genome amplified handle sequence (5-WGAH); e) Interval sequences (SS); f) The second amplified handle sequence (2AH); and g) A sequence that is inversely complementary to the annealing sequence (AS-RC) or a sequence that is inversely complementary to the binder barcode sequence (BAB-RC).

22. The method of claim 1, wherein the labeled oligonucleotide comprises at least the following from 5' to 3': a) The first labeled oligonucleotide amplification sequence (5-TOS) of the nucleic acid corresponding to the first amplification handle sequence (1AH); b) Payload sequence (PL); c) The oligonucleotide amplification sequence (3-TOS) of the second labeling nucleic acid corresponding to the second amplification handle sequence (2AH); At least one first primer is used for whole-genome amplification, and at least one second primer and at least one third primer are used for the amplification of labeled oligonucleotides; At least one second primer has the same sequence as the first amplification handle sequence (1AH), and at least one third primer has a sequence that is inversely complementary to the second amplification handle sequence (2AH-RC).

23. The method according to claim 22, wherein, In step d), add at least one second primer and at least one third primer.

24. The method according to claim 19, wherein, Step e) of preparing a massively parallel sequencing library from the amplified labeled oligonucleotides is performed by a PCR reaction using at least one first library primer and at least one second library primer, wherein the at least one first library primer contains a 3' sequence corresponding to a first amplification handle sequence (1AH), and the at least one second library primer contains a 3' sequence corresponding to a sequence that is inversely complementary to the second amplification handle sequence (2AH-RC).

25. A kit for carrying out the method of claim 1, comprising: a) at least one binding agent targeting at least one target molecule in a biological sample conjugated to a labeled oligonucleotide, said labeled oligonucleotide comprising: i) The nucleic acid payload sequence (PL), which includes a binding barcode sequence (BAB) and a unique molecular identifier sequence (UMI). ii) At least one first-labeled oligonucleotide amplification sequence (5-TOS); (b) at least one primer for whole-genome amplification and at least one primer for amplifying labeled oligonucleotides, wherein the at least one primer for whole-genome amplification and the at least one primer for amplifying labeled oligonucleotides have the same sequence.

26. The kit of claim 25 further comprises an oligonucleotide for extension labeling.

27. The kit according to claim 25, wherein, At least one labeled oligonucleotide has a sequence corresponding to SEQ ID NO: 1, and two primers are used for amplifying the labeled oligonucleotide, having sequences SEQ ID NO: 2 and SEQ ID NO: 3, respectively.

28. The kit according to any one of claims 25 to 27, further comprising one or more first library primers and one or more second library primers.

29. The kit according to claim 28, wherein, The first library primer has a sequence selected from the group consisting of SEQ ID NO: 8 to SEQ ID NO: 15, and the second library primer has a sequence selected from the group consisting of SEQ ID NO: 16 to SEQ ID NO:

27.

30. The kit according to claim 25, wherein, At least one labeled oligonucleotide has a sequence corresponding to SEQ ID NO: 28, and the primer used for amplifying the labeled oligonucleotide is one having a sequence corresponding to SEQ ID NO:

29.

31. The kit according to claim 30, further comprising one or more first library primers and one or more second library primers.

32. The kit according to claim 31, wherein, The first library primer has a sequence selected from the group consisting of SEQ ID NO: 30 to SEQ ID NO: 37, and the second library primer has a sequence selected from the group consisting of SEQ ID NO: 38 to SEQ ID NO: 49.

Citation Information

Patent Citations

  • DNA amplification of a single cell

    EP1109938A1

  • Methods and compositions for identifying or quantifying targets in a biological sample

    US20180251825A1

  • Furnace air-feeding mechanism.

    US865868A

  • Protein detection via nanoreporters

    US9714937B2

  • Method and kit for the generation of DNA libraries for massively parallel sequencing

    WO2017178655A1