Sequencing method
By embedding cells and markers in the droplets, the scaffolds prepared by DNA nanotechnology combine fluorescent dyes and coding oligonucleotides to generate sequencing constructs and perform DNA sequencing, the problem of inefficient sequencing of a large number of single cell target gene sequences in the prior art is solved, and high-throughput and efficient data processing and analysis are achieved.
Patent Information
- Application Number
- CN202380079633.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-18
- Filing Date
- 2023-11-14
- Publication Date
- 2025-06-27
AI Technical Summary
It is difficult for the prior art to efficiently sequence target gene sequences of a large number of individual cells, especially when analyzing high numbers of rare cells, existing methods have problems of high labor intensity and low data processing efficiency.
Sequencing constructs were generated by embedding cells and markers in droplets, combining fluorescent dyes and coding oligonucleotides on scaffolds prepared using DNA nanotechnology, and optical readings and data analysis were performed through DNA sequencing technology.
High-throughput sequencing of target gene sequences of multiple cells is achieved, data processing efficiency is improved, and the need for isolation and arrangement of individual cells is reduced. It can effectively identify and sort a large number of clones and retrieve target gene sequences of interest.
Smart Images

Figure CN120225460A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for sequencing target gene sequences of multiple cells. Background Art
[0002] Rare cells, such as adult stem cells, circulating tumor cells, and reactive immune cells (e.g., T cells, B cells, or NK cells reactive to a certain antigen), are of great interest to basic and translational researchers. For example, reactive immune cells (such as B cell clones) that respond to a certain pathogen (such as a virus) and can produce antibodies against the specific pathogen are of great value for generating much-needed therapeutic antibodies. Similarly, reactive T cells are also sought after in the context of personalized medicine and the treatment of cancer and other diseases. Once reactive T cells are identified and isolated, the gene sequences encoding the corresponding T cell receptors can be cloned and used to generate genetically engineered T cells, such as CAR-T cells, which show affinity for the target antigen. Similarly, circulating tumor cells are expected to be of great value in diagnosing cancer, predicting outcomes, managing therapies, and discovering new cancer drugs and cell-based therapies. Cell suspensions containing rare cells are usually obtained from tissue samples or liquid biopsies by separation methods. Therefore, the identification, analysis, and isolation of rare cells in these samples, specifically the analysis of these cells at the single-cell level (single-cell analysis, SCA), are of great value for basic research and translational research, diagnostic and therapeutic applications, and in the biological processes, development, and production of biologics and cell therapies.
[0003] With the expansion of the ability to identify and distinguish various cell types, the identification of cell types has become more refined, i.e., the rare cell populations of interest are smaller and better defined. Therefore, in order to discover rare cells of interest, it is usually necessary to analyze a high (<100,000), very high (<1,000,000), or ultra-high (>1,000,000) number of cells.
[0004] Specifically, there is extensive attention to techniques that can screen large populations of B cell or T cell clones to find clones that exhibit desired characteristics. In many cases, after such clones are identified, the corresponding genomic sequences encoding the variable portions of the antibodies or TCRs are cloned and sequenced. Various methods have been developed to screen single cells or clones and are used to screen B cell / T cell clones. These methods include limited dilution and manual screening in microtiter plates, nanopore imaging methods, microfluidic systems, and other methods based on microfluidic technology and imaging in flow-through.
[0005] While some of these methods are inherently very limited in throughput, a key limitation common to all of them is that cloning genomic sequences from a clone of interest requires physically separating that clone from other clones. This means that at some point in the workflow, all individual cells / clones must be arrayed into wells, nanopores, nanopens, etc., or they must be sorted from the remaining cells. However, it is estimated that ideally, each desired antibody or TCR will be screened at a magnitude of 10 trillion clones, as this number seems to be adequately related to the actual diversity present in the immune repertoire. Sorting or arraying 10 trillion clones per project is prohibitively labor-intensive, which is why most current screening efforts sample a significantly smaller number of clones.
[0006] Accordingly, it is desirable to be able to track individual cells and assign analytical data from different types of analyses to specific cells within a large mixture of cells. Summary of the Invention
[0007] Accordingly, an object of the present application is to provide a method capable of sequencing target gene sequences of a large number of individual cells with high throughput.
[0008] The above object is achieved by the subject matter of the independent claims. Advantageous embodiments are defined in the dependent claims and the following description.
[0009] There is provided a method for sequencing target gene sequences of a plurality of cells, comprising the steps of: providing droplets, each droplet containing at least one of the plurality of cells and at least one optically detectable marker. The marker comprises: a nucleic acid backbone having a plurality of attachment sites at predetermined positions; a plurality of markers for attachment to at least some of the attachment sites; at least a first orientation indicator and a second orientation indicator; wherein each marker comprises: at least one dye, particularly a fluorescent dye; a coding oligonucleotide portion configured to uniquely encode the characteristics of the at least one dye; and an attachment oligonucleotide portion configured to reversibly attach to one of the attachment sites; and the attachment oligonucleotide portion of each marker comprises a unique oligonucleotide sequence configured to bind to a complementary sequence of one of the attachment sites. The method further comprises the steps of: releasing the target gene sequence from the cells; ligating the target gene sequence to the marker of the marker to generate a sequencing construct; and sequencing the sequencing construct.
[0010] Specifically, when connecting the target gene sequence to the label of the marker, the target gene sequence is ligated to the oligonucleotide of the marker. Specifically, the coding oligonucleotide portion of the marker of the marker is ligated to the attachment oligonucleotide portion. Thus, each of the sequencing constructs comprises: at least one gene sequence in the target gene sequence, and the coding oligonucleotide portion and the attachment oligonucleotide portion of one marker in the marker of the marker.
[0011] The arrangement of dyes or dye combinations on scaffolds prepared using structural DNA nanotechnology can achieve the purpose that, through combinatorial encoding, a large number of dye combinations can be generated using a limited amount of dyes. Then, the combinatorial encoding can be combined with the spatial encoding on the DNA nanostructure scaffold, thereby enabling the dyes or dye combinations to be precisely arranged on the scaffold. Such an arrangement can be optically read by reading the markers and orientation indicators, and can also be optically read using DNA sequencing by reading / sequencing the sequence stretches of the attached oligonucleotide part - coding oligonucleotide part, which link the corresponding dyes to the corresponding markers. In this way, a large number of markers can be generated, which: (A) allow for the efficient particle indexing (identification) of a large number of droplets, (B) allow for the optical reading of the indexing, (C) allow for the reading of the indexing based on DNA sequencing, and (D) allow for the physical linkage of a part or the entire indexing (the concatenate of all the attached oligonucleotide part - coding oligonucleotide part sequence stretches required to reliably identify a given marker) to a gene target sequence of interest, thereby successively (E) enabling the sequencing of the particle index - gene target sequence - fusion product, which (F) can easily retrieve the gene target sequence of interest from the clone of interest by sequencing the concatenate. In this case, using imaging - based screening and assays, such as antigen binding, antibody aggregation, specific clone productivity, cytokine secretion, the activity of a reporter gene in target cells, killing assays, growth kinetics assays, etc., the clone of interest can be easily identified. Specifically, the particle indexing provided by the markers allows for the reliable identification of the same clone during imaging and the reliable three - dimensional orientation from one time point to the next in time - series acquisition or sequential assays. In this way, a large number of clones can be identified and ranked according to multiple relevant criteria. The method also allows for the efficient retrieval of the gene target sequence of interest from these prioritized clones. At this point, two alternative exemplary methods can be provided. The first method is based on ligating a part of the sequence stretches required to read the complete index to a given target gene sequence to provide the sequencing construct. In this case, by sequencing a sufficient number of ligates, the complete index / code is read to read the complete index, and the sequencing must be performed at the single - droplet level.
[0012] In the second alternative exemplary method, the complete concatenate between all the coding oligonucleotide part - attached oligonucleotide part - sequence stretches belonging to a given marker in a given droplet is ligated (concatenated) to the target gene sequence to provide the sequencing construct.
[0013] Such target gene sequences can be single or multiple sequences. In a particularly preferred embodiment, the target gene sequence encodes an antibody or a T cell receptor, or a portion thereof (e.g., a region encoding a framework, variable region, hypervariable region, complementarity determining region).
[0014] Details of suitable markers are disclosed in the patent application with the application number EP 22153210.4, the content of which is incorporated herein by reference.
[0015] The orientation indicator is configured to attach to the scaffold and can be, for example, a fluorescent dye. The orientation indicator can be used to visually determine the orientation of the marker in space. The coding oligonucleotide portion is a unique sequence of nucleic acid that is unique for the excitation / emission wavelength and / or fluorescence lifetime of the at least one dye of the marker. Thus, through the unique combination of dyes in each marker and their specific attachment sites, the marker can be visually and unambiguously identified. By sequencing the unique coding oligonucleotide portion and the attachment oligonucleotide portion, the same marker can be unambiguously identified.
[0016] By ligating the target gene sequence to the oligonucleotide of the marker, specifically, to the coding oligonucleotide portion and the attachment oligonucleotide portion of the marker of the marker, the target gene sequence is physically linked to the sequence information, thereby enabling the marker to be unambiguously identified and thus linking the target gene sequence to the marker.
[0017] Preferably, the nucleic acid scaffold comprises a scaffold strand and staple strands, the staple strands being configured to bind to the scaffold strand at predetermined positions to cause the scaffold strand to fold into a predetermined shape. The scaffold can be DNA origami. The sizes of these DNA origami structures can range from a few nanometers to micrometers. To fabricate such DNA origami-based structures, a longer DNA molecule (scaffold strand) is folded at positions precisely identified by the so-called staple strands. The DNA origami can be designed to provide a self-assembling scaffold with a specific predetermined shape. This enables easy and reproducible synthesis and assembly of the scaffold. The staple strands can be site-selectively functionalized. In this case, the positional resolution is limited by the nucleotide size (in the nanometer or sub-nanometer range). The prior art has utilized this to generate fluorescent standards, where fluorescent dyes are linked to precisely positioned bands on the DNA origami. These standards are called "nanorulers" and are used for calibration of imaging systems such as confocal microscopy or super-resolution microscopy (e.g., STED), as disclosed, for example, in US2014 / 0057805 A1.
[0018] The DNA origami provides a scaffold for the marker. Preferably, the DNA origami structure comprises at least one scaffold strand and multiple staple strands, wherein the staple strands are complementary to at least a portion of the scaffold strand and are configured to transform the scaffold strand into a predetermined conformation. Specifically, the strands are oligonucleotides. This enables the generation of a scaffold with a predetermined two-dimensional or three-dimensional shape that can self-assemble. Additionally, this enables the specific placement of attachment sites on the scaffold.
[0019] The attachment site is preferably a unique nucleic acid sequence of the staple strand. Preferably, at a predetermined attachment site, the marker can attach to the staple strand of the scaffold. Since the staple strands are located at predetermined positions, the position of the attachment site can also be predetermined. Thus, the attachment site (the attachment site of the nucleic acid scaffold) is a unique oligonucleotide sequence that is complementary to the attachment oligonucleotide portion of a marker.
[0020] Alternatively, the nucleic acid scaffold can comprise or consist of DNA bricks. These DNA bricks are oligonucleotides of shorter length that have overlapping hybridization segments to bind to each other and assemble the scaffold. In this case, the oligonucleotide of the marker can form a component of the nucleic acid scaffold. Attachment sites are on the oligonucleotides of shorter length that bind to the DNA bricks to generate the nucleic acid scaffold.
[0021] Preferably, prior to sequencing, in a specific alternative, prior to ligation or prior to amplification, the marker is released from the nucleic acid scaffold of the marker. This enables particularly robust sequencing.
[0022] Preferably, prior to sequencing, the target gene sequence and the oligonucleotide portion of the marker are amplified, specifically, by adding primers and performing PCR. Specifically, this step is performed after releasing the target gene sequence. The amplification can further include introducing a restriction site or a ligation site in the amplification product. This enables particularly robust sequencing.
[0023] Particularly preferably, the marker of the marker comprises a cleavage site that is configured to separate the dye from the oligonucleotide of the marker, and wherein prior to ligation, or alternatively prior to amplification, the dye of each marker is excised from the oligonucleotide of the marker at the cleavage site. This enables particularly robust sequencing.
[0024] Preferably, the droplets are liquid droplets, specifically, liquid droplets having a discrete volume, liquid droplets (liquid droplets) of a first liquid (or first fluid) in an immiscible second liquid (or second fluid), or solid droplets (solid droplets)). This enables the processing of the plurality of cells in the same volume while maintaining the association of the marker with the corresponding cells within the droplets.
[0025] The liquid droplets can be based on droplet technologies of water-in-oil emulsions. Examples in this regard include microfluidically generated droplets, such as the Bio-Rad droplet digital PCR technology.
[0026] The solid droplets can comprise polymeric compounds, particularly hydrogels. This enables the particularly easy handling of the cells and the marker within the droplets. For other examples of hydrogel droplets or beads containing cells, which include general methods for imaging droplets containing cells, reference is made to patent applications PCT / EP2021 / 058785 and PCT / EP2021 / 061754, the contents of which are hereby incorporated by reference in their entirety.
[0027] Furthermore, document WO 2019 / 028166 A1 discloses hydrogel beads suitable for sequencing gene sequences within the beads.
[0028] Preferably, the step of providing the droplets comprises embedding the cells and the marker within the droplets, preferably by means of a microfluidic device. This enables the particularly effective embedding of the cells and the marker within the droplets.
[0029] Preferably, the cells are cultured within the droplets, particularly before releasing the target gene sequence from the cells. This enables the analysis of the cells during or after a period of time.
[0030] Preferably, an optical readout of at least one of the droplets having the at least one cell and the at least one marker is obtained. Specifically, this is obtained before releasing the target gene sequence. This enables the visual analysis of the cells within the droplets and the identification of the marker.
[0031] The optical readout can be an image-based readout, which can be obtained on a microscope (e.g., point-scanning confocal or a camera- / wide-field imaging system such as a spinning disk microscope, light sheet fluorescence microscope, light field microscope, stereomicroscope). Additionally, the optical readout can be a non-image-based readout, e.g., in a cell counter or a flow-based readout device having at least one point detector or line detector. The readout can consist of discrete readouts, e.g., a single acquisition of an emission spectrum or an image stack; the readout can be a readout data stream, e.g., which is substantially continuous in a point-scanning confocal or a cell counter. Further, the readout can be a series of images, e.g., a spectral or hyperspectral image stack, where each image records the fluorescence emission in a different wavelength band.
[0032] The optical readout can be generated by a readout device for performing fluorescence multicolor reading or imaging. The readout device generally includes at least one excitation light source, a detection system including at least one detection channel, and can further include filters and / or dispersive optical elements such as prisms and / or gratings to transmit the excitation light to the sample and / or transmit the emission light from the sample to the detector or an appropriate area of the detector. The detection system can include multiple detection channels, which can be spectral detectors for parallel detection of multiple spectral bands or hyperspectral detectors for detecting a continuous portion of the spectrum. The detection system includes at least one detector, which can be a point detector (e.g., photomultiplier tube, avalanche diode, hybrid detector), an array detector, a camera, a hyperspectral camera. The detection system can record the intensity of each channel as is typically the case in a cell counter, or can be an imaging detection system that records images as is the case in a plate reader or a microscope. A readout device having one detection channel, e.g., a camera or a photomultiplier tube, can generate a readout having multiple detection channels using, e.g., different excitation and emission bands.
[0033] Particularly preferably, in the optical readout of the at least one droplet in the droplets, at least one dye of each marker based on the marker is used to determine the at least one marker in the droplets. This enables the identification of the marker based on the dye of the marker.
[0034] Preferably, in the sequencing data generated during sequencing of the sequencing construct in at least one of the droplets, the sequences of the attachment oligonucleotide portion and the coding oligonucleotide portion are determined, and wherein, based on the presence of the sequences in the sequencing data and the presence of the marker in the at least one droplet, the sequencing data of the at least one droplet is associated with the optical readout of the at least one droplet. This can effectively link the optical readout with the sequencing data. Thus, for example, the phenotype identified in the optical readout can be linked to the specific genotype identified in the sequencing data.
[0035] Preferably, the target gene sequence is generated by reverse transcription of the mRNA of the at least one cell. This enables sequencing of particularly diverse target gene sequences.
[0036] Preferably, the plurality of cells are immune cells, and wherein the target gene sequence is an immunoglobulin gene sequence, particularly a VDJ sequence. This can effectively identify, for example, the immunoglobulin gene sequences of a large number of immune cells.
[0037] Preferably, each droplet further contains at least one non-immune cell, and wherein the at least one immune cell is specific for the antigen of the non-immune cell. This can effectively identify immunoglobulin gene sequences that are specific for the antigen of the non-immune cell.
[0038] Preferably, the ligation step includes ligating all of the target gene sequences of the cells in one of the droplets together to the label of the marker in one of the droplets to generate the sequencing construct. Specifically, the oligonucleotides that ligate all of the target gene sequences to the label of the marker, particularly the coding oligonucleotide portion and the attachment oligonucleotide portion that ligate to the label of the marker. This can robustly generate the sequencing construct. In addition, this can particularly effectively sequence the sequencing construct without separating the individual droplets before sequencing. Specifically, this can achieve a particularly robust correlation between the sequencing data of the at least one droplet and the optical readout of the at least one droplet.
[0039] Preferably, the sequencing constructs in all of the droplets are pooled and sequenced together. This can particularly effectively sequence the sequencing construct without separating the individual droplets before sequencing. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Specific embodiments are described below with reference to the accompanying drawings, wherein:
[0041] Figure 1 Several nucleic acid backbones with orientation indicators are shown,
[0042] Figure 2 Examples of markers for attachment sites for attachment to a backbone are shown,
[0043] Figure 3 Various rod-shaped markers are shown,
[0044] Figure 4 Details of the rod-shaped markers are shown,
[0045] Figure 5 Steps for preparing markers for sequencing are shown,
[0046] Figure 6 Cube-shaped markers are shown,
[0047] Figure 7 Droplets having cells and markers are shown,
[0048] Figure 8 A flowchart of a method for sequencing a target gene sequence is shown,
[0049] Figure 9 Details of steps for generating random sequencing constructs are shown,
[0050] Figure 10 Details of steps for generating a predetermined sequencing construct are shown;
[0051] Figure 11 Hypothetical examples of database entries for optically detectable markers are shown; and
[0052] Figure 12 The number of possible different marker combinations by using 10 different fluorescent dyes for rod-shaped markers is illustrated. DETAILED DESCRIPTION
[0053] Figure 1 Nucleic acid backbones 100, 102, 104, 106, 108 having different geometries are schematically shown. Generally, the backbones 100, 102, 104, 106, 108 comprise nucleic acids. Specifically, the backbones 100, 102, 104, 106, 108 are based on DNA origami, which allows the generation of predetermined, stable two-dimensional and three-dimensional shapes. In addition, this allows the generation of multiple attachment sites at predetermined positions along the backbones 100, 102, 104, 106, 108. The attachment sites are oligonucleotide segments that are unique and allow complementary oligonucleotides to hybridize, for example, to attach markers to the backbones 100, 102, 104, 106, 108 at these predetermined positions.
[0054] The backbone 100 is linear or rod-shaped. It contains a first orientation indicator 110 and a second orientation indicator 112. The orientation indicators 110, 112 can be used to determine the orientation, directionality, or polarity of the backbone 100. The orientation indicators 110, 112 can contain dyes, especially fluorescent dyes such as fluorescein or fluorescent proteins. In addition, the dye of the first orientation indicator 110 has different characteristics from the dye of the second orientation indicator 112. These characteristics can include fluorescence emission characteristics, excitation characteristics, or lifetime characteristics. This enables the distinction between the first and second orientation indicators 110, 112 in the optical readout of the backbone 100 (e.g., generated by a microscope, cell counter, or imaging cell counter). The orientation indicators 110, 112 are arranged at intervals from each other. Preferably, each orientation indicator 110, 112 is arranged at opposite ends of the backbone 100. Thus, the first and second orientation indicators 110, 112 can distinguish the first end and the second end of the backbone 100. Ultimately, this enables the determination of the orientation, directionality, or polarity of the backbone 100, e.g., from the first orientation indicator 110 at the first end to the second orientation indicator 112 at the second end.
[0055] The backbone 102 is sheet-shaped, which can be a large linear DNA molecule or an assembly of multiple DNA molecules. The sheet-shaped backbone can greatly increase the number of available attachment sites. To be able to determine the orientation of the backbone 102, a third orientation indicator 114 is provided.
[0056] More geometries are possible, such as a tetrahedral backbone 104, a cubic backbone 106, or a polyhedral backbone 108. These backbones can contain a fourth orientation indicator 116 to determine their orientation.
[0057] Figure 2 A marker 200 for the attachment sites for attachment to the backbones 100, 102, 104, 106, 108 is schematically shown. The marker 200 contains a plurality of fluorescent dyes 202a, 202b, 202c, 202d, 202e. The fluorescence characteristics of the fluorescent dyes 202a to 202e can be different, such as excitation wavelength, emission wavelength, and fluorescence lifetime characteristics. Preferably, the dyes 202a to 202e of the marker 200 can be individually identified in a specific readout of the marker 200. Depending on the number of dyes used when generating the marker 200, a certain number of unique markers can be generated. Generally, the total number of different dyes used can be, for example, in the range of 5 - 50. For 20 dyes and using 5 dyes for each marker, it is possible to easily generate a set of markers with 15504 unique markers.
[0058] Dyes 202a to 202e are each attached to the marker carrier 204. The marker carrier 204 can be an oligonucleotide, and each of the dyes 202a to 202e can be specifically attached to the marker carrier 204 via a unique hybridization moiety 206a, 206b, 206c, 206d, 206e. In addition, as described in more detail below, one end 208 of the marker carrier 204 can be specifically attached to the scaffolds 100 to 108. A cleavage site 210 can be provided to cleave the marker carrier 204. This enables the removal of dyes 202a to 202e from the end 208 of the marker carrier 204, for example when the marker 200 is attached to one of the scaffolds 100 to 108. As described with respect to the marker 200, the orientation indicators 110, 112, 114, 116 preferably have the same or a similar structure.
[0059] Figure 3 Schematically shown are various rod-shaped markers 300, 302, 304, 306, 308. Each of the markers 300 to 308 includes a first orientation indicator 310 and a second orientation indicator 312 at a first attachment site and a second attachment site of the scaffold 314, respectively. In addition, there are another ten attachment sites, one of the attachment sites being designated by the reference numeral 316. In the case of the marker 308, markers 318, 320, 322 are attached at the other three attachment sites 316. The attachment sites 316 of each scaffold 314 are unique such that the markers 318, 320, 322 can be specifically attached to a specific attachment site 316. The length of the markers 300, 302, 304, 306, 308 is preferably between 1 and 2 μm.
[0060] The orientation indicators 310, 312 generate a relative coordinate system for the markers 300, 302, 304, 306, 308, and each attachment site 316 can be located on this coordinate system. For example, an index n can be assigned to each attachment site 316 based on the unique position of the corresponding attachment site 316, where n = 1, 2, 3,.... Thus, the different markers 300 to 308 can be visually distinguished because they use markers with different characteristics and / or markers that are attached (or not attached) to different, distinguishable attachment sites 316 along the scaffold 314.
[0061] Figure 4The details of the rod-shaped marker 308 are schematically shown, specifically, the marker 320 and its attachment to the backbone 314. At a specific attachment site 400 on the backbone 314, the marker 320 is attached to the backbone 314. The attachment site 400 has a unique oligonucleotide sequence that allows hybridization with the complementary attachment oligonucleotide portion 402 of the marker 320. Thus, each attachment site of the marker 308 has a unique oligonucleotide sequence that can specifically target the marker for each attachment site.
[0062] In addition, the marker 320 contains a coding oligonucleotide portion 404. The coding portion 404 is an oligonucleotide sequence unique to one or more fluorescent dyes 408 of the marker 320. This means that the dye of a specific marker can be identified by the sequence of the coding portion.
[0063] The marker 320 also contains a cleavage site 406 for removing the dye 408 from the coding portion 404 and the attachment portion 402.
[0064] Thus, each marker has: a unique sequence for a specific dye (such as the coding portion 404), and another unique sequence (such as the attachment portion 402) that hybridizes to a specific one of the attachment sites on the backbone.
[0065] The first and second orientation indicators 310, 312 are similarly constructed. For example, the orientation indicator 310 is attached to the attachment site 410 of the backbone 314 through a unique and complementary attachment oligonucleotide portion 412. In addition, the orientation indicator 310 contains a coding portion 414 unique to the dye 418 of the orientation indicator. By cleavage of the cleavage site 416, the dye 418 can be removed from the coding portion 414 and the attachment portion 412.
[0066] Figure 5 The steps for preparing the marker 320 for sequencing are shown, for example, to read the information of the attachment portion 402 and / or the coding portion 404. First, the marker 320 is removed from the backbone, for example, by heating to unwind the hybridized attachment portion 402 from the backbone. Then, the dye 408 is removed from the marker 320 by cleavage of the cleavage site 406 (for example, by enzymatic cleavage if the cleavage site 406 is a restriction site). Alternatively, the cleavage site 406 can be cleaved by light or temperature. The remaining attachment portion 402 and coding portion 404 can be separated (such as by chromatography) and then sequenced. For this purpose, a universal primer can be ligated to the remaining attachment portion 402 and coding portion 404. Alternatively, the universal primer can be provided together with the marker vector.
[0067] Figure 6An example of a cube-shaped marker 600 with a three-dimensional array scaffold 601 is schematically shown. The marker 600 includes three orientation indicators 602. In addition, the marker 600 also includes a set of different markers 604, 606, 608. Each marker 604, 606, 608 is attached to a specific attachment site of the marker 600. For example, the attachment site can be located at the corners or edges of the array of the scaffold 601. The markers 604, 606, 608 have different fluorescence properties, such as fluorescence lifetime, emission wavelength, and excitation wavelength.
[0068] Thus, similar to the markers 300 to 308 in Figure 3 , the marker 600 can be distinguished from other markers in optical readout by placing specific markers at specific attachment sites of the scaffold 601. Compared with the rod-shaped markers 300 to 308, this example of the cube-shaped marker 600 provides approximately 360 attachment sites for the markers, which increases the number of possible marker combinations and their positions on the scaffold 601. The physical size of the cube-shaped marker 600 is such that each side of the cube is in the range of 2.5 to 10 μm.
[0069] Figure 7 A cell 700 embedded in a droplet 702 is schematically shown. The droplet 702 can be a hydrogel bead. The droplet 702 also includes cube-shaped markers 704, 706. The markers 704, 706 can be structurally similar to the marker 600. However, the markers 704, 706 can be different from each other in terms of the specific implementation of the scaffold and / or the specific markers attached to the corresponding scaffold and / or the specific attachment sites to which the markers are attached.
[0070] Figure 8 A flowchart of an example of a method for sequencing the target gene sequences of multiple cells is shown. The cells are preferably immune cells, for example, containing target gene sequences encoding immunoglobulins.
[0071] In the first step S800 of the method, a plurality of droplets having cells and markers are provided. To provide or generate the droplets, the cells can be individually embedded in the droplets together with at least one marker. Preferably, each droplet embeds an immune cell having a unique target gene sequence. The droplets can be solid or solidified droplets, including hydrogels. These droplets can also be referred to as hydrogel beads. Alternatively, the droplets can be liquid droplets, such as water-oil emulsions.
[0072] This embedding can be carried out using a microfluidic device, which embeds the marker and the cells as the droplets are generated. Preferably, other cells, such as non-immune cells, can be included in the droplets, which serve as targets for the immune cells, enabling the immune cells to interact with the non-immune cells. After embedding, the droplets can be cultured together in a liquid medium in the same container.
[0073] In step S802, the cells in each droplet, especially the immune cells, are lysed to release their genetic contents, especially the target gene sequences. Optionally, the corresponding markers of the marker can be released from the backbone of the marker, especially the oligonucleotides of the marker.
[0074] In step S804, the target gene sequences and the oligonucleotides can be amplified by polymerase chain reaction (PCR). For this purpose, suitable primers are introduced into the droplets. This results in multiple copies of each of the oligonucleotides of the marker and the target gene sequences. The primers can contain restriction sites to be introduced into the amplification products.
[0075] In step S806, the amplified target gene sequences and the oligonucleotides of the marker are ligated together to generate a sequencing construct. This can be achieved by cutting the amplification products of step S804 with a restriction enzyme (suitable for the restriction sites introduced in step S804). Thus, the generated sticky ends can enable the amplification products to ligate to each other. For example, each target gene sequence and each oligonucleotide of the marker can have complementary restriction sites, such that a target gene sequence randomly binds to the labeled oligonucleotide. To ensure that a specific target gene sequence in a droplet can be unambiguously associated with the corresponding marker in the droplet, a representative number of random sequencing constructs must be generated and sequenced. This ensures that each combination of unique labeled oligonucleotides ligated to the copies of the target gene sequence is present in the random sequencing constructs.
[0076] For an alternative step to step S806, in step S808, a sequencing construct can be generated from the amplified target gene sequences and the labeled oligonucleotides. In step S808, the sequencing construct is generated such that all the oligonucleotides of the marker of the marker in a specific droplet are ligated to the target gene sequence together. Specifically, the ligation order can be predetermined. For example, certain restriction sites can be introduced during amplification in step S804, which allows all the oligonucleotides of the marker in the droplet to be combined with the target gene sequence in a certain order. This ensures that each predetermined sequencing construct carries all the coding parts and attachment parts of the marker in a specific droplet and the target gene sequence. This enables the unambiguous association of a specific target gene sequence with the corresponding marker in the droplet based on the combination of the coding parts and attachment parts in each predetermined sequencing construct.
[0077] In step S810, the sequencing construct is sequenced to generate sequencing data of the sequencing construct. The predetermined sequencing constructs generated in step S808 can be pooled for all droplets because each sequencing construct carries a unique combination of the coding part and the attachment part of the marker in a specific droplet as well as the target gene sequence. In contrast, the random sequencing constructs generated in step S806 need to be sequenced individually for each droplet because the random sequencing constructs only collectively carry the unique combination of the coding part and the attachment part of the marker in a specific droplet, rather than individually. The method ends in step S814.
[0078] The method may include other steps. For example, in an additional step before step S804, an optical readout of the cells and the markers in each droplet is generated, preferably an image or an image stack (e.g., generated with a microscope). The optical readout can be analyzed to determine the markers associated with each droplet. Each marker can be identified by a specific label attached to a specific attachment site. This includes the fluorescence properties of the dye of the label.
[0079] In an additional step after step S812, the sequences of the attachment oligonucleotide part and the coding part are determined to be present in the sequencing data. This enables the determination of the presence of the corresponding marker. Based on the sequences present in the droplet and the individual markers present, the sequencing data is associated with the cells in a specific droplet. For each marker, it is known which labels are attached to which attachment sites, and thus, in the optical readout, the markers present can be unambiguously identified. At the time of sequencing, by determining the presence of the corresponding coding oligonucleotide sequence and the attachment oligonucleotide sequence, the identity of the markers present in the sequencing sample can be similarly identified. By comparing these, the optical readout of the droplet with a specific marker can be associated with or assigned to the sequencing data of that droplet containing these specific markers. This enables the direct linking of the data on the phenotype in the optical readout to the data on the genotype in the sequencing data.
[0080] Figure 9 The detailed steps of generating the random sequencing construct 922 are shown. The oligonucleotide 900 of the label 902 of the marker 904 is embedded in the droplet 906 together with the cell 908. The oligonucleotide 900 and the target gene sequence 910 can be amplified by PCR. The oligonucleotide 900 contains the attachment part 402 and the coding part 404. The target gene sequence 910 may contain multiple unique sequences of interest 912, 914. Similarly, the marker 904 may contain multiple unique labels 902, and for simplicity, Figure 9 only one unique label 902 is shown.
[0081] The amplified tag 916 can contain a restriction site 918 introduced by a primer used for amplification. Similarly, the amplified target gene sequence 920 can also contain the restriction site 918. When digested with the corresponding restriction enzyme, the amplified tag 916 and the amplified target gene sequence 920 can be ligated to form a random sequencing construct 922. These combinations, which are one of the amplified tag 916 and the amplified target gene sequence 920, are exemplified in Figure 9 In the case where the marker 904 contains a plurality of unique markers 902, the random sequencing construct 922 can contain a plurality of combinations of the corresponding amplified tags and the amplified target gene sequence 920.
[0082] Figure 10 The detailed steps for generating a predetermined sequencing construct 1000 are shown. In comparison with Figure 9 The amplified tag 1002 is generated from a marker containing a plurality of unique markers. Thus, the amplified tag 1002 contains a plurality of unique attachment portions 402, 402a, 402b and a plurality of unique coding portions 404, 404a, 404b. Therefore, the amplified tag 1002 can be used to unambiguously identify the corresponding marker based on the combination of the portions 402, 402a, 402b, 404, 404a, 404b.
[0083] The amplified target gene sequence 1004 is exemplarily shown as having a copy of the unique sequence 912. To generate the predetermined sequencing construct 1000, copies of each amplified tag 1002 are ligated to copies of the amplified target gene sequence 1004. Specifically, this can be achieved by selecting a restriction site 918 for amplification such that the portions of the predetermined sequencing construct are ligated in a predetermined order.
[0084] Figure 11 A hypothetical example of a database entry for an optically detectable marker l is shown. The marker is incorporated into a given droplet or particle together with a single cell / clone and is used to assign a particle index l to the specific droplet or particle and the contained single cell / clone. In this example, the optically detectable marker has a combined marker, each containing two dyes, attached to positions p1, p10, p4, p3, which are addressed via corresponding attachment oligonucleotide portions. Each dye is linked to the marker via an oligonucleotide having at least a coding oligonucleotide portion. Thus, the dye species can be determined optically or by reading / sequencing the coding oligonucleotide portion. Depending on how the marker is designed, the position of the dye can be determined optically in association with an orientation marker, or by reading the attachment oligonucleotide portion physically linked to the coding oligonucleotide portion or a complementary sequence. Additionally, as Figure 11As shown, the clone l in droplet l marked by marker l has the target gene sequence shown in the table. This target gene sequence can be linked or joined to a part or the complete complement of the attached oligonucleotide portion - coding oligonucleotide portion.
[0085] In the sense of this document, reading or sequencing of the attached oligonucleotide portion or the coding oligonucleotide portion can be carried out on the sense strand or the complementary antisense strand. Thus, both the sense and the complementary antisense of the attached oligonucleotide portion or the coding oligonucleotide portion are meant in the sense of this document, as both are equally applicable for identifying a specific position on the marker or a specific dye of the marker.
[0086] Figure 12 Another example is provided where microscopic readout is used, which can distinguish a set of 10 ATTO dyes, where markers containing unique combinations of 2 out of these 10 dyes are generated, resulting in 45 unique combinations. A scaffold with 6 positions and two orientation markers is used, and these positions and orientation markers are spaced approximately 200 nm apart on a rod-shaped structure with a length of approximately 1600 nm. In the characteristics of the dyes (such as spectrum and lifetime), the combination of two-layer coding and spatial coding on the nanostructure allows for the generation of a very large number of different markers or codes. This means that a simple rod-shaped structure in the micron range can be used in combination with a limited number of fluorescent dyes to encode billions of particles.
[0087] In all the figures, elements having the same or similar functions are denoted by the same reference numerals. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items and can be abbreviated as " / ". All combinations between the various features of the embodiments and between the various features of the embodiments, as well as combinations with the various features or groups of features in the foregoing description and / or claims, are considered to be disclosed.
[0088] Although certain aspects have been described in the context of a device, it is clear that these aspects also represent a description of the corresponding method, where the module or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also represent a description of the corresponding module or item or feature of the corresponding device.
[0089] List of Reference Signs 100, 102, 104, 106, 108, 314, 601 Nucleic acid scaffold 110, 112, 114, 116, 310, 312, 602 Orientation indicator 200, 318, 320, 322, 604, 606, 608 and 902 markers 202a, 202b, 202c, 202d, 202e 408 and 418 fluorescent dyes 204 marker carrier 206a, 206b, 206c, 206d, 206e hybridization portions End of the 208 marker carrier 210, 406, 416 cleavage sites 300, 302, 304, 306, 308, 600 704, 706, 904 markers 316, 400, 410 attachment sites 402, 402a, 402b, 412 attachment oligonucleotide portions 404, 404a, 404b, 414 coding oligonucleotide portions 700, 908 cells 702, 906 droplets Oligonucleotide of the 900 marker 910 target gene sequence Unique sequences of the 912 and 914 target gene sequences 916 and 1002 amplified markers 918 restriction site 920 and 1004 amplified target gene sequences 922 random sequencing construct 1000 predetermined sequencing construct
Claims
1. A method for sequencing the target gene sequences (910) of multiple cells (700, 908), comprising the following steps: Providing droplets (702, 906), each droplet containing at least one of the multiple cells (700, 908) and at least one optically detectable marker (300, 302, 304, 306, 308, 600, 704, 706, 904), the marker comprising: A nucleic acid backbone (100, 102, 104, 106, 108, 314, 601) having a plurality of attachment sites (316, 400, 410) at predetermined positions; A plurality of markers (200, 318, 320, 322, 604, 606, 608, 902) for attachment to at least some of the attachment sites (316, 400, 410); At least a first orientation indicator and a second orientation indicator; Wherein each marker (200, 318, 320, 322, 604, 606, 608, 902) comprises: at least one dye (202a, 202b, 202c, 202d, 202e, 408, 418), an encoding oligonucleotide portion (404, 414) configured to encode the characteristics of the at least one dye (202a, 202b, 202c, 202d, 202e, 408, 418), and an attachment oligonucleotide portion (402, 412) configured to reversibly attach to one of the attachment sites (316, 400, 410), and Wherein the attachment oligonucleotide portion (402, 412) of each marker (200, 318, 320, 322, 604, 606, 608, 902) comprises a unique oligonucleotide sequence configured to bind to a complementary sequence of one of the attachment sites (316, 400, 410); Releasing the target gene sequence (910) from the cells (700, 908); Linking the target gene sequence (910) to the markers (200, 318, 320, 322, 604, 606, 608, 902) of the markers (300, 302, 304, 306, 308, 600, 704, 706, 904) to generate a sequencing construct (922, 1000); and Sequencing the sequencing construct (922, 1000).
2. The method according to claim 1, wherein prior to sequencing, the markers (200, 318, 320, 322, 604, 606, 608, 902) are released from the nucleic acid backbone (100, 102, 104, 106, 108, 314, 601) of the markers (300, 302, 304, 306, 308, 600, 704, 706, 904).
3. The method according to any one of the preceding claims, wherein prior to sequencing, the oligonucleotide (900) portions of the target gene sequence (910) and the markers (200, 318, 320, 322, 604, 606, 608, 902) are amplified.
4. The method according to any one of the preceding claims, wherein the markers (200, 318, 320, 322, 604, 606, 608, 902) of the markers (300, 302, 304, 306, 308, 600, 704, 706, 904) comprise cleavage sites (210, 406, 416) configured to separate the dyes (202a, 202b, 202c, 202d, 202e, 408, 418) from the oligonucleotide (900) of the marker (200, 318, 320, 322, 604, 606, 608, 902), and wherein prior to ligation, the dyes (202a, 202b, 202c, 202d, 202e, 408, 418) of each marker (200, 318, 320, 322, 604, 606, 608, 902) are excised from the oligonucleotide (900) of the marker at the cleavage sites (210, 406, 416).
5. The method according to any one of the preceding claims, wherein the droplets (702, 906) are liquid droplets of a first liquid in an immiscible second liquid, or solid droplets.
6. The method according to any one of the preceding claims, wherein the step of providing the droplets (702, 906) comprises embedding cells (700, 908) and markers (300, 302, 304, 306, 308, 600, 704, 706, 904) in the droplets (702, 906), preferably by means of a microfluidic device.
7. The method according to any one of the preceding claims, wherein the cells (700, 908) are cultured in the droplets (702, 906).
8. The method according to any one of the preceding claims, wherein an optical readout of at least one of the droplets (702, 906) having the at least one cell (700, 908) and the at least one marker (300, 302, 304, 306, 308, 600, 704, 706, 904) is obtained.
9. The method according to claim 8, wherein in the optical readout of the at least one droplet among the droplets (702, 906), the at least one marker (300, 302, 304, 306, 308, 600, 704, 706, 904) in the droplets (702, 906) is determined based on at least one dye (202a, 202b, 202c, 202d, 202e, 408, 418) of each marker (200, 318, 320, 322, 604, 606, 608, 902) of the marker (300, 302, 304, 306, 308, 600, 704, 706, 904).
10. The method according to claim 8 or 9, wherein in the sequencing data generated during sequencing of the sequencing construct (922, 1000) in at least one of the droplets (702, 906), the sequences of the attachment oligonucleotide portion (402, 412) and the coding oligonucleotide portion (404, 414) are determined, and wherein, Based on the presence of the sequence in the sequencing data and the presence of the marker (300, 302, 304, 306, 308, 600, 704, 706, 904) in the at least one droplet (702, 906), the sequencing data of the at least one droplet (702, 906) is associated with the optical readout of the at least one droplet (702, 906).
11. The method according to one of the preceding claims, wherein the target gene sequence (910) is generated by reverse transcription of mRNA of the at least one cell (700, 908).
12. The method according to one of the preceding claims, wherein the plurality of cells (700, 908) are immune cells, and wherein the target gene sequence (910) is an immunoglobulin gene sequence.
13. The method according to claim 12, wherein each droplet (702, 906) further comprises at least one non-immune cell, and wherein the at least one immune cell is specific for an antigen of the non-immune cell.
14. The method according to one of the preceding claims, wherein the ligation step comprises ligating all the target gene sequences (910) of the cell (700, 908) in one droplet among the droplets (702, 906) together to the marker (200, 318, 320, 322, 604, 606, 608, 902) of the marker (300, 302, 304, 306, 308, 600, 704, 706, 904) in one droplet among the droplets (702, 906) to generate the sequencing construct (922, 1000).
15. The method according to claim 14, wherein the sequencing constructs (922, 1000) in all the droplets (702, 906) are pooled and sequenced together.
Citation Information
Patent Citations
DNA-origami-based standard
US20140057805A1
Hydrogel beads for nucleotide sequencing
WO2019028166A1