A method and apparatus for sequence-verified gene synthesis and assembly

The method enriches target sequences using barcode-identified selection probes to purify high-fidelity nucleic acid molecules, addressing high error rates in DNA synthesis and enabling affordable, accurate longer DNA segment production.

WO2026003843A1PCT designated stage Publication Date: 2026-01-02SHILO SHAY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/IL2025/050551
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-26
Filing Date
2025-06-26
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Current DNA synthesis methods face high error rates, leading to costly production of short, low-accuracy DNA segments, while long, accurate segments are expensive, hindering high-throughput and affordable synthesis of longer DNA segments.

Method used

A method involving sequencing reads with barcode sequences to enrich target sequences using selection probes, allowing for the purification of highly error-free nucleic acid molecules by identifying and selecting probe-bound molecules with high fidelity.

Benefits of technology

Enables the production of high-quality, uniform, and cost-effective nucleic acid molecules with low error rates, facilitating the synthesis of longer DNA segments with high accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IL2025050551_02012026_PF_FP_ABST
    Figure IL2025050551_02012026_PF_FP_ABST
Patent Text Reader

Abstract

Methods of enriching a target sequence, comprising receiving sequencing reads derived from a plurality of nucleic acid molecule, wherein nucleic acid molecules of the plurality each comprise a barcode sequence; identifying in the sequencing reads a barcode sequence present in nucleic acid molecule comprising the target sequence; contacting the plurality of nucleic acid molecule with a selection probe that hybridized to the identified barcode sequence and comprising a selection moiety; and selecting a probe-bound target nucleic acid molecule by selecting the selection moiety are provided. Systems for performing the methods of the invention are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

A METHOD AND APPARATUS FOR SEQUENCE-VERIFIED GENE SYNTHESISAND ASSEMBEYCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 664,230, filed June 26, 2024, the content of which is incorporated herein by reference in its entirety.FIELD OF INVENTION

[0002] The present invention is in the field of DNA synthesis and purification.BACKGROUND OF THE INVENTION

[0003] Synthetic biology is an emerging field that uses genes, gene variants, genetic pathways, and other biological components in diverse applications such as health, agriculture, and bioproduction. High-throughput DNA synthesis is a necessity for the synthetic biology revolution. The relatively high error rate found in current DNA oligos prevents them from being used as-is for gene synthesis and this impacts cost. Synthesis costs and the length of the synthesized DNA segment are closely coupled. Low-cost, low-accuracy DNA segments will be rather short. Long accurate segments are expensive. There is a great need to synthesize longer DNA segments at an affordable cost with high throughput and accuracy.SUMMARY OF THE INVENTION

[0004] The present invention provides methods of enriching a target sequence, comprising receiving sequencing reads derived from a plurality of nucleic acid molecule, wherein nucleic acid molecules of the plurality each comprise a barcode sequence; identifying in the sequencing reads a barcode sequence present in nucleic acid molecule comprising the target sequence; contacting the plurality of nucleic acid molecule with a selection probe that hybridized to the identified barcode sequence and comprising a selection moiety; and selecting a probe-bound target nucleic acid molecule by selecting the selection moiety are provided. Systems for performing the methods of the invention are also provided.

[0005] According to a first aspect, there is provided a method of enriching a target sequence, the method comprising: a. receiving sequencing reads derived from a plurality of nucleic acid molecules, wherein nucleic acid molecules of the plurality each comprise abarcode sequence, and wherein the plurality comprises a nucleic acid molecule comprising the target sequence and a nucleic acid molecule devoid of the target sequence; b. identifying in the sequencing reads a barcode sequence that is present only in a nucleic acid molecule of the plurality that comprises the target sequence; c. contacting the plurality of nucleic acid molecules with a selection probe under conditions and for a sufficient time to hybridize the selection probe to the identified barcode sequence to produce a probe-bound target nucleic acid molecule, wherein the selection probe comprises: i. a sequence complementary to the identified barcode sequence; and ii. a selection moiety; d. selecting from the plurality of nucleic acid molecule the probe-bound target nucleic acid molecule by selecting the selection moiety; thereby enriching a target sequence.

[0006] According to some embodiments, the method is a method of enriching a sequence- verified target sequence and wherein the nucleic acid molecules devoid of the target sequence comprise the target sequence with errors.

[0007] According to some embodiments, a sequence verified target sequence is a target sequence with an error rate below a predetermined threshold and the target sequence with errors is a target sequence with an error rate above a predetermined threshold.

[0008] According to some embodiments, the method comprises before step (a): a. providing a plurality of solid supports, wherein each solid support is conjugated to a plurality of nucleic acid molecules, wherein each nucleic acid molecule comprises: i. a barcode sequence specific to the solid support and not shared by other solid supports of the plurality; and ii. a capturing sequence; b. providing a pool of nucleic acid molecules comprising a capture sequence, wherein the provided pool comprises a nucleic acid molecule comprising the target sequence and a nucleic acid molecule devoid of the target sequence, and wherein the capturing sequence and the capture sequence are complementary sequences capable of hybridizing together; c. contacting the provided pool of nucleic acid molecules to the provided plurality of solid supports under conditions and sufficient time to hybridizethe capture sequence to the capturing sequence to produce at least one hybridized solid support; d. contacting the plurality of hybridized solid supports with a polymerase to elongate ends of nucleic acid molecules to produce a plurality of elongated solid supports; e. dissociating elongated nucleic acid molecules comprising the capture sequence from elongated nucleic acid molecule comprising the capturing sequence and conjugated to the solid support; and f. sequencing the dissociated elongated nucleic acid molecule comprising the capture sequence to produce the sequencing reads.

[0009] According to some embodiments, the contacting the plurality of nucleic acid molecules with a selection probe is contacting the elongated nucleic acid molecules comprising the capturing sequence and conjugated to the solid support with the selection probe.

[0010] According to some embodiments, the solid supports are beads.[Oi l] According to some embodiments, the method further comprises amplifying the elongated nucleic acid molecules to produce copies of the elongated nucleic acid molecules and wherein the sequencing further comprises sequencing the copies.

[0012] According to some embodiments, the plurality of nucleic acid molecules is between 100-10,00,000 nucleic acid molecules per solid support.

[0013] According to some embodiments, the plurality of nucleic acid molecules are covalently linked to the solid support.

[0014] According to some embodiments, the plurality of nucleic acid molecules are conjugated to the solid support at a 5’ end of each nucleic acid molecule of the plurality.

[0015] According to some embodiments, each nucleic acid molecule of the plurality of nucleic acid molecules is conjugated to the solid support at a 5’ end and wherein the barcode sequence is 5’ to the capturing sequence, optionally, wherein each nucleic acid molecule comprises a unique molecular identifier (UMI) sequence and wherein the UMI sequence is 5’ to the capturing sequence.

[0016] According to some embodiments, the method comprises before step (a): a. providing a pool of nucleic acid molecules comprising a nucleic acid molecule comprising the target sequence and a nucleic acid molecule devoid of the target sequence, and wherein each oligonucleotide of the poolcomprises a barcode sequence that is specific to the nucleic acid molecule and is not shared by other nucleic acid molecules of the pool; b. amplifying the pool of nucleic acid molecules to produce copies of the nucleic acid molecules; and c. sequencing at least a portion of the copies to produce the sequencing reads.

[0017] According to some embodiments, the contacting the plurality of nucleic acid molecules with a selection probe is contacting the pool of nucleic acid molecules, contacting a portion of the copies that were not sequenced or both.

[0018] According to some embodiments, each nucleic acid molecule of the plurality of nucleic acid molecules further comprises a UMI sequence.

[0019] According to some embodiments, the identifying comprises identifying a plurality of continuous nucleotide sequences that comprise the target sequence and a barcode sequence and wherein step (c) comprises contacting the plurality of nucleic acid molecules with a plurality of selection probes, wherein the plurality of selection probes comprises a probe to each identified barcode sequence.

[0020] According to some embodiments, the selection moiety is a detectable moiety or isolatable moiety selected from a fluorophore, a radioactive moiety, a nanopore-detectable moiety, biotin, a click reaction substrate, an antigen, a magnetic moiety and a bead.

[0021] According to some embodiments, the selection moiety is a fluorophore.

[0022] According to some embodiments, the enriching comprises fluorescence activated cell sorting (FACS).

[0023] According to some embodiments, the plurality of nucleic acid molecules comprising a nucleic acid molecule comprising a target sequence comprises fragments of the target sequence that when assembled produce the target sequence.

[0024] According to some embodiments, the identifying comprises identifying a barcode sequence that is present in continuous nucleotide sequences that comprise all of the fragments needed to assemble the target sequence.

[0025] According to some embodiments, the method further comprises after the selecting, assembling selected target sequence fragments into a complete target sequence.

[0026] According to some embodiments, the target sequence is a sequence of a gene of interest.

[0027] According to another aspect, there is provided a system comprising: a. a thermomixer; b. a DNA sequencer;c. a control unit to analyze DNA sequencing results from the DNA sequencer, and select a barcode sequence or its complement sequence to be part of a selection probe; and d. a computer program product comprising a non-transitory computer-readable storage medium having program code embodied thereon, the program code executable by at least one hardware processor to apply the selection probe to the thermomixer.

[0028] According to some embodiments, the computer program product comprises program code executable by at least one hardware processor to perform the method of the invention.

[0029] According to some embodiments, the system further comprises a sorting device capable of sorting solid supports comprising the selection probe from solid supports devoid of the selection probe.

[0030] According to some embodiments, the sorting device is a FACS machine.

[0031] According to some embodiments, the system further comprises an oligo or DNA molecule printer capable of producing the selection probe.

[0032] Further embodiments and the full scope of applicability of the present invention will become apparent from the detailed description given hereinafter. However, it should be understood that the detailed description and specific examples, while indicating preferred embodiments of the invention, are given by way of illustration only, since various changes and modifications within the spirit and scope of the invention will become apparent to those skilled in the art from this detailed description.BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1: A schematic diagram showing beads structure.

[0034] Figure 2: A schematic diagram showing FACS machine operation principle in the context of the disclosed invention.

[0035] Figure 3: A schematic diagram showing applying oligos over beads.

[0036] Figure 4: A schematic diagram of the steps of an embodiment of a method of the invention.

[0037] Figure 5 : A schematic description of gene synthesis main steps.

[0038] Figure 6: A schematic block diagram of a device in line with the disclosed invention.DETAILED DESCRIPTION OF THE INVENTION

[0039] The present invention, in some embodiments, provides methods of enriching a target sequence, comprising receiving sequencing reads derived from a plurality of nucleic acid molecule, wherein nucleic acid molecules of the plurality each comprise a barcode sequence; identifying in the sequencing reads a barcode sequence present in nucleic acid molecule comprising the target sequence; contacting the plurality of nucleic acid molecule with a selection probe that hybridized to the identified barcode sequence and comprising a selection moiety; and selecting a probe-bound target nucleic acid molecule by selecting the selection moiety are provided. Systems for performing the methods of the invention are also provided.

[0040] By a first aspect, there is provided a method of enriching a target sequence, the method comprising: a. receiving sequencing reads derived from a plurality of nucleic acid molecules, wherein nucleic acid molecule of the plurality each comprise a barcode sequence; b. identifying in the sequencing reads a barcode sequence that is present in a nucleic acid molecule of the plurality that comprises the target sequence; c. contacting the plurality of nucleic acid molecules with a selection probe, wherein the selection probe comprises a sequence that hybridizes to the identified barcode sequence and a selection moiety; and d. selecting from the plurality of nucleic acid molecules a probe-bound nucleic acid molecule by selecting the selection moiety; thereby enriching a target sequence.

[0041] The difficulty in producing high quality nucleic acid molecules is well known in the art. Very expensive synthesis methods can be employed to ensure high sequence fidelity or cheaper methods can be employed but they produce a high error rate. The instant method presents a way to start with a cheaply produced starting composition of nucleic acid molecules and purifying such that the final product contains sequenced verified highly error-free molecules only. This method is cheap and thus enables the production of highly uniform, high fidelity nucleic acid molecules without the need for the expensive synthesis methods. In some embodiments, the target sequence is a sequence-verified target sequence.

[0042] In some embodiments, the method is an in vitro method. In some embodiments, the method is an ex vivo method. In some embodiments, the method is a method of enriching a target sequence. In some embodiments, the method is a method of enriching nucleic acid molecules comprising the target sequence. In some embodiments, enriching is selecting. In some embodiments, enriching is isolating. In some embodiments, enriching is purifying. Insome embodiments, the plurality of nucleic acid molecules comprises a nucleic acid molecule comprising the target sequence and a nucleic acid molecule devoid of the target sequence. In some embodiments, the method is a method of removing the nucleic acid molecule with the target sequence form the nucleic acid molecule devoid of the target sequence.

[0043] In some embodiments, the target sequence is a gene. In some embodiments, the target sequence is a fragment of a gene. In some embodiments, the target sequence is a sequence of interest. In some embodiments, the gene is a gene of interest. In some embodiments, the target sequence is a sequence-verified target sequence. In some embodiments, the target sequence is a sequence with an error rate below a predetermined threshold. In some embodiments, a sequence desired to be produced is provided and the threshold error rate is an error rate deemed acceptable in the sequence to be produced and the target sequence is the sequence desired to be produced with an error rate below the threshold. In some embodiments, a molecule devoid of the target sequence comprises the sequence desired to be produced but with an error rate above the threshold.

[0044] In some embodiments, the plurality of nucleic acid molecules is an error-prone plurality of nucleic acid molecules. In some embodiments, error prone comprises an error rate greater than a predetermined threshold. In some embodiments, error prone comprises an error rate of greater than 1 error per 50, 100, 200, 250, 300, 400, 500, 600, 750, 800, 900, 1000, 1200, 1400, 1500, 1600, 1800, 2000, 2200, 2400, 2500, 2600, 2800, 3000, 3200, 3400, 3500, 3600, 3800, 4000, 5000, 10000, 50000, 100000, 500000, 1000000, 5000000, 10000000, 50000000, 100000000, 500000000 or 1000000000 bases. Each possibility represents a separate embodiment of the invention. In some embodiments, error prone comprises an error rate of greater than 1 error per 1000 bases. In some embodiments, error prone comprises an error rate of greater than 1 error per 2000 bases. In some embodiments, error prone comprises an error rate of greater than 1 error per 3000 bases. In some embodiments, error prone comprises an error rate of greater than 1 error per 1000000 bases. In some embodiments, error prone comprises an error rate of greater than 1 error per 1000000000 bases. In some embodiments, the target sequence is a low error target sequence. In some embodiments, the selected nucleic acid molecule is a low error nucleic acid molecule. In some embodiments, low error comprises an error rate less than a predetermined threshold. In some embodiments, low error is an error rate lower than 1 error per 100, 200, 250, 300, 400, 500, 600, 750, 800, 900, 1000, 1200, 1400, 1500, 1600, 1800, 2000, 2200, 2400, 2500, 2600, 2800, 3000, 3200, 3400, 3500, 3600, 3800, 4000, 5000, 10000, 50000, 100000, 500000, 1000000, 5000000, 10000000, 50000000, 100000000, 500000000 or 1000000000 bases. Each possibilityrepresents a separate embodiment of the invention. In some embodiments, low error is an error rate of less than 1 error per 1000 bases. In some embodiments, low error is an error rate of less than 1 error per 2000 bases. In some embodiments, low error is an error rate of less than 1 error per 3000 bases. In some embodiments, low error is an error rate of less than 1 error per 4000 bases. In some embodiments, low error is an error rate of less than 1 error per 1000000 bases. In some embodiments, low error is an error rate of less than 1 error per 1000000000 bases.

[0045] In some embodiments, the plurality of nucleic acid molecules comprises fragment of the target sequence. In some embodiments, the fragments of the target sequence can be assembled to produce the target sequence. In some embodiments, the fragments of the target sequence when assembled produce the target sequence. In some embodiments, the method further comprises assembling the selected fragments into a complete target sequence. In some embodiments, the assembling is after the selecting. In some embodiments, the method further comprises isolating the selected fragments and assembling the isolated fragments into a complete target sequence. In some embodiments, the method further comprises isolating / selecting sequence-verified fragments and assembling the isolated / selected sequence-verified fragments into a complete sequence-verified target sequence. Methods of sequence assembly are well-known in the art and any method of assembly may be used. This includes for example Golden Gate assembly, polymerase chain assembly (PCA), and Gibson assembly to name but a few. In some embodiments, the assembly is reference-based assembly to produce the target sequence. In some embodiments, the assembling is performed in an emulsion. In some embodiments, assembling comprises emulsion PCR (emPCR). In some embodiments, assembling comprises amplifying fragments with a tag. In some embodiments, each fragment is amplified with a tag. In some embodiments, the tags are used for assembly. In some embodiments, tags are hybridized to produce combined fragments.

[0046] As used herein, the term “sequencing reads” refers to sequences of bases obtained from a starting DNA indicating the sequence of nucleotides in that starting DNA. Sequencing reads are the output of having sequenced the plurality of nucleic acid molecules. Any type of sequencing that produces reads of individual molecules may be used to produce the sequencing reads and this includes next generation sequencing (NGS), massively parallel sequencing, sequencing by synthesis, nanopore sequencing, Twist sequencing or any sequencing method known in the art. Methods of sequencing which pool reads together to provide a consensus sequence (e.g., Sanger Sequencing) may not be used, as it is necessary toidentify individual molecules comprising the target sequence and getting the sequence of those molecules barcode.

[0047] In some embodiments, the plurality of nucleic acid molecules are the starting nucleic acid molecules. In some embodiments, the plurality of nucleic acid molecules is a group of nucleic acid molecules. In some embodiments, the plurality of nucleic acid molecules is a pool of nucleic acid molecules. In some embodiments, the plurality of nucleic acid molecules comprises at least 10, 100, 1000, 10000, 100000 or 1000000 nucleic acid molecules. Each possibility represents a separate embodiment of the invention. In some embodiments, the starting nucleic acid molecules are produced by error-prone amplification. In some embodiments, error prone amplification is error-prone synthesis. Other methods that may produce the starting nucleic acid molecules include, but are not limited to chemical phosphonamidite synthesis and enzymatic synthesis.

[0048] In some embodiments, nucleic acid molecules of the plurality comprise a barcode. In some embodiments, each nucleic acid molecule of the plurality comprises a barcode. In some embodiments, all nucleic acid molecules of the plurality comprise a barcode. In some embodiments, the barcode is adjacent to the target sequence. In some embodiments, the barcode is 5’ to the target sequence. In some embodiments, the barcode is 3’ to the target sequence. In some embodiments, the barcode is proximal to the target sequence. In some embodiments, proximal is sufficiently close that both the barcode and target sequence are present in the same sequencing read. In some embodiments, the same sequencing read is in a continuous sequencing read. It will be understood that for a barcode to identify a sequencing read / the target sequence it must be that the continuous sequence of the molecule is read and it contains the barcode and the target sequence. Thus, methods of sequencing in which sequences are pooled (e.g., sanger sequencing and pyro sequencing) cannot be used. Reads of individual molecules must be received so that barcodes in molecules with the target sequence can be identified.

[0049] In some embodiments, the barcode is a nucleic acid barcode. In some embodiments, the barcode is a barcode sequence. In some embodiments, the barcode sequence is a sequence and its compliment. In some embodiments, complement is reverse complement. As used herein, the terms "barcode", and "barcode sequence", are used interchangeably and refer to a nucleic acid sequence that that uniquely identifies the DNA molecule either as a specific molecule or as part of a group of molecules (i.e., molecules comprising the target sequence, molecules from a bead). Barcodes are well known in the art and many commercial kits are available that provide barcodes and specifically barcodes for sequencing. In someembodiments, the barcode does not comprise a sequence of more than 10, 12, 15, 17, 20, 22, 25, 27, 30, 35, 40, 45 or 50 nucleotides found in nature. Each possibility represents a separate embodiment of the invention. In some embodiments, the barcode does not comprise a sequence of more than 10, 12, 15, 17, 20, 22, 25, 27, 30, 35, 40, 45 or 50 nucleotides found in a mammalian genome. Each possibility represents a separate embodiment of the invention. In some embodiments, a mammalian genome is a human genome. In some embodiments, the barcode comprises at least 2, 5, 8, 10, 12, 15, 17, 20, 22, 25, 27, 30, 35, 40, 45, 50, 60, 70, 75, 80, 90 or 100 nucleotides. Each possibility represents a separate embodiment of the invention. In some embodiments, the barcode comprises or consists of at most 20, 25, 30, 40, 50, 60, 70, 75, 80, 90, 100, 200, 250, 300, 400, 500, 600, 750, 800, 900 or 1000 nucleotides. Each possibility represents a separate embodiment of the invention. In some embodiments, the barcode comprises or consists of at most 500 nucleotides. In some embodiments, the barcode is a random barcode. In some embodiments, the barcode is a native barcode. It will be understood that the exact sequence of the barcode is not important. The barcode is merely used to identify molecules comprising the correct target sequence and then by hybridizing the barcode with a selection probe those correct molecules can be identified / selected. In some embodiments, the barcode is a mixture of predetermined short sequences. In some embodiments, the barcode is a continuous sequence. In some embodiments, the barcode is not a continuous sequence. In some embodiments, the barcode is split into a plurality of non- continuous coordinates on the molecule. In some embodiments, the barcode pieces are split by constant regions. In some embodiments, the constant regions are common to all beads.

[0050] In some embodiments, identifying is identifying at least one read that comprises the target sequence. In some embodiments, the identifying is identifying in the sequencing reads. In some embodiments, the identifying is identifying a barcode sequence present with the target sequence in a read. In some embodiments, present is present in a continuous read. In some embodiments, a barcode sequence is identifying that is present in at least one nucleic acid molecule comprising the target sequence. In some embodiments, a barcode sequence is identified that is present only in a nucleic acid molecule that comprises the target sequence. In some embodiments, a barcode sequence is identified that is present only in nucleic acid molecules that comprise the target sequence. Since the barcode will be used for molecule selection, it is preferred that the barcode not be present in molecules that lack the target sequence as these undesired molecules would be selected as well. In some embodiments, the identifying is identifying a barcode that is present only in molecules from the plurality comprising the target sequence. In some embodiments, the identifying is identifying allbarcodes present in nucleic acid molecules comprising the target sequence. In some embodiments, the identifying is identifying all barcodes only present in nucleic acid molecules comprising the target sequence.

[0051] In some embodiments, the identifying is identifying a continuous nucleotide sequence that comprises the target sequence and a barcode. In some embodiments, the identifying is identifying a plurality of continuous nucleotide sequences that comprise the target sequence and a barcode. In some embodiments, the identifying is identifying a plurality of barcodes. In some embodiments, the plurality of probes is a plurality of probes in nucleic acid molecules comprising the target sequence. In some embodiments, the plurality of probes is a plurality of probes that are only in nucleic acid molecules comprising the target sequence.

[0052] In some embodiments, the plurality of nucleic acid molecules is contacted with a selection probe. It will be understood that the sequencing reads are derived from the plurality of nucleic acid molecules, but that the plurality of molecules need not be sequenced themselves. Generally, the plurality of nucleic acid molecules will have undergone amplification to produce copies or are copied into a complementary strand and it is these copies that are sequenced. Thus, the original plurality remains and can be contacted with the selection probe. Alternatively, the plurality can be directly sequenced but retained or captured after sequencing and then contacted with the selecting probe. Similarly, after amplification the amplification products can be split and while some are sequenced others are retained for contacting with the selection probe.

[0053] In some embodiments, the selection probe comprises a sequence complementary to the identified barcode sequence. In some embodiments, complementary is reverse complementary. In some embodiments, the selection probe comprises a sequence that hybridizes to the identified barcode sequence. In some embodiments, complementary is at least 70, 75, 80, 85, 90, 92, 95, 97, 99 or 100% complementary. Each possibility represents a separate embodiment of the invention. In some embodiments, complementary is 100% complementary. In some embodiments, the method further comprises producing the selection probe. In some embodiments, producing is synthesizing. In some embodiments, the method further comprises selecting a selection probe with the required sequence (i.e., one complementary to the identified barcode) from a library of selection probes. Once the barcode sequence(s) present with the target sequence is / are identified the sequence(s) for the selection probe become(s) known. This probe(s) can either be synthesized at this point, or a library of probes complementary to the barcodes used may already exist in which case the proper probe(s) can simply be selected from the library.

[0054] The term “complementary” refers to the ability of polynucleotides to form base pairs with one another. Base pairs are typically formed by hydrogen bonds between nucleotide units in antiparallel polynucleotide strands. Complementary polynucleotide strands can base pair in the Watson-Crick manner (e.g., A to T, A to U, C to G), or in any other manner that allows for the formation of duplexes. As persons skilled in the art are aware, when using RNA as opposed to DNA, uracil rather than thymine is the base that is considered to be complementary to adenosine. However, when a U is denoted in the context of the present invention, the ability to substitute a T is implied, unless otherwise stated. Perfect complementarity or 100% complementarity refers to the situation in which each nucleotide unit of one polynucleotide strand can hydrogen bond with a nucleotide unit of a second polynucleotide strand. Less than perfect complementarity refers to the situation in which some, but not all, nucleotide units of two strands can hydrogen bond with each other. For example, for two 20-mers, if only two base pairs on each strand can hydrogen bond with each other, the polynucleotide strands exhibit 10% complementarity. In the same example, if 18 base pairs on each strand can hydrogen bond with each other, the polynucleotide strands exhibit 90% complementarity.

[0055] In some embodiments, the selection probe comprises a selection moiety. In some embodiments, the selection probe is at least two probes, wherein a first probe comprises a sequence complementary to the barcode and a non-barcode complementary sequence and a second probe is complementary to the non-barcode complementary sequence. In some embodiments, the first probe does not comprise a selection moiety. In some embodiments, the second probe comprises a selection moiety. The term "moiety", as used herein, relates to a part of a molecule that may include either whole functional groups or parts of functional groups as substructures. As used herein, the term “selection moiety” refers to a moiety that can be selected, identified or isolated from a solution. In some embodiments, a selecting moiety is a detectable moiety. As sued herein, a “detectable moiety” is a molecule or part thereof that can or does generate a signal which can be measured (i.e., detected). In some embodiments, the detectable moiety is a fluorescent moiety. In some embodiments, the detectable moiety is a nanopore detectable moiety. In some embodiments, the selection moiety is fluorescent and can be separated / isolated by fluorescent activated cell sorting (FACS). In some embodiments, a selecting moiety is an isolatable moiety. In some embodiments, a selecting moiety is a separatable moiety. Moieties which can be bound and captured, such as to a resin or magnet for example, can be considered separatable / isolatable. For example, a biotin moiety can be isolated / separated by binding to streptavidin. In some embodiments, the selection moiety is a capture moiety and it is isolated / separated by binding to a capturingmoiety. In some embodiments, the selection moiety is selected from the group consisting of: a fluorophore, a radioactive moiety, a nanopore-detectable moiety, biotin, a click reaction substrate, an antigen, a magnetic moiety and a bead. In some embodiments, the selection moiety is a fluorophore. In some embodiments, In some embodiments, the selectable moiety can be selected / isolated by a microfluidics channel. In some embodiments, selection of the selection moiety is performed in a microfluidic sorter.

[0056] As used herein, a “capture moiety” and “capturing moiety” are configured to bind to each other. In some embodiments, the capturing moiety specifically binds to the capture moiety. In some embodiments, the capturing moiety and capture moiety specifically bind to each other. In some embodiments, contacting the capturing moiety with the capture moiety produces a complex. In some embodiments, the complex is a capturing moiety-capture moiety complex. Binding pairs that can be used as capture and capturing moieties (i.e., as selection moieties and for detecting / selecting / isolating the selection moiety) are well known in the art and any such pairs can be used. Indeed, essentially any protein and a highly specific antibody against that protein can be used as the binding pair. Examples of binding pairs that can be used as the capture and capturing moiety include, but are not limited to: biotin and streptavidin or avidin, digoxigenin (DIG) and anti-DIG antibody, MS2 and MS2 binding protein (MCP), fluorescein (FAM) and anti-FAM antibody, Cy3 dye and anti-Cy3 antibody, Cy5 dye and anti- Cy5 dye, dinitrophenol (DNP) and anti-DNP antibody, acrydite and acrylamide or acrylamide-containing matrices, aptamers and their binding target, and an antibody and anti- Fc antibody. In some embodiments, the capture moiety is a protein or proteinaceous moiety. In some embodiments, the capture moiety comprises biotin and the capturing moiety comprises avidin. In some embodiments, avidin is selected from streptavidin and neutravidin. In some embodiments, the capture moiety and capturing moiety bind covalently. In some embodiments, the binding pair is not a nucleic acid sequence and its reverse complement. In some embodiments, the capturing moiety is a nucleic acid molecule reverse complementary to the sequence of the capture moiety. In some embodiments, the capturing moiety is perfectly complementary to the capture moiety. In some embodiments, the capturing moiety comprises at least 80, 85, 90, 92, 95, 97, 99 or 100% complementarity to the capture moiety. Each possibility represents a separate embodiment of the invention. In some embodiments, the capturing moiety comprises at least 5, 7, 10, 12, 15, 17, 20 or 25 complementary bases to the capture moiety. Each possibility represents a separate embodiment of the invention.

[0057] In some embodiments, the contacting is under conditions sufficient for hybridizing the selection probe to another sequence. In some embodiments, the contacting is for a timesufficient for hybridizing the selection probe to another sequence. In some embodiments, the another sequence is a barcode sequence. In some embodiments, the another sequence is the identified barcode sequence. In some embodiments, hybridizing is binding. Conditions for nucleotide strand hybridizing are well known and a skilled artisan can select the conditions in order to bind the probe to the identified barcode. In some embodiments, the contacting hybridizes the selection probe to the identified barcode sequence. In some embodiments, the contacting produces probe-bound target nucleic acid molecules. In some embodiments, the hybridizing produces probe-bound target nucleic acid molecules. In some embodiments, target nucleic acid molecules are nucleic acid molecules comprising the target sequence.

[0058] In some embodiments, selecting probe-bound nucleic acid molecules is selecting probe-bound target nucleic acid molecules. In some embodiments, selecting probe-bound nucleic acid molecules is separating probe-bound nucleic acid molecules. In some embodiments, selecting probe-bound nucleic acid molecules is isolating probe-bound nucleic acid molecules. In some embodiments, the selecting / separating / isolating is from a solution. In some embodiments, the solution is a solution comprising the plurality of nucleic acid molecules. In some embodiments, the contacting is in the solution. In some embodiments, the selection probe is added to the solution. In some embodiments, the solution is a solution comprising probe-bound nucleic acid molecules. In some embodiments, the selecting / separating / isolating is by use of the selection moiety. In some embodiments, the selecting / separating / isolating is via the selection moiety. In some embodiments, the selecting / separating / isolating is selecting / separating / isolating the selection moiety. It will be understood that by selecting / separating / isolating the selection moiety, the nucleic acid molecules bound to the selection moiety are also selected / separated / isolated.

[0059] In some embodiments, nucleic acid molecules of the plurality further comprise a unique molecular identifier (UMI). In some embodiments, each nucleic acid molecule of the plurality further comprises an UMI. In some embodiments, each molecule comprises a unique UMI that is different from all other UMIs in the plurality. In some embodiments, a molecule and all copies of that molecule share the same UMI that is different from all other UMIs in the plurality. In some embodiments, an UMI is a sequence of an UMI. In some embodiments, an UMI is a random barcode sequence. In some embodiments, an UMI is a random sequence. The technology for including random sequence when synthesizing nucleic acid molecules is well known, and thus a set (e.g., plurality, group or pool) of nucleic acid molecules wherein each contains a different unique sequence, can be produced by a skilled artisan or purchased commercially. In some embodiments, the UMI comprises at least 5, 6, 7, 8, 9 or 10 randomnucleotides. Each possibility represents a separate embodiment of the invention. In some embodiments, the UMI comprises at least 5 random nucleotides. In some embodiments, the UMI comprises at least 6 random nucleotides. In some embodiments, the UMI comprises 5- 10 random nucleotides. In some embodiments, the UMI comprises 6-10 random nucleotides. In some embodiments, the UMI comprises at most 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 75, 80, 90 or 100 random nucleotides. Each possibility represents a separate embodiment of the invention. In some embodiments, the UMI is a continuous sequence. In some embodiments, the UMI is a non-continuous sequence. The UMI can be used when analyzing the sequencing reads to identify clones (copies) all originating from the same original sequence.

[0060] The positioning of the barcode sequence relative to the target sequence is not relevant and it can be 5’ or 3’. Similarly, the positioning of the UMI is not significant and it can be 5’ or 3’ to the barcode sequence as well as the target. It may be between the target sequence and the barcode sequence. However, primer binding sequences if present for amplification should be near or proximal to the ends so that the full molecule can be amplified. Thus, the binding site for the forward primer and reverse primers should be at the 5’ and 3’ ends or proximal thereto. In some embodiments, proximal is within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 75, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450 or 500 nucleotides. Each possibility represents a separate embodiment of the invention.

[0061] In some embodiments, the method comprises before step (a) producing the sequencing reads. Two methods for producing the sequencing reads, one that uses supports and one that does not, are provided hereinbelow. Either method may be used, though the method using solid supports has advantages related to the ability to more easily isolate the supports as well as bringing together various sequences. Especially when fragments of a large target sequence are being used, the support method is especially advantageous as barcodes from beads that contain all of the fragments, and all with the correct sequence, may be selected / separated / isolated.

[0062] Producing sequencing reads using supports: In some embodiments, the plurality of nucleic acid molecules is conjugated or hybridized to supports. In some embodiments, the sequencing reads are derived from a plurality of nucleic acid molecule conjugated or hybridized to supports. In some embodiments, the method comprises before step (a) producing the sequencing reads.

[0063] In some embodiments, the method comprises providing a plurality of supports, wherein each support is conjugated to a nucleic acid molecule. In some embodiments, eachsupport is conjugated to a plurality of nucleic acid molecules. In some embodiments, the support is a solid support. In some embodiments, the support is a physical support. In some embodiments, the support is a bead or resin. In some embodiments, the support is not a flat surface. In some embodiments, beads are not localized at discrete locations. In some embodiments, beads are not localized at discrete locations on a surface. In some embodiments, beads are in solution. In some embodiments, resin is in solution. It will be understood that molecules of a resin do not have a discrete location. In some embodiments, the solid support is a bead. In some embodiments, the bead is an avidin bead. In some embodiments, the bead is a magnetic or paramagnetic bead. In some embodiments, the nucleic acid molecules are immobilized on the support. In some embodiments, immobilized on is conjugated to. In some embodiments, conjugated is covalently conjugated. In some embodiments, the nucleic acid molecule is conjugated at its 5’ end to the support. In some embodiments, the nucleic acid molecule is conjugated at its 3’ end to the support. In some embodiments, a support is such as is depicted in Figure 1. In some embodiments, a support comprises all of the components depicted in Figure 1, though the UMI is optional.

[0064] In some embodiments, a nucleic acid molecule on the support comprises a barcode sequence and a capturing sequence. In some embodiments, each nucleic acid molecule on the support comprises a barcode sequence and a capturing sequence. In some embodiments, a capturing sequence is a capturing moiety that is a nucleic acid sequence. In some embodiments, the barcode is specific to the support. In some embodiments, the barcode sequence uniquely identifies the support. In some embodiments, the barcode sequence of a support is not shared by other supports of the plurality of supports. The barcode on the support is thus common to all nucleic acid molecules on the support, but is unique as compared to other supports. Thus, the barcode uniquely identifies the support. In some embodiments, there is a group or plurality of supports with the same barcode and thus the barcode identifies that group. In some embodiments, the plurality of supports comprises a first support comprising a first barcode and second support comprising a second barcode and the first and second barcodes are distinct from each other. In some embodiments, the plurality of supports comprises a first group of supports comprising a first barcode and second group of supports comprising a second barcode and the first and second barcodes are distinct from each other. In some embodiments, a solid support, e.g., a bead, comprises at least 50, 100, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000, 5,000,000, 10,000,000, 50,000,000, 100,000,000, 500,000,000 or 1,000,000,000 nucleic acid molecules immobilized on it. Each possibility represents a separate embodiment of the invention. In some embodiments, a solidsupport, e.g., a bead, comprises 100-1,000,000 nucleic acid molecules. In some embodiments, a solid support, e.g., a bead, comprises 100-10,000,000 nucleic acid molecules. In some embodiments, a solid support, e.g., a bead, comprises 1,000-1,000,000 nucleic acid molecules. In some embodiments, a solid support, e.g., a bead, comprises 1,000-10,000,000 nucleic acid molecules.

[0065] In some embodiments, the nucleic acid molecules of the plurality comprise an UMI. In some embodiments, the UMI is a molecule specific UMI. In some embodiments, each nucleic acid molecule of the plurality comprises an UMI.

[0066] In some embodiments, a nucleic acid molecule of the plurality is conjugated to the support by its 5’ end and the barcode sequence is 5’ to the capturing sequence. In some embodiments, a nucleic acid molecule of the plurality is conjugated to the support by its 3’ end and the barcode sequence is 3’ to the capturing sequence. In some embodiments, a nucleic acid molecule of the plurality is conjugated to the support by its 5’ end and the UMI sequence is 5’ to the capturing sequence. In some embodiments, a nucleic acid molecule of the plurality is conjugated to the support by its 3 ’ end and the UMI sequence is 3 ’ to the capturing sequence. The relative positioning of the UMI and the barcode sequences is not significant. The UMI can be 5’ or 3’ to the barcode if an UMI is present. In some embodiments, a nucleic acid molecule of the plurality is conjugated to the support by an internal modification. In some embodiments, the internal modification is a modification of the backbone of the nucleic acid molecule. In some embodiments, the nucleic acid molecule is double stranded and is conjugated to the support at either end.

[0067] In some embodiments, the method further comprises providing a group of nucleic acid molecules comprising a capture sequence. In some embodiments, the nucleic acid molecules are oligonucleotides. In some embodiments, the group is in solution. In some embodiments, a group is a pool. In some embodiments, a group is a plurality. In some embodiments, each nucleic acid of the group comprises a capture sequence. In some embodiments, the capture sequence and the capturing sequence are complementary. In some embodiments, the capture sequence and the capturing sequence hybridize to each other. In some embodiments, the capture sequence and the capturing sequence are capable of hybridizing to each other. In some embodiments, the group comprises a nucleic acid molecule comprising the target sequence. In some embodiments, the group comprises a nucleic acid molecule devoid of the target sequence. In some embodiments, the pool of nucleic acid molecules is a pool of oligonucleotides. In some embodiments, the pool is a pool of molecules with essentially the same sequence and some comprise the sequence with an error rate below a predeterminedthreshold (i.e., the target sequence) and some comprise the sequence with an error rate above the predetermined threshold.

[0068] In some embodiments, the method further comprises contacting the provided group of nucleic acid molecules to the provided plurality of supports. In some embodiments, the contacting is under condition sufficient to hybridize the capture sequence to the capturing sequence. In some embodiments, the contacting is for a time sufficient to hybridize the capture sequence to the capturing sequence. In some embodiments, to hybridize is to bind. In some embodiments, the hybridizing produces at least one hybridized support. In some embodiments, at least one hybridized support is a plurality of hybridized supports. In some embodiments, the capturing sequence is common to all supports and the capture sequence is common to all nucleic acid molecule of the group. In some embodiments, the plurality of supports comprises a support with a first capturing sequence and a support with a second capturing sequence and the group of nucleotide molecules comprises a nucleic acid molecule comprising a first capture sequence and a nucleic acid molecule comprising a second capture sequence. In some embodiments, the first capturing sequence is complementary to the first capture sequence and the second capturing sequence is complementary to the second capture sequence. In some embodiments, the first capturing sequence cannot hybridize with the second capture sequence and the second capturing sequence cannot hybridize with the first capture sequence. In some embodiments, the nucleic acid molecules hybridized to the support is as depicted in Figure 3.

[0069] In some embodiments, the method further comprises contacting the hybridized support with a polymerase. In some embodiments, the solution is contacted with a polymerase. In some embodiments, the polymerase extends ends of nucleic acid molecules. In some embodiments, ends are free ends. In some embodiments, ends are 3’ ends. In some embodiments, ends are 5’ ends and ends are extended by ligation or chemical reaction. Polymerases for extending nucleic acid molecules with free ends are well known in the art and a skilled artisan can select the proper polymerase. In some embodiments, the polymerase is a high-fidelity polymerase. In some embodiments, a high-fidelity polymerase produces a low error rate. In some embodiments, contacting with a polymerase produces an elongated solid support. In some embodiments, a plurality of elongated solid supports are produced. It will be understood that for a polymerase to be active it needs a complementary strand to make a copy of. Thus, single stranded molecules in solution will not be extended by the polymerase. However, when the nucleic acid molecules in solution hybridize to the supports via the capturing-capture sequences there is now a template for elongation. The nucleic acid moleculeconjugated to the support is elongated to now include the sequence of the hybridized nucleic acid molecule and the nucleic acid molecule that hybridized to the support is now elongated to have the barcode sequence (or rather its complement which also works as a barcode sequence).

[0070] In some embodiments, the method further comprises amplifying elongated nucleic acid molecules. In some embodiments, primers are added to amplify the elongated nucleic acid molecules. In some embodiments, a single primer is added to produce a linear amplification. In some embodiments, a nucleic acid molecule immobilized on the bead further comprises a first primer binding sequence. In some embodiments, the first primer binding sequence is 5’ to the barcode sequence. In some embodiments, the first primer binding sequence is 3’ to the barcode sequence. In some embodiments, the first primer binding sequence is 5’ to the capturing sequence, the first primer binding sequence is 3’ to the capturing sequence. In some embodiments, the group of nucleic acid molecules comprise a second primer binding sequence. In some embodiments, the second primer binding sequence is 5’ to the target sequence. In some embodiments, the second primer binding sequence is 3’ to the target sequence. In some embodiments, the second primer binding sequence is 5’ to the capture sequence. In some embodiments, the second primer binding sequence is 3’ to the capture sequence. In some embodiments, the first and second primer binding sequences bind a forward primer and a reverse primer. In some embodiments, the first and second primer binding sequences facilitate amplification of the elongated molecules. In some embodiments, amplification produces copies of the elongated nucleic acid molecules. In some embodiments, the sequencing is sequencing the copies.

[0071] In some embodiments, the method further comprises dissociating the hybridized strands. In some embodiments, the method further comprises dissociating the elongated nucleic acid molecule comprising the capture sequence from the elongated nucleic acid molecule comprising the capturing sequence. In some embodiments, the method further comprises dissociating the elongated nucleic acid molecule comprising the capture sequence the elongated nucleic acid molecules immobilized on the supports.

[0072] In some embodiments, the method further comprises sequencing the dissociated molecules. In some embodiments, the sequencing is sequencing the dissociated elongated nucleic acid molecules. In some embodiments, the sequencing is sequencing the molecules in solution. It will be understood that during amplification dissociation inherently happens. This results in copies of both strands being in solution. In some embodiments, the method comprises sequencing the copies produced by amplification. If there is no amplification onlyone strand will be in solution, however, amplification will produce tens or hundreds of strands (both positive and negative strands) and thus provides more template for sequencing. In some embodiments, the sequencing produces the sequencing reads. In some embodiments, the amplification product undergoes library preparation before sequencing. Methods of prepping libraries for sequencing are well known in the art and any such method may be used. These preparation steps increase the read length and / or read base quality. In some embodiments, library preparation comprises producing synthetic long reads. In some embodiments, the sequencing is synthetic long reads sequencing.

[0073] In some embodiments, contacting the plurality of nucleic acid molecules with a selection probe is contacting the elongated nucleic acid molecule with a selection probe. In some embodiments, the contacting is contacting the elongated nucleic acid molecules comprising the capturing sequence with the selection probe. In some embodiments, the contacting is contacting the elongated nucleic acid molecules immobilized on the support with the selection probe. In some embodiments, the contacting is contacting the elongated nucleic acid molecules comprising the capture sequence with the selection probe. In some embodiments, the contacting is contacting the elongated nucleic acid molecules comprising the capture sequence with a first selection probe and a second selection probe, wherein the first selection probe hybridizes to the elongated nucleic acid molecule and the second selection probe hybridizes to the first selection probe. In some embodiments, the second selection probe comprises the selection moiety.

[0074] Producing sequencing reads without supports: In some embodiments, the plurality of nucleic acid molecules is a pool of nucleic acid molecules. In some embodiments, the method further comprises providing a pool of nucleic acid molecules. In some embodiments, the pool comprises a nucleic acid molecule comprising the target sequence and a nucleic acid molecule devoid of the target sequence. In some embodiments, nucleic acid molecules of the pool comprise a barcode sequence. In some embodiments, each nucleic acid molecule of the pool comprises a barcode sequence. In some embodiments, the barcode is specific or unique to the nucleic acid molecule and is not shared by other nucleic acid molecule of the pool.

[0075] In some embodiments, the nucleic acid molecules of the pool are linked. In some embodiments, linked is connected. In some embodiments, a portion of the nucleic acid molecule of the pool are linked. In some embodiments, the pool comprises aggregates, wherein the aggregates comprise linked nucleic acid molecules. In some embodiments, an aggregate comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or 100 molecules. Each possibility represents a separate embodimentof the invention. In some embodiments, linked is by a chemical link. In some embodiments, linked is by a linker. In some embodiments, linked is by a biomolecule. In some embodiments, linked is concatenation. In some embodiments, the barcode is unique to the linked molecules. In some embodiments, the probe is unique to an aggregate. In some embodiments, each molecule of an aggregate comprises its own unique UMI. In some embodiments, each molecule of an aggregate comprises its own unique barcode.

[0076] In some embodiments, the method further comprises amplifying the pool of nucleic acid molecules to produce copies of the nucleic acid molecules. In some embodiments, the amplifying comprises contacting with a polymerase. In some embodiments, the amplifying comprises contacting with primers. In some embodiments, the molecule of the pool comprise primer binding sites. In some embodiments, the method further comprises amplifying the nucleic acid molecules in aggregates to produce copies of the nucleic acid molecules in solution. In some embodiments, the method further comprises sequencing at least a portion of the copies. In some embodiments, the method further comprises sequencing the copies. In some embodiments, the sequencing produces sequencing reads.

[0077] In some embodiments, the contacting the plurality of nucleic acid molecules with the selection probe comprises contacting the pool of nucleic acid molecules. In some embodiments, the contacting the plurality of nucleic acid molecules with the selection probe comprises contacting the copies. In some embodiments, the contacting the plurality of nucleic acid molecules with the selection probe comprises contacting a portion of the copies. In some embodiments, the contacted portion is the portion that was not sequenced. In some embodiments, the contacting the plurality of nucleic acid molecules with the selection probe comprises contacting the aggregates. In some embodiments, the aggregates are not sequenced. In some embodiments, the method further comprises dissociating the aggregates. In some embodiments, dissociating is separating. In some embodiments, the aggregates are separated into individual molecules.

[0078] By another aspect, there is provided a system for performing a method of the invention.

[0079] By another aspect, there is provided a system comprising: a. a thermomixer; b. a DNA sequencer: c. a control unit; and d. a computer program product comprising a non-transitory computer- readable storage medium having program code embodied thereon, theprogram code executable by at least one hardware processor to apply a detectable probe to the thermomixer.

[0080] In some embodiments, the system is for use in performing a method of the invention. In some embodiments, the system is suitable for performing a method of the invention. In some embodiments, the computer program product comprises program code executable by at least one hardware processor to perform a method of the invention. In some embodiments, the control unit analyzes DNA sequencing results. In some embodiments, the DNA sequencing results are sequencing reads. In some embodiments, the DNA sequencing results are from the DNA sequencer. In some embodiments, the DNA sequencer produces sequencing results and transmits them to the control unit. In some embodiments, the control unit selects a barcode sequence. In some embodiments, a barcode sequence that is in a continuous molecule with a target sequence is selected. In some embodiments, the barcode sequence or its complement sequence is to be part of a selection probe. In some embodiments, the complement is the reverse complement. In some embodiments, the complement of the barcode sequence is to be part of the selection probe. In some embodiments, the control unit selects a selection probe based on the selected barcode.

[0081] In some embodiments, the selection probe identified by the control unit is applied to the thermomixer. In some embodiments, the selection probe selected by the control unit is applied to the thermomixer. In some embodiments, the system further comprises a nucleic acid molecule printer. In some embodiments, the nucleic acid molecule is a DNA molecule. In some embodiments, the nucleic acid molecule is an oligo. In some embodiments, the system further comprises a primer. In some embodiments, the primer is an oligo primer. In some embodiments, the primer is a pair of primers. In some embodiments, the primers are amplification primers. In some embodiments, the primers can be added to the thermomixer for amplification. In some embodiments, the printer is capable of producing the selection probe. In some embodiments, the control unit transmits the sequence of the selection probe to the printer that prints the probe. In some embodiments, the printer transmits to the computer program product that the probe has been printed. In some embodiments, the at least one hardware processor applies the printed probe to the thermomixer.

[0082] In some embodiments, the thermomixer mixes the probe and nucleic acid molecules in the mixer such that the probe and molecules can hybridize. In some embodiments, the thermomixer produces conditions sufficient for hybridizing the probe to a nucleic acid molecule comprising the barcode. In some embodiments, the thermomixer is programed to mix for a time sufficient for the probe to hybridize to a nucleic acid molecule comprising thebarcode. In some embodiments, the thermomixer comprises an inlet configured to intake a solution comprising nucleic acid molecules.

[0083] In some embodiments, the system further comprises a sorting device. In some embodiments, the sorting device is capable of sorting molecules comprising the selection probe from molecule devoid of the selection probe. In some embodiments, the sorting is sorting supports comprising the selection probe from supports devoid of the selection probe. In some embodiments, the selection probe comprises a selection moiety and the sorting is sorting the selection moiety. In some embodiments, the selection moiety comprises a fluorophore and the sorting device is a FACS machine. In some embodiments, the device comprises a connector connecting the thermomixer to the sorting device. In some embodiments, the connector is a tube. In some embodiments, the connector is a channel. In some embodiments, the connector is configured to deliver a solution from the thermomixer to the sorting device. In some embodiments, the connector is configured to deliver nucleic acid molecules hybridized to the selection probe from the thermomixer to the sorting device. In some embodiments, the sorting device is a microfluidics sorter.

[0084] As used herein, the term "about" when combined with a value refers to plus and minus 10% of the reference value. For example, a length of about 1000 nanometers (nm) refers to a length of 1000 nm+- 100 nm.

[0085] It is noted that as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a polynucleotide" includes a plurality of such polynucleotides and reference to "the polypeptide" includes reference to one or more polypeptides and equivalents thereof known to those skilled in the art, and so forth. It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as "solely," "only" and the like in connection with the recitation of claim elements, or use of a "negative" limitation.

[0086] In those instances where a convention analogous to "at least one of A, B, and C, etc." is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., "a system having at least one of A, B, and C" would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms.For example, the phrase "A or B" will be understood to include the possibilities of "A" or "B" or "A and B."

[0087] It is appreciated that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable subcombination. All combinations of the embodiments pertaining to the invention are specifically embraced by the present invention and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all sub-combinations of the various embodiments and elements thereof are also specifically embraced by the present invention and are disclosed herein just as if each and every such sub-combination was individually and explicitly disclosed herein.

[0088] As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents, unless the context clearly dictates otherwise. The terms “a” (or “an”) as well as the terms “one or more” and “at least one” can be used interchangeably.

[0089] Furthermore, “and / or” is to be taken as specific disclosure of each of the two specified features or components with or without the other. Thus, the term “and / or” as used in a phrase such as “A and / or B” is intended to include A and B, A or B, A (alone), and B (alone). Likewise, the term “and / or” as used in a phrase such as “A, B, and / or C” is intended to include A, B, and C; A, B, or C; A or B; A or C; B or C; A and B; A and C; B and C; A (alone); B (alone); and C (alone).

[0090] Wherever embodiments are described with the language “comprising,” otherwise analogous embodiments described in terms of “consisting of’ and / or “consisting essentially of’ are included.

[0091] Additional objects, advantages, and novel features of the present invention will become apparent to one ordinarily skilled in the art upon examination of the following examples, which are not intended to be limiting. Additionally, each of the various embodiments and aspects of the present invention as delineated hereinabove and as claimed in the claims section below finds experimental support in the following examples.

[0092] Various embodiments and aspects of the present invention as delineated hereinabove and as claimed in the claims section below find experimental support in the following examples.EXAMPLES

[0093] Generally, the nomenclature used herein and the laboratory procedures utilized in the present invention include molecular, biochemical, microbiological and recombinant DNA techniques. Such techniques are thoroughly explained in the literature. See, for example, "Molecular Cloning: A laboratory Manual" Sambrook et al., (1989); "Current Protocols in Molecular Biology" Volumes I-III Ausubel, R. M., ed. (1994); Ausubel et al., "Current Protocols in Molecular Biology", John Wiley and Sons, Baltimore, Maryland (1989); Perbal, "A Practical Guide to Molecular Cloning", John Wiley & Sons, New York (1988); Watson et al., "Recombinant DNA", Scientific American Books, New York; Birren et al. (eds) "Genome Analysis: A Laboratory Manual Series", Vols. 1-4, Cold Spring Harbor Laboratory Press, New York (1998); methodologies as set forth in U.S. Pat. Nos. 4,666,828; 4,683,202; 4,801,531; 5,192,659 and 5,272,057; "Cell Biology: A Laboratory Handbook", Volumes I-III Cellis, J. E., ed. (1994); "Culture of Animal Cells - A Manual of Basic Technique" by Freshney, Wiley- Liss, N. Y. (1994), Third Edition; "Current Protocols in Immunology" Volumes I-III Coligan J. E., ed. (1994); Stites et al. (eds), "Basic and Clinical Immunology" (8th Edition), Appleton & Lange, Norwalk, CT (1994); Mishell and Shiigi (eds), "Strategies for Protein Purification and Characterization - A Laboratory Course Manual" CSHL Press (1996); all of which are incorporated by reference. Other general references are provided throughout this document.Materials and methods

[0094] Tagged beads: High-quality DNA synthesis comprises the following main steps:Step #0 - barcoded bead design and productionStep #1 - oligos designStep #2 - oligos productionStep #3 - oligos validation via sequencing, checking whether the oligos match the design sequence or notStep #4 - pulling out validated or filtering out unvalidated oligosStep #5 - assembly of validated oligo to genes

[0095] The preliminary step includes the preparation of the materials for the process. The main component in the preparation step is the design and production of a component that will enable the tagging of a single molecule from the oligos for identification during sequence verification, which may support possible clustering of different oligos (optional), and then selection of the sequence-verified oligos. This component will be referred to hereafter as a solid support, although some embodiments may not include the use of beads and may be encoded directly in the oligo sequence.

[0096] Step #0 - the design and production of barcoded beads. The barcoded beads (see Fig. 1) may include a 5' primer binding site 102- unique molecular identifier (UMI) 103- a barcode sequence 104 - gene capturing sequence 105 - cleavage site (deoxy uracil or type II restriction enzyme recognition sequence, ribo base, phosphorothioate bond, photocleavable moieties or other cleavage mediators that may be added to the bead or to support the addition in a later stage, such as during amplification through the primers).

[0097] The gene synthesis process starts with the definition of a target gene or other desired nucleic acid sequence. If needed, the gene is then divided into segments that can be synthesized as oligos. Step #1 starts with segments of the target gene to which auxiliary tags and sequences are added - primer binding site on one end (5') and gene capture complementary sequence to the other. The gene capture sequence is specific for every group of segments composing a desired sequence.

[0098] In step #2 the oligos are produced. This phase is performed using methods known in the art, either enzymatic or chemical, for example, on DNA microarrays. Error-prone PCR can be used to create copies of the sequence which will include variants of the sequence do to the introduction of errors. The synthesis of DNA is more tolerated for errors than by common standards in the field, which gives the opportunity to increase length over accuracy. The tolerated error rate depends on several factors downstream of the process, including the number of oligos required for the final sequence.

[0099] In step #3, the oligos are attached through the capture sequence hybridization that may be followed by ligation or polymerization to the barcoded beads. The oligos concentration will provide around a single representation of each segment of the gene to a bead. The sequences on the bead are amplified to a solution. The amplicons are sequenced by known sequencing methods such as sequencing by synthesis (SBS), nanopore sequencing, or sequencing by ligation (SBL) to identify the correct sequences matching the sequences designed in step #1. This step may include a repair iteration after the sequencing for beads that do not have all the segments to assemble the full-length gene.

[0100] In step #4, beads with validated oligos are pulled out based on their bead-barcode, for example using fluorescent probes attached to the beads barcode of the bead with correct sequences. The probes can be taken from a library or synthesized on the spot. The pulled beads with validated oligos serve as an input to step #5, where they undergo amplification, for example, with emulsion PCR. The amplified beads are collected and undergo assembly. Methods for oligos assembly to genes are known in the art (golden-gate, Gibson, polymerase chain assembly (PCA)). However, the disclosed invention also facilitates in-situ synthesis,making use of the content of validated beads holding multiple different DNA section types, while pulling is implemented post-synthesis.

[0101] Gene segments, oligo pool: Each gene variant may be split into gene segments depending on its length. If short enough, the gene variant will not be split at all and will hold the entire segment. Gene segments are manufactured as an oligo pool in high throughput in processes known in the art, such as microarray synthesis (multiple options).

[0102] The oligos are designed to include the following two parts. First, is a sequence complementary to the capture sequence and the cleavage site - facilitating hydrogen bonding between the bases of the oligos and the complementary bases in the capture sequence attached to the beads. The capture sequence is unique for a given gene variant and common to all the segments originating from the same gene variant. The paring will mediate the transformation of the oligo to follow the bead DNA through elongation or ligation. Second, is the desired sequence.

[0103] An optional cleavage and primer binding site is optional. An additional primer binding site can be added, and it is used to mediate the exponential PCR amplification to prepare an amplicon for sequencing or to mediate circularization of the molecule for amplification.

[0104] DNA sequencer: Gene segments together with their attached tags are sequenced in a way known in the art such as SBS or nanopore (multiple options). The sequencing may be applied in two steps of the flow. In the first step, the goal of this sequencing is twofold. To identify error-free molecules or error-free beads, i.e., beads not holding any incorrect sequence segment, and to make sure that these error-free beads hold all the relevant gene segments required for gene assembly. In some embodiments of the disclosed invention, only error-free criterion is used. Bead tags associated with the above requirements are stored for the following steps. These bead tags are referred to as the selected bead tags.

[0105] A second DNA sequencing step may take place post-sorting or post-assembly. After the beads are selected, DNA sequencing may be used for the identification of the selected bead, identification of missing gene variants or to verify correct assembly in embodiments which include assembly.

[0106] In some applications, the amplification process will introduce the sequencing adaptors with a 5' tail of P5 and P7 to prepare the library for sequencing.

[0107] Probe synthesis

[0108] DNA printer: The DNA printer is used for on-the-fly probe printing. These probes are printed to fit the selected bead tags or the unique UMI sequence of error-free oligos. The printer may print the probes with fluorescent tags or another selection moiety, or just print themarkers while selection elements may be later added via a separate process such as PCR amplification or tailing with terminal transferase (TdT) with one or more types of fluorophore or other modification of one or more of the triphosphate nucleotides (dNTPs).

[0109] The probes can be produced in high amounts or can be produced in low amounts and amplified to a higher amount. The product of the amplification can be used as dsDNA or turned to ssDNA for example by protecting one strand from 5' digestion with exonucleases using one primer with phosphorothioate at the 5' end. The amplification can be linear and produce ssDNA probes.

[0110] To undergo amplification, probes can be synthesized with amplification- supporting sequences, i.e., primer binding sites. The sequences may also be added in a second stage, for example with ligation. These sequences may include a primer binding site and the complementary sequence to the primer binding site of the other strand for exponential amplification or a single primer binding site or self-priming with nickase for linear amplification.

[0111] Probe synthesis via ligation of predesigned sequences: In the case where the bead barcode is made through "split and pool" ligation of predesigned sequences from a library, those sequences can also be used to produce the probe. After the identification of beads containing correct sequences, probes matching the bead barcode of the beads with correct sequences can be produced through ligation of the sequences comprising the bead barcode. For example, if the bead barcode is a result of 3 steps of split and pool of a library of 384 sequences, after the identification of a bead barcode with the correct sequence this sequence can be produced via ligation of the 3 sequences from the library, in a single or two-step flow.

[0112] Another option in cases where the hybridization of each sequence from the library is stable enough to be used as a probe, the probes are used with 3 different fluorescent tags and only beads carrying the 3 colors are selected. This will work only in cases where there are no other beads with the same probes in different combinations.

[0113] Optical bead sorter: An optical bead sorter works on a single-bead basis. Beads are sorted based on the light emitted from the fluorescent tags. One example of such a device is a fluorescence-activated cell sorter (FACS) machine, which is usually used for cell sorting, but can be also used for bead sorting.

[0114] Reference is now made to Figure 2 which presents the operation principle of a FACS machine used in accordance with the disclosed invention. Beads holding only error-free gene segments (good beads) 201 are marked with a fluorescent probe while beads holding errored gene segments (bad beads) 202 are not marked. Good beads will ideally contain the full set ofsegments needed to assemble the target sequence and all the segments will be error- free. When the beads are passed via a pipe 203, optical sensor 204 identifies the light emitted from good beads and applies an electrical field 205 to pull these beads in one direction. Bad beads on the other hand are pulled in a different direction such that in the end all the "good" beads 206 are separated from the "bad" or errored beads 207. The good beads can be sorted one by one for downstream applications or in batches for other applications. For example, in an accurate oligo pool production, the beads don’t have to be sorted one by one. All the correct beads can be sorted in a batch to a single tube.

[0115] These principles of action can be implemented in a microfluidic device as well, or any apparatus that can separate the good beads from the bad. The apparatus selected will depend on the type of tag added to the error- free beads.

[0116] Droplet generation: Once a good bead with more than a single oligo is selected, it can undergo assembly. One of the assembly methods that supports high throughput reaction is emulsion. The droplets for the assembly process can be generated in several ways. Two well- established methods are microfluidics and vortex for an unspecific emulsification. Post assembly the droplets can be placed in specific locations for further validation or higher level of assembly.Example 1: Tagged bead production

[0117] Reference is now made to Figure 1, which includes an illustration of an embodiments of bead structure of the invention.

[0118] The beads are a prerequisite element. They may be prepared in advance, in large volume, may be used over multiple production cycles, and may be restored and corrected for multiple uses.

[0119] The bead material may include a large variety of materials such as polystyrene and hydrogel, or it may not be a bead like a dendrimer polymer or a direct encoding of identification tag that enables selection to the oligo sequence. The bead size may also vary according to several different factors of the system flow, such as the sensitivity to reagents, its ability to support the selection process, the efficacy of biochemical processes, and the cost.

[0120] The beads hold 2 basic elements: the physical bead 101 and multiple DNA tags or barcodes 102-105 attached to it. This structure of DNA tags repeats itself multiple times on the bead. For simplicity, the figure depicts only 3 tags attached to the bead. However, the actual number of tags may range from a few 10s to hundreds of millions, all attached to the same bead. It is preferred that the attachment to the bead will be through the 5' end of thenucleic acid molecule (e.g., DNA molecule) and the free end of the tags will be the 3’ end, but this is not mandatory. Each DNA tag attached to a bead is structured in the following way:

[0121] Primer binding sequence (BS) 102 - this segment is common to all the beads. It is attached directly or through a tether to the bead via its 5’ end. The primer length is typically a few 10s of bases. The primer binding sequence is used for amplification. Illumina's TruSeql is an example of a primer binding site that may be used as part of preparing the oligos for sequencing.

[0122] Unique Molecule Identifier - UMI 103 - this tag is downstream to the primer sequence and is unique for every molecule on the bead. It allows one to identify amplified gene segments that have a common origin, i.e., a result of amplification of a common source, in a single DNA molecule. The UMI enables the quantification of gene segments originating from different sources, namely identical oligos on the same bead, which will be easily identified based on the different UMI attached to them. As the UMI context is local to the bead, a smaller number of bases should suffice, and the number of bases may be in the 10-20 range or even less. Statistically, an even smaller number of bases will enable the identification of the different molecules on the same bead in the great majority of cases. Alternatively, the UMI sequence does not have to be part of the bead, it can be synthesized as part of the DNA oligo and to be cleaved off before assembly.

[0123] Bead barcode 104 - this tag follows the UMI tag though the order of these two elements is not essential. It typically holds a few 10s of bases. The size of the barcode is proportional to the batch size to enable the stable and specific hybridization of the selection probe. This tag is unique per bead, i.e., each bead has a substantially different barcode so that a specific bead can be identified based on its barcode sequence. The order of the UMI and bead barcode is flexible. In some cases, the UMI will be closer to the bead and in some cases it will be reversed.

[0124] Gene tag or capturing sequence 105 - this tag is the sequence following the bead barcode / UMI and typically holds a few 10s of bases to enable specific attachment to the corresponding tag (capture sequence) on the oligo. This sequence is common to a given bead and is designed to capture gene segments of a specific gene variant. Unlike the bead barcode, which is unique per bead, the capture sequence may repeat itself over multiple beads. For example, if we want to produce 100 different gene variants, we need to have 100 different capture sequences, one capture sequence per gene variant, in proportion to the error rate and the predicted molecules with the correct sequence. If we prepare 100,000 beads, we should have approximately 1,000 beads sharing a common capture sequence. Alternatively, if theentire target length of the gene is captured to the bead, a single capturing tag can be used. The capture sequence is followed by a cleavage site.

[0125] Cleavage site - The capture sequence may end with a cleavage site (not shown in Figure 1). The cleavage site's purpose is to remove the tag from the oligo sequence. Once the oligo sequence is removed it is available for delivery of downstream processes, such as assembly. The cleavage sequence may act on dsDNA or ssDNA. The cleavage site may be a modified base, such as deoxy Uracil, ribonucleotide, phosphorothioate bond, photocleavable moieties, or sequence that can mediate the cleavage, such as type Ils restriction enzyme binding site (less preferred, since it limits the sequence space of the synthesized oligos and makes a constant end to the bead tag that can result in cross-hybridization between different gene tags) or with the use of CAS9 and CRISPR with appropriate guide RNA. In some embodiments, the cleavage site may work on amplified DNA attached to the bead, while in other embodiments, it may work on amplified DNA that is free in the solution. The cleavage sequence may work by supporting the introduction of the cleavage mediator through the amplification primer. The cleavage site is optional and may be replaced with direct amplification of the oligo from the bead using appropriate primers.

[0126] It should be noted that beads are just one example of using a cluster holding the content previously defined. Other methods may exist (e.g., other solid supports) to cluster the content without having what is commonly called a bead. To emphasize the point that there are many ways to implement the clustering functionality that the bead takes, we define the term G-bead. G-bead stands for a generalized bead and covers any implementation of the bead functionality. Also, the sequence of tags above may be changed, and some tag segments may be omitted or change in order in some embodiments of the disclosed invention.

[0127] The production of the bead tags can be done by several different methods known in the art. One example is the "split and pool" methodology:

[0128] Nucleotide synthesis - The synthesis can be done chemically in a 3' to 5' or a 5' to 3' direction or enzymatically. The primer binding sequence - this part of the tag can be presynthesized, not on the bead, with a modification to mediate attachment to the bead. The attached sequence may include not only the primer binding sequence but also the UMI. Alternatively, the primer binding sequence can be synthesized directly on the bead. All the beads will undergo the same processes through the primer synthesis. In this way, the synthesis of the UMI tag will include the synthesis of random sequences all over the bead by applying a mixture of bases instead of a single base. The synthesis of the bead tags with the "split and pool" methodology - the beads are mixed and redivided arbitrarily to add one of the fourpossible bases for each redivided group. This methodology is applied for the synthesis of the bead tag. The gene tag will also be synthesized with split and pool that will be followed by sequencing to identify the synthesized gene tags, or by applying constant synthesis for batches of the beads synthesized until this stage or by attaching constant sequences of the gene tags to the end of the DNA attached to the bead by ligation elongation or chemical reaction.

[0129] Addition of pre- synthesized oligos to create the barcodes and the tags can be done by chemical or enzymatical (ligation for example) methods. The primer binding sequence is constant for all the beads and it can be followed by a random sequence - UMI. The UMI may contain a handle sequence to mediate the addition of the following sequence - the bead barcode. The bead barcode is composed of a certain number of sub-barcodes (96 in the length of 4 bp, for example) that will create diversity by the split and pool methodology. This approach might be beneficial in the creation of fluorescent probes in the later stage. Eventually, the gene tags will be added from a library of sequences.Example 2: Method for production of an error-free pool of oligos from a starting pool with errors

[0130] Having given this overview of the new production method, a step-by-step review of various methods of the invention are provided below. Reference is now made to Figure 3, which presents the basic concept behind the synthesis flow of the disclosed invention. The beads 301 (equivalent to bead 101 of Figure 1, indeed all bead components of this figure are equivalent to their counterparts in Figure 1) are produced in advance holding multiple instances of DNA segments 302, 303, 304 and 305. Oligo instances 308, 309 and 310 are separately synthesized. These oligos comprise 2 sections: a gene section and a gene tag.

[0131] These gene sections may be identical, i.e., holding the same gene segment (see 308 and 310), or may be different, i.e., holding a different segment of the same gene (see 309). Once the oligos are applied over the beads, in due time oligo instances are attracted via hydrogen bonds to corresponding beads in accordance with their gene tags. This means for example that oligo sequence 308 which holds a capture sequence 306 will attach to capturing sequence 305 of the bead as both sequences are complementary. Over time each bead will attract oligos with capture sequences complementary to the capturing sequences attached to the bead. At the end of this process beads should hold at least one full set of the segments relating to a given gene variant.

[0132] Beads holding a complete set of error- free gene segments (oligos) are identified and fluorescent selection probes or markers 311 corresponding to these error-free beads (e.g., theywill hybridize to the bead tag 304) are applied to the beads. These selection probes are used in the following steps to filter out bad beads and keep error-free good beads.

[0133] Figure 4 provides a step-by-step view of an embodiment of the method of the invention. In Step 1 a pool of oligos is generated that contains the gene of interest. Or if the gene is large, the pool contains fragments of the gene of interest that can be assembled to produce the full gene. This pool is made by the cheap and rapid methods known in the art. These methods are heavily error prone and such pools cannot be used for most downstream assays that rely on a correct sequence. The next steps of the method of the invention will take this error-heavy pool and produce an error- free oligo pool or an error free assembled sequence of the gene.

[0134] At the end of each oligo of the pool is a capture sequence that hybridizes to the capturing sequence at the end of nucleic acid molecule on the bead. Generally speaking, the capture sequence on the oligo is reverse complementary to the capturing sequence on the bead and they will thus hybridize by hydrogen binding between complementary bases. The sequences need not be 100% reverse complementary to each other, but rather need only be sufficiently complementary as to hybridize. If the nucleic acid molecule is linked to the bead in a 5’ to 3’ orientation, the capturing sequence will be at the 3’ end. The capture sequence on the oligo will also then be at the 3’ end of the oligo. These two sequences make up a capture sequence-capturing sequence pair.

[0135] In Step 2, the hybridization of the capture sequence on the oligo to the capturing sequence on the bead captures many oligos from the initial pool to each bead. For the sake of simplicity, a single captured oligo is shown. After oligos are captured to the bead, an extension / amplification is performed. By extending the molecule linked to the bead, a copy of the oligo (its reverse compliment) is made. This copy is also linked to the bead and the bead will now have many extended nucleic acid molecules attached to it, each one with a sequence (reverse complement) of an oligo from the initial pool. The hybridized oligo will of course also be extended and in so doing will gain the bead barcode, UMI and primer sequence (its reverse complement).

[0136] If only extension is performed, a single extended oligo will be made. However, as shown in Step 3, primers can be added to perform amplification (linear or exponential) and this will result in many copies of the oligo, capturing sequence, bead barcode, and UMI being produced. These copies will not be bound to the bead and can be sequenced. During sequencing oligo sequences that do not have any errors can be identified. The sequencing will also provide the bead barcode, so beads can be identified that bear error-free oligo sequences.When the gene of interest has been broken down into fragments on multiple oligos, beads can be identified with all of the fragments that make up the gene of interest and only beads with all the fragments which are all error free are selected.

[0137] As shown in Step 4, once a bead is validated, a probe is made which will hybridize to the bead barcode (reverse complement of the bead barcode). Since each bead has a unique bead barcode this probe will only bind to the selected bead. The probe will contain a selection moiety that will allow for selection of the good beads after hybridization of the probe. For simplicity, Figure 4 shows a fluorescent probe, though any moiety that allows for selection may be used. A single probe is shown hybridizing to the bead barcode region, but in reality each nucleic acid molecule of the bead has the barcode region and so many probes would actually bind the bead. For simplicity, only one selected bead is shown, but in fact many good beads without errors and ideally with all the fragments needed to assemble the gene of interest will be selected and probes will be designed for each selected beads barcode. The selection mode, however, can be the same for all the probes so that all the desired beads are pooled together.

[0138] In Step 5, the beads are sorted based on the selection moiety. Since a fluorescent probe is exemplified a fluorescent sorter is used and all the selected “good beads” are pooled together. Finally, in Step 6 the oligos can be removed from the beads (e.g., by cleavage at the cleavage site if it is included, or by amplification) and delivered as a pool of error-free oligos or assembled into the full gene of interest and delivered as an error free full-length gene.Example 3: Synthesis and assembly flow to produce a full target sequence from an error-free pool

[0139] Reference is now made to Figure 5 which presents the main steps in synthesis and assembly of genes according to the disclosed invention.

[0140] The process starts with oligo segment design 401 before any of the above steps are performed. The target gene variants are evaluated for their length. If too long, gene variants are split into shorter segments that can fit an oligo pool synthesis. The oligos include these gene segments which are appended with additional tags to help with the production flow in validation, sorting, and assembly. The result defines the content of an oligo pool that is synthesized 402 using methods known in the art (the cheap, fast, error-prone pool used as input for the method of the invention. The synthesis can be done with any method: chemical, phosphoramidite, or enzymatic, TdT, for example. The product of this stage can also be ordered from commercial vendors. Each oligo contains the following elements: a. Gene segment or a complete gene if segmentation is not needed.b. A recognition sequence (capture sequence) for a specific bead. c. Sequence for mediating assembly (For example a type IIS restriction enzyme recognition site) in downstream processes. d. Primer binding site (Illumina's Truseq read 2 for example) or a handle sequence to bind or add a primer.

[0141] Usually, oligos are created as single- stranded DNA (ssDNA). If amplification is applied the product of the amplification is double- stranded DNA (dsDNA). dsDNA can be ligated to the 3' or the 5' end of the ssDNA on the bead or be elongated with polymerase. ssDNA can also serve as a substrate for elongation or ligation.

[0142] Other advantages of dsDNA are the possibility to reduce errors by using the complement strand. This can contribute to high-quality sequences by methods that exploit both strands for error detection and elimination.

[0143] Beads that were previously prepared in accordance with the oligo pool design are mixed with the oligos 403. Hybridization of the oligos to the matching sequence (binding to the capturing sequence) on the bead and ligation (or polymerization in some embodiments or chemical attachment that enables amplification such as click chemistry) are taking place during this step. The oligos may undergo amplification before the mixing step and / or after appending the oligo sequence to the bead. The next step 404 includes amplification of the sequences after hybridization-elongation / ligation between the bead and the oligos. The amplification can be linear (more accurate, less efficient) or exponential (better amplification).

[0144] In step 405 the amplified tags with their attached oligos are sequenced. The sequencing data is used to identify beads which hold only error-free gene segments and have all the relevant segments attached to it. The tags of these beads are the selected barcode which are selected / isolated in the coming steps.

[0145] In step 406 fluorescent markers are constructed for the selected barcode. These markers are complementary to the bead tags. In step 407 these markers are added to the mixture of beads. The markers can be collected or injected from a pre-synthesized collection of markers or be synthesized ad-hoc for this purpose.

[0146] The marked beads containing sequenced and verified oligos are sorted in step 408 into an individual tube or well on a plate for further amplification and this is followed with assembly in step 409. Sorting is done based on the markers. If fluorescent markers are used, FACS can serve as a good solution for the sorting operation.

[0147] After amplification and assembly, the wells may undergo a 2ndsequencing cycle 410. The goals of this second sequencing cycle are:a. Find the identity of the oligos / bead in each well. b. Confirm the correct assembly. c. Further reduce the error rate.

[0148] At this stage the assembled sequence can be delivered, cloned, or subjected to further concatenation of the assembled products.

[0149] Optional step 411 may be used to trigger a smaller scale production cycle to recover for missing or errored gene variants. This additional (optional) production cycle may be also integrated with a new production cycle that may commence.Example 4: Description of an apparatus for performance of methods of the invention

[0150] Reference is now made to Figure 5 which presents a schematic description of a device implemented according to the disclosed invention. The device comprises multiple modules, all assembled on a robotic platform 507. The device makes use of an oligo pool including DNA segments. These segments may either be complete genes, if the genes are short enough, segments of long genes or any other DNA content. The oligo pool may either be an input to the device or printed by the device itself using the optional printing module 501. Tagged beads together with the oligo pool and the required enzymes for biochemical processes such as elongation and amplification are placed on the reagent rack module 505 and applied as an input to the thermomixer 502, which is capable for temperature control including thermocycling and mixing, in accordance with the requirements of the process. After the barcoded beads undergo elongation or other process to get the oligo sequence on the bead, the sample is amplified and taken from the thermomixer 502 for sequencing by module 503. Sequencing results are analyzed by the control unit 504. This control unit is responsible for managing the device. It instructs the robotic platform to transfer material or samples between the modules and is also responsible for configuring the different modules, i.e., the printer, and for analyzing sequencing results. The control unit may include the entire computational requirements of the process, or it may interact with additional computational servers. These sequencing results are evaluated and beads holding only error-free oligos are identified. In accordance with these identified beads, their related bead tags are read from the control unit database. The printer 501 is then configured by the control unit 504 to print markers which are complementary to the identified bead tags (having only error-free oligos attached to them) or the markers can be selected from an already produced library of markers corresponding to the bead tags. The markers are mixed with the beads with or without further amplification. The result is run through the selection module 506. This module may include a microfluidics selection device (with an optical camera for example). The output of the selection moduleholds error-free beads and is then sorted and collected. The beads may be placed separately one by one, or in a pool, for further use. The thermomixer 502 may receive the sorted library for further amplification and sequencing for identification and isolation of the molecules of the bead. Another option is that assembly will be followed by amplification and / or sequencing.

[0151] It should be noted that the device may make use of various kinds of beads, implementing the bead functionality, as previously explained and referred to in general as G- beads.Example 5: Additional embodiments of the invention

[0152] Accurate oligo pool without beads: This embodiment of the disclosed invention produces an accurate oligo pool from an oligo pool produced by methods known in the art. Current art oligo pool production is limited in its error rate (relatively high error-rate) which limits the ability of using it for gene synthesis. This embodiment does not involve beads. The oligos are synthesized as follows: primer binding sitel - large UMI (>20b) - cleavage site - desired sequence - cleavage site - primer binding site 2 (optional).

[0153] The oligo pool is amplified and sequenced. Correct sequences are identified according to their UMI sequence. A probe is designed from the UMI sequence (the UMI sequence may be larger than the probe sequence to enable flexibility in the probe sequence). Biotinylated probes are produced and hybridize to the oligos with correct sequence and are selected with streptavidin coated magnetic beads or other selection procedures. Then, the desired sequence is separated at the cleavage sites and the oligo pool can be assembled or delivered.

[0154] Accurate oligo pool with beads: This embodiment of the disclosed invention produces an accurate oligo pool from an oligo pool produced by methods known in the art with their rather high error-rate. This embodiment makes use of beads. In this embodiment the assembly phase is skipped, and a high-quality oligo pool is produced.

[0155] The process starts with marking a population of DNA molecules with a tag. The tag has a double purpose - identification of the molecule through sequencing, and mediating selection of the molecule. The tag can be for a single molecule or used to join a plurality of DNA molecules that can be later assembled post-selection to a longer DNA, such as a gene.

[0156] The tag enables the identification of the molecule through sequencing and selection of the molecule. After identification of the DNA molecules with a correct sequence, molecules with a correct sequence are selected with a probe targeting the tag. The selection is done on individual molecules, a pool of individual molecules, or a set of molecules that can be subjected to downstream processes, such as assembly, or a pool of sets of molecules.

[0157] Accurate pools of long oligos: This embodiment of the disclosed invention produces an accurate long oligo pool from a long oligo pool produced by methods known in the art. Current oligos pools are limited in their length to 200-300 bases. The error rate for longer oligos, e.g., 1,000 bases long, is very high making it not feasible for practical use.

[0158] With this embodiment of the disclosed invention, a high error rate long oligo pool is the starting input, and the disclosed method is used to efficiently filter out errored (long) oligos, resulting in a high-quality error-free long oligo pool.

[0159] Gene assembly methods: This embodiment of the disclosed invention relates to the assembly phase. It is disclosed that different assembly options are available: a. Gene assembly in a pool. With this approach, different gates of the golden gate assembly have to be specified. This requires a specific design for gene variants to ensure gate specificity and prevent gate overlap; b. Gene assembly as individual genes with microwell assembly. With this embodiment, correct beads carrying segments to be assembled are placed into array of microwells. The assembly process takes place in each well independently. Then a sample from each specific product from a specific well is collected and sequence verified. The desired wells are collected one by one; c. Gene assembly as individual genes in emulsion assembly.

[0160] Although the invention has been described in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications and variations will be apparent to those skilled in the art. Accordingly, it is intended to embrace all such alternatives, modifications and variations that fall within the spirit and broad scope of the appended claims.

Claims

CLAIMS:

1. A method of enriching a target sequence, the method comprising: a. receiving sequencing reads derived from a plurality of nucleic acid molecules, wherein nucleic acid molecules of said plurality each comprise a barcode sequence, and wherein said plurality comprises a nucleic acid molecule comprising said target sequence and a nucleic acid molecule devoid of said target sequence; b. identifying in said sequencing reads a barcode sequence that is present only in a nucleic acid molecule of said plurality that comprises said target sequence; c. contacting said plurality of nucleic acid molecules with a selection probe under conditions and for a sufficient time to hybridize said selection probe to said identified barcode sequence to produce a probe-bound target nucleic acid molecule, wherein said selection probe comprises: i. a sequence complementary to said identified barcode sequence; and ii. a selection moiety; d. selecting from said plurality of nucleic acid molecule said probe-bound target nucleic acid molecule by selecting said selection moiety; thereby enriching a target sequence.

2. The method of claim 1, wherein said method is a method of enriching a sequence- verified target sequence and wherein said nucleic acid molecules devoid of said target sequence comprise said target sequence with errors.

3. The method of claim 2, wherein a sequence verified target sequence is a target sequence with an error rate below a predetermined threshold and said target sequence with errors is a target sequence with an error rate above a predetermined threshold.

4. The method of any one of claims 1 to 3, comprising before step (a): a. providing a plurality of solid supports, wherein each solid support is conjugated to a plurality of nucleic acid molecules, wherein each nucleic acid molecule comprises: i. a barcode sequence specific to said solid support and not shared by other solid supports of said plurality; and ii. a capturing sequence; b. providing a pool of nucleic acid molecules comprising a capture sequence, wherein said provided pool comprises a nucleic acid molecule comprising said target sequence and a nucleic acid molecule devoid of said target sequence, andwherein said capturing sequence and said capture sequence are complementary sequences capable of hybridizing together; c. contacting said provided pool of nucleic acid molecules to said provided plurality of solid supports under conditions and sufficient time to hybridize said capture sequence to said capturing sequence to produce at least one hybridized solid support; d. contacting said plurality of hybridized solid supports with a polymerase to elongate ends of nucleic acid molecules to produce a plurality of elongated solid supports; e. dissociating elongated nucleic acid molecules comprising said capture sequence from elongated nucleic acid molecule comprising said capturing sequence and conjugated to said solid support; and f. sequencing said dissociated elongated nucleic acid molecule comprising said capture sequence to produce said sequencing reads.

5. The method of claim 4, wherein said contacting said plurality of nucleic acid molecules with a selection probe is contacting said elongated nucleic acid molecules comprising said capturing sequence and conjugated to said solid support with said selection probe.

6. The method of claim 4 or 5, wherein said solid supports are beads.

7. The method of any one of claims 4 to 6, further comprising amplifying said elongated nucleic acid molecules to produce copies of said elongated nucleic acid molecules and wherein said sequencing further comprises sequencing said copies.

8. The method of any one of claims to 4 to 7, wherein said plurality of nucleic acid molecules a. is between 100-10,00,000 nucleic acid molecules per solid support; b. are covalently linked to said solid support; c. are conjugated to said solid support at a 5’ end of each nucleic acid molecule of said plurality; or d. a combination thereof.

9. The method of any one of claims 4 to 8, wherein each nucleic acid molecule of said plurality of nucleic acid molecules is conjugated to said solid support at a 5’ end and wherein said barcode sequence is 5’ to said capturing sequence, optionally, wherein each nucleic acid molecule comprises a unique molecular identifier (UMI) sequence and wherein said UMI sequence is 5’ to said capturing sequence.

10. The method of any one of claims 1 to 3, comprising before step (a):a. providing a pool of nucleic acid molecules comprising a nucleic acid molecule comprising said target sequence and a nucleic acid molecule devoid of said target sequence, and wherein each oligonucleotide of said pool comprises a barcode sequence that is specific to said nucleic acid molecule and is not shared by other nucleic acid molecules of said pool; b. amplifying said pool of nucleic acid molecules to produce copies of said nucleic acid molecules; and c. sequencing at least a portion of said copies to produce said sequencing reads.

11. The method of claim 10, wherein said contacting said plurality of nucleic acid molecules with a selection probe is contacting said pool of nucleic acid molecules, contacting a portion of said copies that were not sequenced or both.

12. The method of any one of claims 1 to 11, wherein each nucleic acid molecule of said plurality of nucleic acid molecules further comprises a UMI sequence.

13. The method of any one of claims 1 to 12, wherein said identifying comprises identifying a plurality of continuous nucleotide sequences that comprise said target sequence and a barcode sequence and wherein step (c) comprises contacting said plurality of nucleic acid molecules with a plurality of selection probes, wherein said plurality of selection probes comprises a probe to each identified barcode sequence.

14. The method of any one of claims 1 to 13, wherein said selection moiety is a detectable moiety or isolatable moiety selected from a fluorophore, a radioactive moiety, a nanopore-detectable moiety, biotin, a click reaction substrate, an antigen, a magnetic moiety and a bead.

15. The method of claim 14, wherein said selection moiety is a fluorophore.

16. The method of claim 15, wherein said enriching comprises fluorescence activated cell sorting (FACS).

17. The method of any one of claims 1 to 16, wherein said plurality of nucleic acid molecules comprising a nucleic acid molecule comprising a target sequence comprises fragments of said target sequence that when assembled produce said target sequence.

18. The method of claim 17, wherein said identifying comprises identifying a barcode sequence that is present in continuous nucleotide sequences that comprise all of the fragments needed to assemble said target sequence.

19. The method of claim 17 or 18, further comprising after said selecting, assembling selected target sequence fragments into a complete target sequence.

20. The method of any one of claims 1 to 19, wherein said target sequence is a sequence of a gene of interest.

21. A system comprising: a. a thermomixer; b. a DNA sequencer; c. a control unit to analyze DNA sequencing results from said DNA sequencer, and select a barcode sequence or its complement sequence to be part of a selection probe; and d. a computer program product comprising a non-transitory computer-readable storage medium having program code embodied thereon, the program code executable by at least one hardware processor to apply said selection probe to said thermomixer.

22. The system of claim 21, wherein said computer program product comprises program code executable by at least one hardware processor to perform the method of any one of claims 1 to 20.

23. The system of claim 21 or 22, further comprising a sorting device capable of sorting solid supports comprising said selection probe from solid supports devoid of said selection probe.

24. The system of claim 23, wherein said sorting device is a FACS machine.

25. The system of any one of claims 21 to 24, further comprising an oligo or DNA molecule printer capable of producing said selection probe.

Citation Information

Patent Citations

  • Nucleic acid library construction method, obtained nucleic acid library and use thereof

    CN111379031B

  • Method for producing insulating glass unit and method for producing glass window

    US20200010361A1

  • AU2021204364A1