Multiplexed generation and barcoding of genetically engineered cells

Through RNA-guided nuclease and barcoding technology, precise editing of target chromosomal sites and generation of high-throughput variant libraries are achieved, solving the problem of insufficient genome editing efficiency and flexibility in the prior art.

CN111344403BActive Publication Date: 2025-05-06THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV +1
View PDF 49 Cites 0 Cited by

Patent Information

Application Number
CN201880073960.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-09-15
Filing Date
2018-09-14
Publication Date
2025-05-06
Estimated Expiration
2038-09-14

AI Technical Summary

Technical Problem

The prior art is difficult to achieve efficient, flexible precise editing and generation of high-throughput variant libraries in genome editing, especially in metazoan cells.

Method used

The RNA-guided nuclease and barcode multivariate high-throughput genome editing system was used to accurately edit the target chromosomal loci through homologous mediation repair and integrate the genomic barcode at different chromosomal loci to identify and validate individual variants.

Benefits of technology

Efficient and accurate genome editing in multiple cells is achieved, and variants are verified in large scale parallelly through barcode programming methods, improving the flexibility and efficiency of genome editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111344403B_ABST
    Figure CN111344403B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the use of RNA-guided nucleases and genomic barcode multiplexity to produce genetically engineered cells and perform phenotyping. Specifically, a system that promotes precise genome editing at the desired target chromosomal locus through homology-mediated repair is used to achieve high-throughput multiplex genome editing. The guide RNA and donor DNA sequences are integrated as genomic barcodes at different chromosomal loci, allowing the identification, isolation and large-scale parallel verification of individual variants from the transformant pool. Strains can be arranged according to their precise genetic modifications, as specified by the incorporation of donor DNA into heterologous or natural genes. The present disclosure also relates to a method for codon editing outside a typical guide RNA recognition region, which enables full saturation mutation of protein-coding genes, a marker-based internal cloning method that removes background caused by oligonucleotide synthesis errors and incomplete vector backbone cutting, and a method for improving homology-mediated repair by active donor recruitment.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Application No. 62 / 559,493, filed September 15, 2017, which is incorporated herein by reference in its entirety.

[0003] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0004] This invention was made with Government support under Contract HG000205 awarded by the National Institutes of Health and Contract 70NANB15H268 awarded by the National Institute of Standards and Technology. The Government has certain rights in the invention. Technical Field

[0005] The present disclosure generally relates to the field of genome engineering using RNA-guided nucleases. In particular, the present disclosure relates to compositions and methods for high-throughput generation and verification of genetically engineered cells using RNA-guided nucleases and barcode multiplexing. Background of the Invention

[0007] The advent of programmable genome editing via the CRISPR / Cas9 system has enabled rapid advancement of synthetic biology and genetic engineering. The Streptococcus pyogenes type II clustered regularly interspaced short palindromic repeats (CRISPR)-associated protein 9 (Cas9) is the first RNA-guided nuclease (RGN) shown to cleave any genomic location using a guide RNA (gRNA) with homology to the target region. 2 . Using the conservative homologous recombination-based DNA repair pathway present in the host cell, donor DNA with homology flanking the cleavage site can be used to repair the break and introduce the genetic change of interest. The short specificity determining region of the gRNA (typically 20 nucleotides in length) and the length of the donor DNA (~100-150 nt) are compatible with highly parallel array-based oligonucleotide library synthesis, enabling the easy generation of gRNA-donor libraries against thousands of targets. 1,3-8 However, to date, the generation of variant libraries has been limited to pools, which greatly limits the options for characterizing the phenotype of individual variants. For example, microscopy, metabolomics, and many enzyme assay reporter molecules are not suitable for pooled formats.

[0008] CRISPR editing is particularly efficient in yeast because it strongly prefers to use homologous recombination (HR) to repair double-strand breaks in the presence of donor DNA, eliminating the need for selectable markers when editing the genome. 9-11 In contrast to the nearly 100% Cas9 editing efficiency reported in yeast 12-14Gene editing in metazoan cells is perturbed by preferential nonhomologous end joining (NHEJ) over HR, and HR editing in human cells only reaches a maximum efficiency of approximately 10-60%. 15,16 Thus, in addition to extending the long-established use of yeast as a model system for eukaryotic biology, the Cas9 system amplifies the value of yeast as a host for engineering heterologous proteins and pathways.

[0009] Therefore, there remains a need for more efficient and flexible genome editing methods that enhance the repair of RGN-mediated double-strand breaks via the HR mechanism to allow modification of the genome with the precise genetic changes desired, as well as improved methods for high-throughput generation of variant libraries. Summary of the invention

[0010] The present disclosure relates to multiplex generation and verification of genetically engineered cells using RNA-guided nucleases and barcodes. Specifically, high-throughput multiplex genome editing is achieved with a system that promotes precise genome editing at the desired target chromosomal site through homology-mediated repair. Guide RNA and donor DNA sequences are integrated as genomic barcodes at chromosomal loci different from the target locus, allowing easy identification, separation and large-scale parallel verification of individual variants from a pool of transformants. Strains can be arranged according to their precise genetic modifications, as specified by the incorporation of donor DNA into heterologous or natural genes. The present disclosure also relates to a method for codon editing outside a typical guide RNA recognition region, which enables fully saturated mutations of protein-coding genes, a marker-based internal cloning method that removes background caused by oligonucleotide synthesis errors and incomplete vector backbone cutting, and a method for improving homology-mediated repair by active donor recruitment.

[0011] Provided herein is a method for multiplexing and generating genetically engineered cells, the method comprising: (a) transfecting a plurality of cells with a plurality of different recombinant polynucleotides, each recombinant polynucleotide comprising a genome editing cassette, the genome editing cassette comprising a first nucleic acid sequence encoding a first guide RNA (gRNA) that can hybridize at a genomic target locus to be modified and a donor polynucleotide, thereby forming a gRNA-donor polynucleotide combination, wherein each recombinant polynucleotide comprises a different genome editing cassette, the different genome editing cassettes comprising a different gRNA-donor polynucleotide combination, and allowing each cell to express the first nucleic acid sequence, thereby forming a gRNA; and (b) introducing an RNA-guided nuclease into each of the plurality of cells, wherein the RNA-guided nuclease in each cell forms a complex with the gRNA, thereby forming a gRNA-RNA-guided nuclease complex, and allowing the gRNA-RNA-guided nuclease complex to modify the genomic target locus by integrating the donor polynucleotide into the genomic target locus, thereby generating a plurality of genetically engineered cells.

[0012] On the other hand, a method for generating genetically engineered cells in a multiplex manner is provided, the method comprising: (a) transfecting a plurality of cells with a plurality of different recombinant polynucleotides, each recombinant polynucleotide comprising a unique polynucleotide barcode and a genome editing cassette comprising a first nucleic acid sequence and a donor polynucleotide, the first nucleic acid sequence encoding a first guide RNA (gRNA) that can hybridize at a genomic target locus to be modified, thereby forming a gRNA-donor polynucleotide combination, wherein each recombinant polynucleotide comprises a different genome editing cassette, the different genome editing cassettes comprise different gRNA-donor polynucleotide combinations, and allowing each cell to express the first nucleic acid sequence, thereby forming a gRNA; and (b) introducing an RNA-guided nuclease into each of the plurality of cells, wherein the RNA-guided nuclease in each cell forms a complex with the gRNA, thereby forming a gRNA-RNA-guided nuclease complex, and allowing the gRNA-RNA-guided nuclease complex to modify the genomic target locus by integrating the donor polynucleotide into the genomic target locus, thereby generating a plurality of genetically engineered cells.

[0013] In an embodiment, the method further comprises sequence verification and arraying of a plurality of genetically engineered cells, the method comprising: (c) inoculating a plurality of genetically engineered cells in an ordered array in a culture medium suitable for the growth of genetically engineered cells; (d) culturing a plurality of genetically engineered cells under certain conditions, wherein each genetically engineered cell generates a clonal colony in the ordered array; (e) introducing a genome editing cassette from the ordered array colony into a barcoded cell, wherein the barcoded cell comprises a nucleic acid, the nucleic acid comprises a recombination target site of a site-specific recombinase, and a barcode sequence, the barcode sequence identifying the position of the colony in the ordered array corresponding to the genome editing cassette; (f) using a site-specific The heterotropic recombinase system transposes the genome editing cassette to a position adjacent to the barcode sequence of the barcoded cell, wherein site-specific recombination with the recombination target site of the barcoded cell produces a nucleic acid comprising the barcode sequence linked to the genome editing cassette; (g) sequencing the nucleic acid comprising the barcode sequence of the barcoded cell linked to the genome editing cassette to identify the guide RNA sequence and the donor polynucleotide sequence from the genome editing cassette in the colony, wherein the barcode sequence of the barcoded cell is used to identify the colony position in the ordered array from which the genome editing cassette originates; and (h) picking a clone comprising the genome editing cassette of the colony in the ordered array identified by the barcode of the barcoded cell.

[0014] On the other hand, a method for localizing a donor polynucleotide to a genomic target locus in a cell is provided, the method comprising: (a) transfecting a cell with a recombinant polynucleotide, the recombinant polynucleotide comprising a genomic editing box comprising a donor polynucleotide and a DNA binding sequence known to bind a DNA binding domain; (b) introducing a nuclease into the cell, wherein the nuclease recognizes and causes double-stranded DNA breaks at the genomic target locus; (c) introducing a donor recruiting protein into the cell, the donor recruiting protein comprising a DNA binding domain and a DNA break site localization domain and allowing the donor recruiting protein to selectively recruit double-stranded DNA breaks, thereby localizing the donor polynucleotide to the genomic target locus.

[0015] On the other hand, a library of gene editing vectors is provided, each gene editing vector comprising a genome editing box, including (i) a barcode, (ii) a first nucleic acid sequence encoding a first guide RNA (gRNA) that can hybridize at a target locus of the genome to be modified, and (iii) a donor polynucleotide, thereby forming a barcode-gRNA-donor polynucleotide combination; wherein each recombinant polynucleotide comprises a different genome editing box comprising a different barcode-gRNA-donor polynucleotide combination.

[0016] In another aspect, a gene editing vector is provided that includes a donor polynucleotide and a first nucleic acid sequence that encodes a first guide RNA (Guide X) that can hybridize with the vector at a target site, such that when the cell expresses Guide X, Guide X hybridizes with the vector and generates a double-stranded DNA break at the target site.

[0017] In another aspect, a kit is provided, comprising: (a) a gene editing vector described herein, including embodiments thereof; and (b) a nuclease or a polynucleotide encoding a nuclease.

[0018] In another aspect, a kit is provided, comprising: (a) a gene editing vector described herein, including embodiments thereof; and (b) a reagent for genetically modifying a cell.

[0019] On the other hand, a library of gene editing vectors is provided, each gene editing vector comprising a genome editing cassette comprising: (i) a first nucleic acid sequence encoding a first guide RNA (gRNA) that can hybridize at a target locus of the genome to be modified, and (ii) a donor polynucleotide, thereby forming a gRNA-donor polynucleotide combination; wherein each recombinant polynucleotide comprises a different genome editing cassette comprising a different gRNA-donor polynucleotide combination.

[0020] In embodiments, each recombinant polynucleotide further comprises a second nucleic acid sequence encoding an RNA-guided nuclease.

[0021] In one aspect, the present disclosure includes a method for multiplexed genetic modification and barcoding of a cell, the method comprising: a) providing a plurality of recombinant polynucleotides, wherein each recombinant polynucleotide comprises a genome editing cassette, the genome editing cassette comprising a polynucleotide encoding a guide RNA (gRNA) that can hybridize at a genomic target locus to be modified and a donor polynucleotide, the donor polynucleotide comprising a 5' homology arm that hybridizes to a 5' genomic target sequence and a 3' homology arm that hybridizes to a 3' genomic target sequence, the 5' homology arm and the 3' homology arm flanking a nucleotide sequence comprising a desired edit to be integrated into the genomic target locus, wherein each recombinant polynucleotide comprises a different genome editing cassette comprising a different guide RNA-donor polynucleotide combination, such that the plurality of recombinant polynucleotides can produce a plurality of different desired modifications at one or more genomic target loci. Editing; and (b) transfecting cells with a plurality of recombinant polynucleotides; c) culturing the transfected cells under conditions suitable for transcription, wherein guide RNA is generated from each genome editing cassette; d) introducing RNA-guided nucleases into the cells, wherein the RNA-guided nucleases form a complex with the guide RNA generated in the cells, the guide RNA directing the complex to one or more genomic target loci, wherein the RNA-guided nucleases generate double-strand breaks in the genomic DNA of the cells at the one or more genomic target loci, and the donor polynucleotides present in each cell are integrated at the genomic target loci identified by its 5' homology arm and the 3' homology arm by homology-directed repair (HDR), thereby generating a plurality of genetically modified cells; and e) barcoding the plurality of genetically modified cells by integrating the genome editing cassettes present in each genetically modified cell at the chromosomal barcode locus. In certain embodiments, the method further comprises performing additional rounds of genetic modification and genomic barcoding on the genetically modified cells by repeating steps (a)-(e) using different genome editing cassettes.

[0022] In certain embodiments, each recombinant polynucleotide is provided by a vector. For example, the vector can be a plasmid or a viral vector. In certain embodiments, the vector is a high copy number vector.

[0023] In certain embodiments, each recombinant polynucleotide is provided as a linear DNA. For example, the method further comprises amplifying a recombinant polynucleotide comprising a genome editing cassette, which is provided as a PCR product.

[0024] In certain embodiments, the RNA-guided nuclease is also provided by a vector. In certain embodiments, the genome editing cassette and RNA-guided nuclease are provided by a single vector or separate vectors. In another embodiment, the recombinant polynucleotide encoding the RNA-guided nuclease is integrated into the genome of the host cell.

[0025] Transcription of the guide RNA generally depends on the presence of a promoter, which can be incorporated into a genome editing box, or incorporated into a vector or a chromosomal locus (such as a chromosomal barcode locus), into which the genome editing box is inserted. The promoter can be a constitutive or inducible promoter. In certain embodiments, each genome editing box includes a promoter operably connected to a guide RNA encoding polynucleotide. In other embodiments, the chromosomal barcode locus includes a promoter that is operably connected to a polynucleotide encoding any genome editing box guide RNA, and the genome editing box is integrated at the chromosomal barcode locus. In another embodiment, each recombinant polynucleotide is provided by a vector, wherein the vector includes a promoter operably connected to a guide RNA encoding polynucleotide.

[0026] In certain embodiments, the multiple recombinant polynucleotides can generate mutations at multiple sites within a single gene. In other embodiments, the multiple recombinant polynucleotides can generate mutations at multiple sites in different genes or genomes anywhere. For example, each donor polynucleotide can introduce different mutations into a gene, such as insertions, deletions or substitutions. In another embodiment, at least one donor polynucleotide introduces a mutation that inactivates a gene. In another embodiment, at least one donor polynucleotide removes a mutation from a gene. In another embodiment, at least one donor polynucleotide inserts an accurate genetic change in genomic DNA.

[0027] In certain embodiments, the integration of the genome editing box present in the genetically modified cell at the chromosome barcode locus is implemented with HDR. Each recombinant polynucleotide may also include a pair of universal homology arms flanking the genome editing box, which can hybridize the complementary sequence at the chromosome barcode locus to allow the genome editing box to be integrated at the chromosome barcode locus by HDR. In addition, each recombinant polynucleotide may also include a second guide RNA that can hybridize at the chromosome barcode locus, wherein the RNA-guided nuclease further forms a complex with the second guide RNA, and the second guide RNA guides the complex to the chromosome barcode locus, wherein the RNA-guided nuclease forms a double-strand break at the chromosome barcode locus, and the genome editing box is integrated into the chromosome barcode locus by HDR.

[0028] In other embodiments, the genome editing box present in the genetically modified cell is integrated at the chromosome barcode locus, and is implemented with a site-specific recombinase system. Exemplary site-specific recombinase systems include Cre-loxP site-specific recombinase system, Flp-FRT site-specific recombinase system, PhiC31-att site-specific recombinase system, and Dre-rox site-specific recombinase system. In certain embodiments, the chromosome barcode locus also includes a first recombination target site of the site-specific recombinase and the recombinant polynucleotide also includes a second recombination target site of the site-specific recombinase, and the site-specific recombination between the first recombination target site and the second site-specific recombination site causes the genome editing box to be integrated at the chromosome barcode locus.

[0029] In certain embodiments, the method further comprises using a selectable marker that selects clones that have undergone successful integration of the donor polynucleotide at the genomic target locus or successful integration of the genome editing cassette at the chromosomal barcode locus.

[0030] In certain embodiments, the cell to be genetically modified is a eukaryotic or prokaryotic cell. In some embodiments, the cell is a yeast cell, which may be a haploid or diploid yeast cell.

[0031] In certain embodiments, each recombinant polynucleotide also includes a pair of restriction sites flanking the genome editing box. In some embodiments, the restriction sites are recognized by a meganuclease (such as SceI) that produces a double-strand break in the DNA. The expression of the meganuclease can be controlled by an inducible promoter.

[0032] In another embodiment, the genome editing cassette further comprises a tRNA sequence at the 5' end of the nucleotide sequence encoding the guide RNA.

[0033] In another embodiment, the genome editing cassette further comprises a nucleotide sequence encoding a hepatitis delta virus (HDV) ribozyme at the 5' end of the nucleotide sequence encoding the guide RNA.

[0034] In another embodiment, the RNA-guided nuclease is a Cas nuclease (such as Cas9 or Cpf1) or an engineered RNA-guided FokI-nuclease.

[0035] In another embodiment, the genome editing cassette is flanked by restriction sites recognized by a meganuclease.

[0036] In certain embodiments, each genome editing box also includes a unique barcode sequence for identifying the guide RNA and donor polynucleotides encoded by each genome editing box. The unique barcode can replace the guide RNA and the donor polynucleotide for sequencing to identify the genetic modification of the cell. In another embodiment, the method also includes deleting the polynucleotide encoding the guide RNA and the donor polynucleotide integrated at the chromosome barcode locus, while retaining the unique barcode at the chromosome barcode locus representing the deleted sequence. In another embodiment, the method also includes sequencing the barcode of the chromosome barcode locus of at least one genetically modified cell to identify the genome editing box used to genetically modify the cell.

[0037] In certain embodiments, the method further comprises sequencing each genome editing cassette. The genome editing cassette can be sequenced to link the barcode to a specific gRNA-donor polynucleotide combination, for example, in an intermediate cloning step, and then the genome editing cassette is linked to a vector or then transfected into a cell. Alternatively or additionally, sequencing the genome editing cassette integrated at the chromosomal barcode locus can be used to determine the genome editing performed on the genetically modified cell.

[0038] In certain embodiments, the method further comprises sequence verification and arranging a plurality of genetically modified cells, the method comprising: a) placing a plurality of genetically modified cells in an ordered array in a culture medium suitable for growth of the genetically modified cells; b) culturing the plurality of genetically modified cells under certain conditions so that each genetically modified cell produces a clonal colony in the ordered array; c) introducing a genome editing cassette from a colony in the ordered array into a barcoded cell, wherein the barcoded cell comprises a nucleic acid comprising a recombination target site for a site-specific recombinase, and a barcode sequence, the barcode sequence identifying the position of the colony in the ordered array corresponding to the genome editing cassette; d) culturing the plurality of genetically modified cells in an ordered array ... with a site-specific recombinase sequence; The method further comprises the steps of: systematically translocating a genome editing cassette to a position adjacent to a barcode sequence of a barcoded cell, wherein site-specific recombination with a recombination target site of the barcoded cell can generate a nucleic acid comprising the barcode sequence linked to the genome editing cassette; e) sequencing the nucleic acid comprising the barcode sequence of the barcoded cell linked to the genome editing cassette to identify a guide RNA sequence and a donor polynucleotide sequence from the genome editing cassette in a colony, wherein the barcode sequence of the barcoded cell is used to identify a colony position in an ordered array from which the genome editing cassette originates; and f) picking a clone comprising the genome editing cassette from a colony in an ordered array that is identified by the barcode of the barcoded cell.

[0039] For example, the genetically modified cell can be a haploid yeast cell, and the barcoded cell can be a diploid yeast cell that can be joined to the genetically modified cell, wherein a genome editing box from an ordered array of genetically modified haploid yeast colonies is introduced into the barcoded haploid yeast cell, including joining a haploid yeast clone from the colony with a barcoded haploid yeast cell to generate a diploid yeast cell. As described herein, subsequent site-specific recombination produces a nucleic acid containing a barcode sequence that is connected to a genome editing box in a diploid yeast cell. The genetically modified cell can be strain MATα and the barcoded yeast cell can be strain MATa. Alternatively, the genetically modified cell can be strain MATa and the barcoded yeast cell can be strain MATα.

[0040] In certain embodiments, the recombinase system of the barcoded cell is a Cre-loxP site-specific recombinase system, a Flp-FRT site-specific recombinase system, a PhiC31-att site-specific recombinase system or a Dre-rox site-specific recombinase system. In one embodiment, the recombination target site of the barcoded cell includes a loxP recombination site.

[0041] In another embodiment, the recombinase system of the barcode cell uses a meganuclease to generate DNA double-strand breaks. In another embodiment, the meganuclease of the barcode cell is a galactose-inducible SceI meganuclease. In another embodiment, the genomic expression cassette is flanked by restriction sites recognized by a meganuclease.

[0042] In another embodiment, the method further comprises using a selectable marker to select clones that have undergone successful site-specific recombination.

[0043] In certain embodiments, the method further comprises inhibiting non-homologous end joining (NHEJ). For example, NHEJ can be inhibited by contacting the cells with a small molecule inhibitor selected from wortmannin and Scr7. Alternatively, RNA interference or CRISPR interference can be used to inhibit the expression of protein components of the NHEJ pathway.

[0044] In other embodiments, the method further comprises using an HDR enhancer or active donor recruitment to increase the HDR frequency of the cell.

[0045] In another embodiment, the method further comprises using a selectable marker to select clones that have undergone successful integration of the donor polynucleotide at one or more genomic target loci via HDR.

[0046] In another embodiment, the method further comprises phenotyping at least one clone in the ordered array.

[0047] In another embodiment, the method further comprises performing whole genome sequencing on at least one clone in the ordered array.

[0048] In another embodiment, the method further comprises repeating steps (a)-(e) for all colonies in the ordered array to identify the sequences of the guide RNA and donor polynucleotide of the genome editing cassette for each colony in the ordered array.

[0049] In another aspect, the disclosure includes ordered arrays of colonies comprising clones of genetically modified cells generated by the methods described herein, wherein the colonies are indexed according to verified sequences of their guide RNA and donor polynucleotides.

[0050] In another aspect, the present disclosure includes a kit for multiplex genetic modification and barcoding of cells, the kit comprising: a) a plurality of recombinant polynucleotides, wherein each recombinant polynucleotide comprises a genome editing cassette comprising a polynucleotide encoding a guide RNA (gRNA) and a donor polynucleotide, the gRNA being capable of hybridizing at a genomic target locus to be modified, the donor polynucleotide comprising a 5' homology arm that hybridizes to a 5' genomic target sequence and a 3' homology arm that hybridizes to a 3' genomic target sequence, the 5' homology arm and the 3' homology arm flanking a nucleotide sequence comprising a desired edit to be integrated into the genomic target locus, wherein each recombinant polynucleotide comprises a different genome editing cassette comprising a different guide RNA-donor polynucleotide combination, such that the plurality of recombinant polynucleotides are capable of producing a plurality of different desired edits at one or more genomic target loci; and b) an RNA-guided nuclease; and c) a cell comprising a chromosomal barcode locus, wherein the barcode locus comprises an integration site for at least one recombinant polynucleotide genome editing cassette. The kit also includes other reagents and instructions for performing genome editing and barcoding as described herein.

[0051] In certain embodiments, each recombinant polynucleotide in the kit further comprises a pair of universal homology arms flanking the genome editing cassette, which can hybridize complementary sequences at the integration site of the chromosomal barcode locus to allow the genome editing cassette to be integrated at the chromosomal barcode locus by homology-directed repair (HDR).

[0052] In another embodiment, each recombinant polynucleotide further comprises a second guide RNA capable of hybridizing at the chromosomal barcode locus.

[0053] In certain embodiments, the kit further comprises a site-specific recombinase system (such as a Cre-loxP site-specific recombinase system, a Flp-FRT site-specific recombinase system, a PhiC31-att site-specific recombinase system, or a Dre-rox site-specific recombinase system). In another embodiment, the chromosomal barcode locus further comprises a first recombination target site for a site-specific recombinase and the recombinant polynucleotide further comprises a second recombination target site for a site-specific recombinase, so that site-specific recombination can occur between the first recombination target site and the second recombination target site to allow integration of the genome editing box at the chromosomal barcode locus.

[0054] In another embodiment, the RNA-guided nuclease in the kit is a Cas nuclease (such as Cas9 or Cpf1) or an engineered RNA-guided FokI-nuclease.

[0055] In certain embodiments, the kit further comprises a fusion protein designed to accomplish donor recruitment as described herein. Such a fusion protein comprises a polypeptide comprising a nucleic acid binding domain linked to a protein that selectively binds to DNA breaks produced by an RNA-guided nuclease. In another embodiment, the donor polynucleotide further comprises a nucleotide sequence having sufficient complementarity to hybridize to a sequence adjacent to the DNA break, and a nucleotide sequence comprising a binding site recognized by the nucleic acid binding domain of the fusion protein. In certain embodiments, the nucleic acid binding domain is a LexA DNA binding domain and the binding site is a LexA binding site, or the nucleic acid binding domain is a forkhead homolog 1 (FKH1) and the DNA binding domain and the binding site is a FKH1 binding site. In some embodiments, the nucleic acid binding domain-containing polypeptide further comprises a forkhead-associated (FHA) phosphothreonine binding domain. In another embodiment, the nucleic acid binding domain-containing polypeptide comprises a LexA DNA binding domain linked to a FHA phosphothreonine binding domain.

[0056] In another aspect, the present disclosure includes a method for promoting homology-directed repair (HDR) by actively recruiting donors to DNA breaks, the method comprising: a) introducing into a cell a donor-recruiting protein comprising a polypeptide that selectively binds to DNA breaks linked to a polypeptide comprising a nucleic acid binding domain; and b) introducing into the cell a donor polynucleotide comprising i) a nucleotide sequence that is sufficiently complementary to hybridize to a sequence adjacent to the DNA break and ii) a nucleotide sequence comprising a binding site recognized by a nucleic acid binding domain of a fusion protein, wherein the nucleic acid binding domain selectively binds to the binding site on the donor polynucleotide to produce a complex between the donor polynucleotide and the fusion protein, thereby recruiting donors to the DNA break and promoting HDR. In one embodiment, the donor-recruiting protein is a fusion protein.

[0057] In certain embodiments, the protein recruited to the DNA break is an RNA-guided nuclease such as a Cas nuclease (eg, Cas9 or Cpf1 nuclease) or an engineered RNA-guided FokI-nuclease.

[0058] In certain embodiments, the DNA break is a single-strand or double-strand DNA break. If the DNA break is a single-strand DNA break, the fusion protein includes a protein that selectively binds to the single-strand DNA break. If the DNA break is a double-strand DNA break, the fusion protein includes a protein that selectively binds to the double-strand DNA break.

[0059] In certain embodiments, the donor polynucleotide is single-stranded or double-stranded.

[0060] In another embodiment, the nucleic acid binding domain is an RNA binding domain and the binding site comprises an RNA sequence recognized by the RNA binding domain.

[0061] In another embodiment, the nucleic acid binding domain of the donor recruitment protein is a DNA binding domain and the binding site comprises a DNA sequence recognized by the DNA binding domain. In one embodiment, the DNA binding domain is a LexA DNA binding domain and the binding site is a LexA binding site. In another embodiment, the DNA binding domain is a forkhead homolog 1 (FKH1) DNA binding domain and the binding site is a FKH1 binding site.

[0062] In another embodiment, the polypeptide comprising a nucleic acid binding domain (donor recruitment protein) further comprises a forkhead-associated (FHA) phosphothreonine binding domain, wherein the donor polynucleotide is selectively recruited to DNA breaks with a protein containing phosphorylated threonine residues located sufficiently close to the DNA break for the FHA phosphothreonine binding domain to bind to the phosphorylated threonine residues. In another embodiment, the polypeptide comprising a nucleic acid binding domain comprises a LexA DNA binding domain linked to a FHA phosphothreonine binding domain.

[0063] In another embodiment, the donor polynucleotide is provided by a recombinant polynucleotide comprising a promoter operably linked to the donor polynucleotide. In another embodiment, the fusion protein is provided by a recombinant polynucleotide comprising a promoter operably linked to a fusion protein encoding polynucleotide. In certain embodiments, the donor polynucleotide and fusion protein are provided by a single vector or separate vectors. In another embodiment, at least one vector is a viral vector or a plasmid.

[0064] In certain embodiments, the donor polynucleotide is RNA or DNA. In another embodiment, the method further comprises reverse transcribing the RNA-containing donor polynucleotide with a reverse transcriptase to generate a DNA-containing donor polynucleotide.

[0065] In certain embodiments, the DNA breaks are produced by site-specific nucleases, such as, but not limited to, Cas nucleases (such as Cas9 or Cpf1), engineered RNA-guided FokI-nucleases, meganucleases, zinc finger nucleases (ZFNs), and transcription activator-like effector nucleases (TALENs).

[0066] These and other embodiments of the present disclosure will be readily apparent to those skilled in the art in light of the disclosure herein.

[0067] BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figures 1A-1C Showing the dual CRISPR / Cas9 editing and barcoding system. Figure 1A Shown is a guide RNA (gRNA)-donor DNA sequence cloned into a high-copy vector, with the tRNA-HDV ribozyme promoter driving gRNA expression. The guide-donor plasmid is then transformed into cells with Cas9 pre-expressed from a strong constitutive promoter. Figure 1B Target locus editing is shown. Cas9-gRNA-induced dsDNA breaks are resolved by donor DNA-mediated HR, NHEJ, or cell death. Figure 1C Barcoding of the REDI locus is shown. Induction of Scel with galactose enabled replacement of the reverse-selectable FCY1 with the guide RNA-donor DNA segment, allowing (1) PCR-based phenotyping of competitive growth pools and (2) REDI-based identification of individual variants.

[0069] Figure 2A-2C Demonstrated proof-of-concept of high-efficiency Cas9 editing and SceI-mediated barcoding. Figure 2A Shown are gRNAs targeting ADE2 cloned into the high-copy vector backbone shown in Figure 1, without (left column) or with (right column) donor DNA. The gRNA vectors were transformed into cells with pre-expressed Cas9 (top row) or PTEF1-Cas9 encoded on the gRNA vector (not pre-expressed; bottom row). Figure 2B The ADE2 locus in selected clones is shown. Sequencing verified the desired loss-of-function edit. Figure 2C Results are shown for pooled cells from plates that were switched to either galactose or glucose media. Individual clones were isolated and screened for guide-donor cassette integration at the REDI locus, which was verified by Sanger sequencing.

[0070] Figure 3 Selected guide-donor cassette plasmids for editing heterologous ORF (mCherry) are shown. A library of guide-donor oligonucleotides was purchased from Agilent Technologies and cloned into a high copy guide expression vector (Figure 1). Some bacterial clones were sequenced for correct incorporation of the guide-donor insert and subsequently transformed into yeast cells carrying pre-expressed Cas9 and mCherry ORFs. Transformation plates are shown after 2 days of growth.

[0071] Figures 4A-4F Active donor recruitment was shown to make high-frequency donor-mediated repair feasible. Figure 4A Demonstrated plasmid-based system for high-throughput editing without donor recruitment. Figure 4B We showed that random diffusion of donor DNA resulted in inefficient homologous recombination repair, and most transformants carrying effective gRNAs underwent cell death. Figure 4C Shown is an improved dual-plasmid based system with LexA binding sites on the guide-donor plasmid and the LexA DNA binding domain (DBD) fused to the Fkh1 protein for ultra-high efficiency high-throughput editing. Figure 4D We show that dsDNA breaks trigger phosphorylation of threonine residues on endogenous cellular proteins near the break. This leads to the recruitment of Fkh1 via its forkhead-associated (FHA) phosphothreonine binding domain, generating a high local concentration of donor DNA to facilitate the search for homologous DNA during DNA repair. This makes precise editing the dominant outcome as opposed to cell death. Figure 4E The LexA DNA binding domain is shown fused to Cas9, replacing Fkh1. Figure 4F We show that the Cas9-LexA DBD enables pre-recruitment of guide-donor plasmids to gRNA target sites, promoting HDR after DNA cleavage by Cas9.

[0072] Figure 5A and 5B Active donor recruitment was shown to make high-frequency donor-mediated repair feasible. Figure 5A Pre-expressed Cas9 (upper left, e.g. Figure 4A ) or Cas9 and LexA DBD-Fkh1 (upper right, e.g. Figure 4C ) cells were transformed with a plasmid pool carrying 85% of sequence-verified guide-donors targeting null mutations in ADE2 and 15% of the same plasmid carrying a mutated ADE2 guide RNA. Figure 5B Cells pre-expressing Cas9 were transformed with a perfect guide-donor (lower left). Cells pre-expressing a high-copy guide-donor plasmid were transformed with a Cas9 plasmid (lower right).

[0073] Fig. 6A and 6B Shown is a dual editing and barcoding system in combination with recombinase-mediated indexing (REDI). Fig. 6A Steps 1-4 of the process are shown. Step 1: A composite library of plasmids carrying guide RNA (gRNA) and donor DNA sequences is transformed into a recipient strain modified to contain a barcode locus with a reverse selectable marker (FCY1) flanked by sites for the meganuclease SceI. Transformed cells are plated on -HIS to select plasmids containing correct internal cloning events, colonies are merged and grown to mid-log phase in rich medium containing G418 to maintain selection for the guide-donor plasmid. Cells are then transformed with Cas9 / SceI plasmids and plated on -LEU-HIS. Step 2: Chromosomal targets are cut with Cas9-gRNA and repaired by homologous recombination (HR) with donor DNA encoded on the plasmid. Colonies are recovered and grown for several generations in rich medium with galactose to induce SceI. dsDNA breaks at the chromosomal barcode locus promote integration of the guide-HIS3-donor cassette, linearizing the plasmid. Step 3: Select for successful integration of the guide-donor barcode and loss of the plasmid by plating in synthetic medium containing 5-fluoro-cytosine (5-FC). Transformants are arrayed on agar plates at a density of 1536 to allow subsequent mating with barcoded strains. At this stage, transformants may contain successful desired editing, unwanted mutations caused by errors derived from oligonucleotide synthesis, or no editing. Step 4: Arrayed strain variants are joined with barcoded strains, which contain LoxP sites, followed by unique positional barcodes specifying colony coordinates on the plate, and the remaining URA3 gene. Cre induction leads to LoxP-mediated reconstruction of separate URA3, which physically connects the guide-donor sequence with the positional barcode for high-throughput paired-end sequencing (HTS) guide-donor-barcode combination (step 5). Two different P5 primers allow the guide and donor sequences to be connected to specific colony complexes through a common positional barcode. Figure 6B Diagram showing the Mat a variant chromosome and the Mat α barcoded chromosome and the results of steps 4 and 5.

[0074] Figures 7A-7C Showing Redi-mediated massively parallel strain validation. Fig. 7A Clones isolated from multiplex precision editing experiments are shown to contain successfully edited target loci (dark grey), unwanted mutations in the target loci caused by synthetically derived errors or errors during homologous recombination (HR) (light grey), or unsuccessful edits caused by inefficient guide RNAs (light grey). Figure 7BIndependent clones for each design variant are shown (as shown in light grey, dark grey and medium grey). Step 1: These replicons are rearranged into arrays on different plates so that each plate contains the mutation targeted within a specified chromosome window or gene, and there is only one colony per plate for each design variant. The colonies are merged and genomic DNA is extracted for each plate. Step 2: PCR and deep amplicon sequencing of the target chromosome locus are performed. Successfully edited variants are expected to exist at a frequency of 1 / 1536. Figure 7C Describing the rearrangement of desired clones for either pooled (top) or spatially separated phenotypic assays.

[0075] Figure 8 Shown is a potential workflow for editing, barcoding, validating, and phenotyping strains.

[0076] Fig. 9 Library cloning strategy showing minimization of non-functional vector background. Step 1: Oligonucleotide library is amplified with primers containing 5′-extensions to facilitate Gibson- or sticky-end mediated cloning into the vector backbone. Step 2: Amplified oligonucleotides contain internal type IIS restriction sites. Cloning vectors are treated with type IIS enzymes and phosphatases. This enables scarless internal cloning of structural guide RNA components, Pol III terminators and selectable markers such as HIS3. Step 3: Constant inserts are treated with BspQI only to maintain 5′-phosphates. Step 4: Inserts are ligated into the vector backbone and can subsequently be transformed into recipient yeast for selection on -HIS medium.

[0077] Fig. 10A and 10B The synonymous codon spreading strategy is shown, enabling amino acid mutations outside the guide recognition region. Saturation mutagenesis of the open reading frame is enabled by engineering synonymous codon mutations (dark grey) between non-synonymous variants (light grey) and the protospacer adjacent motif (PAM, box, NGG in this depiction). A sham-WT control was created by incorporating only synonymous variants (dark grey). Donor DNA and guide RNA sequences are also shown to engineer amino acid mutations within the guide recognition sequence ( Fig. 10A ) or outside ( Fig. 10B ) is a non-synonymous variant of .

[0078] Fig.11 Repair is shown using an integrated editing cassette directly into the genome.

[0079] Fig.12Library clones showing guide-donors linked to unique DNA barcodes. (1) Oligonucleotides encoding guide-donors are synthesized in high-density array mode and excised from the array surface to generate a composite pool. (2) Each oligonucleotide contains a common amplification sequence flanking the guide-donor cassette to enable amplification of specific sub-pools. The forward primer carries a restriction site (AscI) at its 3′-end and the reverse primer encodes a unique restriction site (NotI) at its 5′-end, followed by a degenerate barcode (bc) encoding a pseudorandom sequence (NNNVHTGNNNVHTGNNNVHTGNNNVHTGNNN or NNNTGVHNNNTGVHNNNT GVHNNNTGVHNNNTGVHNNN). The degenerate barcode is flanked by 50 bp downstream homology sequence (DH). NotI and AscI sites enable sticky-end cloning into multi-copy recipient vectors, and the AscI site is at the 3′-end of the guide RNA promoter. The guide and donor sequences are separated by a type IIS restriction site (BspQI), which enables cloning with arbitrary overhangs and, in the case of a GTTT located directly 3' of the guide sequence, enables cloning within the guide RNA constant structure element.

[0080] (3) High throughput sequencing (HTS) of the first step cloning products enables ligation of guide-BspQI-donor sequences with unique barcodes (bc). Paired end sequencing can be used to increase base call confidence after merging read 1 and read 2 based on quality. (4) (a) Structural guide RNA components as well as yeast-specific (e.g., URA3) and bacteria-specific (e.g., kanR) selection markers are amplified using primers carrying BspQI sequences at their 5′-ends. The reverse primer includes an additional barcode (bc*; NNNNNN or NNNNNNHVVNHBBHBHD) located 3′ of the Illumina read 2 priming sequence and is modified to include a G to A SNP at the first position of the BspQI site. (b) The first step cloning products are cleaved with BspQI and then treated with phosphatase to allow for scarless cloning of structural gRNA inserts. These second step libraries are selected with kanamycin to allow enrichment of vectors carrying the inserts. Paired-end HTS of bc*-donor and bc allowed barcode mapping to unique guide-donor combinations.

[0081] Fig.13Simultaneous editing and barcode integration via a self-destructive plasmid is shown. (1) The guide-donor vector after second-step cloning is transformed into yeast and selected with an insert-specific marker (URA3). The recipient strain is modified to carry a barcode integration locus with an inverse selectable marker (FCY1). In addition to the guide sequence from the library, the guide-donor plasmid also carries a Guide X expression unit to promote barcode integration, such as the Guide X cleavage site flanking FCY1. After transformation, the guide-donor plasmid accumulates to a high copy number by outgrowth. The presence of a Guide X cleavage site at the 5′-end of the downstream homology (DH) sequence of the guide-donor plasmid enables later linearization of the plasmid to accelerate plasmid loss after editing. (2) Induction of Cas9 results in Guide X cleavage of the plasmid and the genomic barcode locus, as well as cleavage of the library-derived guide elsewhere in the genome. (a) Guide X cleavage results in integration of the complete guide RNA-bc*-donor DNA-bc cassette into the genome via upstream homology (UH) sequences present in the guide-donor plasmid and chromosomal barcode loci. (b) Editing-directed guide cleavage is followed by donor DNA-mediated homologous recombination to produce the desired genome edit.

[0082] Fig.14A and 14B The Cpf1 guide-donor system was shown to produce high efficiency (>99%) editing, with editing enhanced ∼10-fold with Cpf1, whose extent of donor recruitment was similar to that of Cas9. Fig.14A : A Cpf1 guide-donor plasmid targeting the ADE2 gene (the guide has a Cpf1 scaffold) was transformed into cells pre-expressing Cpf1. The donor DNA encodes a deletion that results in a frameshift. Fig. 14B : Cpf1 guide-donor was mixed with non-editing plasmid at a ratio of 17:3 and transformed into cells expressing Cpf1 without (left) or with (right) LexA-FHA. The ratio of red:white colonies is shown on the y-axis.

[0083] Fig.15A modified version of the multiplex genome editing system is shown, in which Cpf1 and / or Cas9 or other RNA-guided nucleases (RGNs) or site-specific nucleases (such as SceI, other meganucleases, ZFNs or TALENs) are expressed from the REDI locus, optionally together with other multiplex editing components, such as TetR and LexA-FHA and markers for positive and negative selection (URA3 and hphMX). In this arragement, the self-destructive guide-donor vector is integrated into the REDI barcode locus, while removing Cas9, Cpf1, and all genes in between. Genetic removal of Cas9 at the DNA level, followed by sufficient natural evolution to dilute Cas9 mRNA and protein, ensures that subsequent fitness analysis is not confounded by the effects of Cas9::guide editing binding to chromatin. The editing guide can be paired with Cas9 or Cpf1, and similarly, the barcode guide X can be paired with Cas9 or Cpf1. The advantage of having nucleases specific to RGNs is that there is no competition between the editing and barcoding guides for the associated RGNs. This arrangement also increases the flexibility of the multiplex system by allowing targeting of more genomic regions through the use of RNA-guided nucleases with different PAM requirements.

[0084] Fig.16 Plasmid spike-in experiments are shown to demonstrate that LexA-FHA and linearized vectors improve HDR efficiency and editing survival. It should be noted that the circular plasmid LexA-FHA produced the overall highest transformation survival rate.

[0085] Fig.17 Shown are HDR efficiencies at LexA sites with or without the presence of the donor recruitment protein dn53BP1-LexA. Targeting 2 independent genes (CACNA1D (CAC) and PPP1R12C (PPP)). The first panel shows the NHEJ ratio at the cleavage site. The second panel shows the total HDR percentage at the cleavage site, and the third panel shows the ratio of HDR to NHEJ in the cell. DETAILED DESCRIPTION OF THE INVENTION

[0087] Unless otherwise indicated, the practice of the present disclosure will employ conventional methods of genome editing, biochemistry, chemistry, immunology, molecular biology and recombinant DNA technology within the skill of the art. Such techniques are fully explained in the literature. See, for example, Targeted Genome Editing Using Site-Specific Nucleases: ZFNs, TALENs, and the CRISPR / Cas9 System (T. Yamamoto ed., Springer, 2015); Genome Editing: The Next Stepin Gene Therapy (Advances in Experimental Medicine and Biology, T. Cathomen, M. Hirsch, and M. Porteus eds., Springer, 2016); Aachen Press Genome Editing (CreateSpace Independent Publishing Platform, 2015); Handbook of ExperimentalImmunology, Vols.I-IV (DMWeir and CCBlackwell eds., Blackwell Scientific Publications); ALLehninger, Biochemistry (Worth Publishers, Inc., currentaddition); Sambrook, et al., Molecular Cloning: A Laboratory Manual (3rd Edition, 2001); Methods In Enzymology(S.Colowick and N. Kaplan eds., Academic Press, Inc.).

[0088] All publications, patents, and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety.

[0089] I. Definitions

[0090] In describing the present disclosure, the following terminology is employed and is intended to be defined as follows.

[0091] It must be noted that, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the content clearly dictates otherwise. Thus, for example, reference to "a cell" includes a mixture of two or more cells, and so forth.

[0092] The term "about", particularly when referring to a given amount, is intended to encompass a deviation of plus or minus 5 percent.

[0093] "Barcode" refers to one or more nucleotide sequences used to identify nucleic acids or cells associated with the barcode. The barcode length can be 3-1000 or more nucleotides, preferably 10-250 nucleotides in length, more preferably 10-30 nucleotides in length, including any length within these ranges, such as 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1000 nucleotides in length. For example, a barcode can be used to identify a single cell, a subpopulation of cells, a colony, or a sample from which a nucleic acid originates. Barcodes can also be used to identify the location of cells, colonies or samples from which nucleic acids originate (i.e., positional barcodes), such as colony locations in cell arrays, well locations in multiwell plates, or test tubes, flasks or other container locations in racks. In particular, barcodes can be used to identify genetically modified cells from which nucleic acids originate. In some embodiments, barcodes are used to identify specific types of genome editing. For example, the guide RNA-donor polynucleotide box itself can be used as a barcode to identify genetically modified cells from which nucleic acids originate. Alternatively, a unique barcode can be used to identify each guide RNA-donor polynucleotide box used for multivariate genome editing. In addition, multiple barcodes can be used in conjunction to identify different nucleic acid features. For example, positional barcode compilation (such as for identifying cells, colonies, cultures or sample locations in arrays, multiwell plates or racks) can be combined with barcodes for identifying guide RNA-donor polynucleotides used for genome editing. In some embodiments, barcodes are inserted into nucleic acids (such as in "barcode loci") in each round of genome editing to identify guide RNA and / or donor polynucleotides used for genetic modification of cells.

[0094] The term "barcoded cell" refers to a cell comprising a nucleic acid comprising a barcode sequence. In one embodiment, the barcode identifies the location of a colony comprising the barcoded cell.

[0095] The terms "polypeptide" and "protein" refer to polymers of amino acid residues and are not limited to a minimum length. Thus, the definition includes peptides, oligopeptides, dimers, multimers, etc. The definition encompasses full-length proteins and fragments thereof. The term also includes post-expression modifications of the polypeptide, such as glycosylation, acetylation, phosphorylation, hydroxylation, etc. In addition, for the purposes of this disclosure, "polypeptide" refers to a protein including modifications of the native sequence such as deletions, additions, and substitutions, as long as the protein maintains the desired activity. These modifications may be intentional, such as by site-directed mutagenesis, or may be accidental, such as by mutations in the host producing the protein or errors caused by PCR amplification.

[0096] The term "Cas9" as used herein encompasses Cas9 endonucleases of type II clustered regularly interspaced short palindromic repeats (CRISPR) systems from any species, and also includes biologically active fragments, variants, analogs and derivatives thereof that retain Cas9 endonuclease activity (i.e., catalyzing site-directed DNA cleavage to produce double-strand breaks). Cas9 endonucleases bind and cut DNA at sites including sequences complementary to their binding guide RNA (gRNA).

[0097] Cas9 polynucleotide, nucleic acid, oligonucleotide, protein, polypeptide or peptide refers to a molecule derived from any source. The molecule need not be physically derived from an organism, but may be produced synthetically or recombinantly. Cas9 sequences from some bacterial species are well known in the art and listed in the National Center for Biotechnology Information (NCBI) database. See, e.g., NCBI's Cas9 entries from: Streptococcus pyogenes (WP_002989955, WP_038434062, WP_011528583); Campylobacter jejuni (WP_022552435, YP_002344900), Campylobacter coli (WP_060786116); Campylobacter fetus (WP_059434633); Corynebacterium ulcerans (NC_015683, NC_017317); Corynebacterium diphtheriae (NC_015683, NC_017317); diphtheria (NC_016782, NC_016786); Enterococcus faecalis (WP_033919308); Spiroplasma syrphidicola (NC_021284); Prevotella intermedia (NC_017861); Spiroplasma taiwanense (NC_021846); Streptococcus iniae (NC_021314); Belliella baltica (NC_018010); Psychroflexus torquis I (NC_018721); Streptococcus thermophilus (YP_820832), Streptococcus mutans (NC_018721); Streptococcus thermophilus (YP_820832), Streptococcus mutans (NC_018721); Streptococcus syrphidicola (NC_021284); Prevotella intermedia (NC_017861); Spiroplasma taiwanense (NC_021846); Streptococcus iniae (NC_021314); Belliella baltica (NC_018010); Psychroflexus torquis I (NC_018721); Streptococcus thermophilus (YP_820832), Streptococcus mutans (NC_018721); Streptococcus syrphidicola (NC_021284); mutans)(WP_061046374, WP_024786433); Listeria innocua(NP_472073); Listeria monocytogenes(WP_061665472); Legionella pneumophila(WP_062726656); Staphylococcus aureus(WP_001573634);Francisella tularensis (WP_032729892, WP_014548420), Enterococcus faecalis (WP_033919308); Lactobacillus rhamnosus (WP_048482595, WP_032965177); and Neisseria meningitidis (WP_061704949, YP_002342100); all sequences of which (as entered through the filing date of this application) are incorporated herein by reference. Any of these sequences or variants thereof include sequences having at least about 70-100% sequence identity, including any percentage of identity within this range, such as 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99% sequence identity, which sequences or variants thereof can be used for genome editing as described herein, wherein the variant retains biological activity, such as Cas9 site-directed nuclease activity. See also Fonfara et al. (2014) Nucleic Acids Res. 42(4):2577-90; Kapitonov et al. (2015) J. Bacteriol. 198(5):797-807, Shmakov et al. (2015) Mol. Cell. 60(3):385-397 and Chylinski et al. (2014) Nucleic Acids Res. 42(10):6091-6105); for sequence comparisons of Cas9 and discussion of genetic diversity and phylogenetic analysis. ;

[0098] "Derivatives" refer to any suitable modification of a native polypeptide of interest, a fragment of a native polypeptide, or its respective analogs, such as glycosylation, phosphorylation, polymer conjugation (e.g., with polyethylene glycol), or other addition of foreign moieties, as long as the desired biological activity of the native polypeptide is retained. Methods for preparing polypeptide fragments, analogs or derivatives are generally available in the art.

[0099] "Fragment" refers to a molecule consisting of only a portion of the complete full-length sequence and structure. A fragment can include a C-terminal deletion, an N-terminal deletion, and / or an internal deletion of a polypeptide. An active fragment of a particular protein or polypeptide generally includes at least about 5-10 consecutive amino acid residues of the full-length molecule, preferably at least about 15-25 consecutive amino acid residues of the full-length molecule, and most preferably at least about 20-50 or more consecutive amino acid residues of the full-length molecule, or any integer between 5 amino acids and the full-length sequence, as long as the fragment in question retains biological activity, such as Cas9 site-directed endonuclease activity.

[0100] "Substantially pure" generally refers to separation of a substance (compound, polynucleotide, nucleic acid, protein, polypeptide, polypeptide composition) such that the substance comprises the majority of the sample in which it is present. Typically, in a sample, the substantially pure component comprises 50%, preferably 80%-85%, and more preferably 90-95% of the sample. Techniques for purifying polynucleotides and polypeptides of interest are well known in the art and include, for example, ion exchange chromatography, affinity chromatography, and density-based sedimentation.

[0101] When referring to polypeptides, "isolated" means that the specified molecule is separated or discrete from the whole organism (where it is found in nature), or exists in the absence of almost no other biological macromolecules of the same type. When referring to polynucleotides, the term "isolated" refers to a nucleic acid molecule that lacks all or part of the sequence normally associated with it in nature; or a sequence as it exists in nature but with heterologous sequences associated with it; or a molecule separated from a chromosome.

[0102] As used herein, the phrase "heterogeneous cell population" refers to a mixture of at least two cell types, one type being a cell of interest (eg, having a genomic modification of interest). A heterogeneous cell population can be derived from any organism.

[0103] As used herein, the terms "isolating" and "isolation" in the context of selecting cells or cell populations having a genomic modification of interest refer to separating the cells or cell populations having the genomic modification of interest from a heterogeneous cell population, such as by positive or negative selection.

[0104] The term "selectable marker" refers to a marker that can be used to identify or enrich a cell population from a heterogeneous cell population, either by positive selection (selection of cells expressing the marker) or negative selection (exclusion of cells expressing the marker).

[0105] As used herein, the terms "polynucleotide," "oligonucleotide," "nucleic acid," and "nucleic acid molecule" include polymeric forms of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. The term refers only to the primary structure of the molecule. Thus, the term includes triple-stranded, double-stranded, and single-stranded DNA, and triple-stranded, double-stranded, and single-stranded RNA. It also includes modifications such as methylation and / or capping, as well as unmodified forms of the polynucleotide. More specifically, the terms "polynucleotide," "oligonucleotide," "nucleic acid," and "nucleic acid molecule" include polydeoxyribonucleotides (containing 2-deoxy-D-ribose), polyribonucleotides (containing D-ribose), any other type of polynucleotide that is an N- or C-glycoside of a purine or pyrimidine base, other polymers containing non-nucleotide backbones, such as polyamides (such as peptide nucleic acids (PNA)) and polymorpholinos (commercially available from Anti-Virals, Inc., Corvallis, Oreg., as Neugene) polymers, and other synthetic sequence-specific nucleic acid polymers, as long as the polymer contains nucleic acid bases in a configuration that allows base pairing and base stacking, as seen in DNA and RNA. There is no deliberate distinction between "polynucleotide," "oligonucleotide," "nucleic acid," and "nucleic acid molecule" and these terms may be used interchangeably. Thus, these terms include, for example, 3'-deoxy-2',5'-DNA, oligodeoxynucleotide N3'P5' phosphoramidates, 2'-O-alkyl-substituted RNA, double-stranded and single-stranded DNA, and double-stranded and single-stranded RNA, microRNA, DNA:RNA hybrids, and hybrids between PNA and NA or RNA, as well as known types of modifications, such as tags, methylation, "capping", the use of analogs (such as 2-aminoadenylic acid, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, C5-propynylcytidine, C5-propynyluridine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-methylcytidine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6-) ... )-methylguanine and 2-thiocytidine) to replace one or more naturally occurring nucleotides, internucleotide modifications, such as those with uncharged bonds (such as methylphosphonate, phosphotriester, phosphoramidate, carbamate, etc.), negatively charged bonds (such as thiophosphate, dithiophosphate, etc.) and positively charged bonds (such as aminoalkylphosphoramidate, aminoalkylphosphotriester), those containing pendant parts, such as proteins (including nucleases, toxins, antibodies, signal peptides, poly-L-lysine, etc.), those with intercalators (such as acridine, psoralen, etc.), those containing chelators (such as metals, radioactive metals, boron, oxidized metals, etc.), those containing alkylating agents, those with modified bonds (such as alpha anomeric nucleic acids, etc.), and unmodified forms of polynucleotides or oligonucleotides. The term also includes locked nucleic acids (e.g., containing ribonucleotides, which have a methylene bridge between the 2'-oxygen atom and the 4'-carbon atom).See, e.g., Kurreck et al. (2002) Nucleic Acids Res. 30:1911-1918; Elayadi et al. (2001) Curr. Opinion Invest. Drugs 2:558-561; Orum et al. (2001) Curr. Opinion Mol. Ther. 3:239-243; Koshkin et al. (1998) Tetrahedron 54:3607-3630; Obika et al. (1998) Tetrahedron Lett. 39:5401-5404.

[0106] The terms "hybridize" and "hybridization" refer to the formation of a complex between nucleotide sequences that are sufficiently complementary to form a complex via Watson-Crick base pairing.

[0107] The term "homologous region" refers to a nucleic acid region with homology to another nucleic acid region. Therefore, whether a "homologous region" exists in a nucleic acid molecule is determined according to another nucleic acid region in the same or different molecule. In addition, since nucleic acid is usually double-stranded, the term "homologous region" used herein refers to the ability of nucleic acid molecules to hybridize with each other. For example, a single-stranded nucleic acid molecule can have 2 homologous regions that can hybridize with each other. Thus, the term "homologous region" includes nucleic acid segments with complementary sequences. The homologous region length is variable, but is usually 4-500 nucleotides (such as about 4-about 40, about 40-about 80, about 80-about 120, about 120-about 160, about 160-about 200, about 200-about 240, about 240-about 280, about 280-about 320, about 320-about 360, about 360-about 400, about 400-about 440, etc.).

[0108] The term "complementary" or "complementarity" as used herein refers to a polynucleotide that can form a base pair with another. Base pairs are usually formed by hydrogen bonds between nucleotide units, which are in an antiparallel direction between polynucleotide chains. Complementary polynucleotide chains can pair bases in a Watson-Crick manner (such as AT, AU, CG) or in any other manner that allows the formation of a duplex. It is known to those skilled in the art that when RNA is used in contrast to DNA, uracil (U) rather than thymine becomes the base considered complementary to adenosine. However, unless otherwise stated, when uracil is mentioned in the context of this disclosure, the ability to replace thymine is implied. "Complementarity" can exist between two RNA chains, between two DNA chains, or between an RNA chain and a DNA chain. It is generally understood that two or more polynucleotides can be "complementary" and can form a duplex, although not perfect or less than 100% complementary. Two sequences are "perfectly complementary" or "100% complementary" if at least one continuous portion of each polynucleotide sequence contains a complementary region, perfectly base-paired with other polynucleotides, without any mismatches or interruptions in such regions. Two or more sequences are considered "perfectly complementary" or "100% complementary" even if one or both of the polynucleotides contain additional non-complementary sequences, as long as the continuous complementary regions within each polynucleotide are able to hybridize perfectly with each other. "Less than perfect" complementarity refers to a situation in which not all consecutive nucleotides in such a complementary region are able to base pair with each other. Determining the percentage of complementarity between two polynucleotide sequences is within the ordinary skill of the art. For Cas9 targeting purposes, the gRNA may include a sequence that is "complementary" to the target sequence (such as a major or minor allele) and is capable of sufficient base pairing to form a duplex (i.e., the gRNA hybridizes with the target sequence). In addition, the gRNA may include a sequence that is complementary to a PAM sequence, wherein the gRNA also hybridizes with the PAM sequence in the target DNA.

[0109] A "target site" or "target sequence" is a nucleic acid sequence that is recognized (i.e., has sufficient complementarity for hybridization) by a guide RNA (gRNA) or donor polynucleotide homology arms. The target site can be allele specific (e.g., a major or minor allele).

[0110] The term "donor polynucleotide" refers to a polynucleotide that provides the sequence for the desired edits to be integrated into the genome at the target locus via HDR.

[0111] "Homologous arm" refers to a portion of a donor polynucleotide, which is responsible for targeting the donor polynucleotide to the genomic sequence to be edited in the cell. The donor polynucleotide generally includes a 5' homology arm hybridized with a 5' genomic target sequence and a 3' homology arm hybridized with a 3' genomic target sequence, and the target sequence is flanked by the nucleotide sequence that the genomic DNA is intended to edit. Homologous arms refer to 5' and 3' (i.e., upstream and downstream) homology arms in this article, which relate to the relative position of homology arms and nucleotide sequences, and the sequence includes the desired editing in the donor polynucleotide. 5' and 3' homology arms hybridize with the target locus of the genomic DNA to be modified, and the regions are respectively referred to as "5' target sequence" and "3' target sequence" in this article. The nucleotide sequence containing the desired editing is integrated into the genomic DNA at the genomic target locus by HDR, and the locus is identified by 5' and 3' homology arms (i.e., there is enough complementarity for hybridization).

[0112] "Administering" a nucleic acid such as a donor polynucleotide, guide RNA, or Cas9 expression system to a cell includes transduction, transfection, electroporation, translocation, fusion, engulfment, shooting, or bombardment, among others, i.e., any means by which the nucleic acid can be transported across the cell membrane.

[0113] "Selective binding" with respect to a guide RNA refers to a guide RNA that preferentially binds to a target sequence of interest or binds to a target sequence with a higher affinity than other genomic sequences. For example, a gRNA will bind to a substantially complementary sequence and not to unrelated sequences. A gRNA that "selectively binds" to a specific allele, such as a specific mutant allele (e.g., an allele containing a substitution, insertion, or deletion) refers to a gRNA that preferentially binds to a specific target allele and, to a lesser extent, to a wild-type allele or other sequences. A gRNA that selectively binds to a specific target DNA sequence will selectively direct an RNA-guided nuclease (e.g., Cas9) to bind to a substantially complementary sequence at the target site and not to unrelated sequences.

[0114] As used herein, the term "recombination target site" refers to a region of a nucleic acid molecule containing a binding site or sequence-specific motif recognized by a site-specific recombinase, which binds to and catalyzes the recombination of a specific DNA sequence at the target site. Site-specific recombinases catalyze recombination between two such target sites. The relative orientation of the target sites determines the outcome of the recombination. For example, if the recombination target sites are on separate DNA molecules, translocation occurs.

[0115] As used herein, the terms "label" and "detectable label" refer to detectable molecules, including but not limited to radioisotopes, fluorescent agents, chemiluminescent agents, chromophores, enzymes, enzyme substrates, enzyme cofactors, enzyme inhibitors, semiconductor nanoparticles, dyes, metal ions, metal sols, ligands (such as biotin, streptavidin or haptens), etc. The term "fluorescent agent" refers to a substance or a portion thereof that can exhibit fluorescence in a detectable range. Examples of specific labels that can be used to implement the present disclosure include, but are not limited to, SYBR Green, SYBR Gold, CAL Fluor dyes (such as CAL Fluor Gold 540, CAL Fluor Orange 560, CAL Fluor Red 590, CAL Fluor Red 610, and CAL Fluor Red 635), Quasar dyes (such as Quasar 570, Quasar 670, and Quasar 705), Alexa Fluor (such as Alexa Fluor 350, Alexa Fluor 488, Alexa Fluor 546, Alexa Fluor 555, Alexa Fluor 594, Alexa Fluor 647, and Alexa Fluor 784), cyanine dyes (such as Cy 3, Cy3.5, Cy5, Cy5.5 and Cy7), fluorescein, 2', 4', 5', 7'-tetrachloro-4-7-dichlorofluorescein (TET), carboxyfluorescein (FAM), 6-carboxy-4', 5'-dichloro-2', 7'-dimethoxyfluorescein (JOE), hexachlorofluorescein (HEX), rhodamine, carboxy-X-rhodamine (ROX), tetramethylrhodamine (TAMRA), FITC, dansyl, umbelliferone, dimethyl acridinium ester (DMAE), Texas Red, luminol, NADPH, horseradish peroxidase (HRP) and α-β-galactosidase.

[0116] "Homology" refers to the percentage of identity between two polynucleotides or two polypeptide portions. Over a defined molecular length, when the sequences show at least about 50% sequence identity, preferably at least about 75% sequence identity, more preferably at least about 80% 85% sequence identity, more preferably at least about 90% sequence identity, and most preferably at least about 95% 98% sequence identity, then the two nucleic acid sequences or two polypeptide sequences are "substantially homologous" to each other. As used herein, substantially homologous also refers to sequences showing complete identity to a specified sequence.

[0117] Generally, "consistency" refers to the exact nucleotide-to-nucleotide or amino acid-to-amino acid correspondence of two polynucleotides or polypeptide sequences, respectively. The percent consistency can be determined by directly comparing the sequence information between the two molecules, aligning the sequences, counting the number of exact matches between the two aligned sequences, dividing by the shorter sequence length, and multiplying the result by 100. Existing computer programs can be used to assist in the analysis, such as ALIGN, Dayhoff, MO, included in Atlas of Protein Sequence and Structure MO Dayhoff ed., 5 Suppl. 3: 353 358, National biomedical Research Foundation, Washington, DC, which adjusts the local homology algorithm of Smith and Waterman Advances in Appl. Math. 2: 482-489, 1981 for peptide analysis. Programs for determining nucleotide sequence identity are available from the Wisconsin Sequence Analysis Package, Version 8 (available from Genetics Computer Group, Madison, WI), e.g., BESTFIT, FASTA, and GAP programs, which also rely on the Smith and Waterman algorithm. These programs are easy to use, using the default parameters recommended by the manufacturer and described in the Wisconsin Sequence Analysis Package above. For example, the percent identity of a particular nucleotide sequence to a reference sequence can be determined using the Smith and Waterman homology algorithm, using a default scoring table and gap penalties for 6 nucleotide positions.

[0118] In the context of the present disclosure, another method for establishing sequence identity is to use the MPSRCH package of programs copyrighted by the University of Edinburgh, developed by John F. Collins and Shane S. Sturrok, and distributed by IntelliGenetics, Inc. (Mountain View, CA). According to this software package, the Smith Waterman algorithm can be used, wherein default parameters are used for the scoring table (e.g., a gap penalty of 12, a gap extension penalty of 1, and a gap of 6). The "match" value generated according to the data reflects "sequence identity". Other suitable programs for calculating the percentage of identity or similarity between sequences are known in the art, for example, another alignment program is BLAST, using default parameters. For example, BLASTN and BLASTP can use the following default parameters: genetic code = standard; filter = none; chain = both; cutoff = 60; expectation = 10; matrix = BLOSUM62; description = 50 sequences; sorting mode = high score; database = non-redundant, GenBank + EMBL + DDBJ + PDB + GenBank CDS translation + Swiss protein + Spupdate + PIR. Details of these procedures are readily available.

[0119] Alternatively, homology can be determined by hybridizing the polynucleotides under conditions that form stable duplexes between homologous regions, followed by digestion with a single-stranded specific nuclease and size determination of the digested fragments. Substantially homologous DNA sequences can be identified in Southern hybridization experiments, for example, under stringent conditions defined for a particular system. Defining appropriate hybridization conditions is within the skill of the art. See, for example, Sambrook et al., supra; DNA Cloning, supra; Nucleic Acid Hybridization, supra.

[0120] As used herein, "recombinant" to describe a nucleic acid molecule refers to a polynucleotide of genomic, cDNA, viral, semisynthetic or synthetic origin that is unrelated, by its origin or manipulation, to all or part of the polynucleotide with which it is naturally associated. The term "recombinant" used in reference to a protein or polypeptide refers to a polypeptide produced by expression of a recombinant polynucleotide. Typically, a gene of interest is cloned and then expressed in a transformed organism, as further described below. The host organism expresses the foreign gene to produce the protein under expression conditions.

[0121] The term "transformation" refers to the insertion of an exogenous polynucleotide into a host cell, regardless of the method used for insertion. For example, it includes direct uptake, transduction, or f-conjugation. The exogenous polynucleotide can be maintained as a non-integrating vector, such as a plasmid, or can be integrated into the host genome.

[0122] "Recombinant host cell," "host cell," "cell line," "cell culture," and other such terms refer to microbial or higher eukaryotic cell lines cultured as unicellular entities, refer to cells that can or have been used as recipients for recombinant vectors or other transfer DNA, and include the primary progeny of the original cell that has been transfected.

[0123] A "coding sequence" or a sequence "encoding" a selected polypeptide is a nucleic acid molecule that is transcribed (in the case of DNA) and translated (in the case of mRNA) into a polypeptide in vivo when placed under the control of appropriate regulatory sequences (or "control elements"). The boundaries of the coding sequence can be determined by a start codon at the 5' (amino) terminus and a translation stop codon at the 3' (carboxyl) terminus. Coding sequences can include, but are not limited to, cDNA from viral, prokaryotic or eukaryotic mRNA, genomic DNA sequences from viral or prokaryotic DNA, and even synthetic DNA sequences. A transcription termination sequence may be located 3' to the coding sequence.

[0124] Typical "control elements" include, but are not limited to, transcriptional promoters, transcriptional enhancer elements, transcriptional termination signals, polyadenylation sequences (located 3' to the translation stop codon), sequences for optimizing translation initiation (located 5' to the coding sequence), and translation termination sequences.

[0125] "Operably linked" refers to an arrangement of elements in which the described components are configured to perform their usual functions. Thus, a given promoter operably linked to a coding sequence is capable of affecting expression of the coding sequence in the presence of appropriate enzymes. A promoter need not be contiguous with a coding sequence, as long as it serves to direct its expression. Thus, for example, an inserted untranslated but transcribed sequence can be present between a promoter sequence and a coding sequence, and the promoter sequence can still be considered "operably linked" to a coding sequence.

[0126] An "expression cassette" or "expression construct" refers to a collection that can direct the expression of a sequence or gene of interest. An expression cassette generally includes the above-mentioned control elements, such as a promoter that is operably linked (for directing transcription) to a sequence or gene of interest, and generally also includes a polyadenylation sequence. In certain embodiments of the present disclosure, the expression cassette described herein may be contained in a plasmid or viral vector construct (such as a vector containing a genome editing cassette for genome modification, wherein the editing cassette contains a promoter that is operably linked to a polynucleotide encoding a guide RNA and a donor polynucleotide). In addition to the components of the expression cassette, the construct may also include one or more selective markers, a signal that allows the construct to exist as a single-stranded DNA (such as an M13 origin of replication), at least one multiple cloning site, and a "mammalian" origin of replication (such as an SV40 or adenovirus origin of replication).

[0127] "Purified polynucleotide" refers to a polynucleotide of interest or a fragment thereof that is substantially free of proteins naturally associated with the polynucleotide, e.g., contains less than about 50%, preferably less than about 70%, and more preferably less than about 90% of the protein. Techniques for purifying polynucleotides of interest are well known in the art and include, for example, disruption of cells containing the polynucleotide with a chaotropic agent and separation of polynucleotides from proteins by ion exchange chromatography, affinity chromatography, and density sedimentation.

[0128] The term "transfection" is used to refer to the uptake of foreign DNA into a cell. A cell is transfected when foreign DNA is introduced into the cell membrane. Several transfection techniques are generally known in the art. See, for example, Graham et al. (1973) Virology, 52:456, Sambrook et al. (2001) Molecular Cloning, a laboratory manual, 3rd edition, Cold Spring Harbor Laboratories, New York, Davis et al. (1995) Basic Methods in Molecular Biology, 2nd edition, McGraw-Hill, and Chu et al. (1981) Gene 13:197. These techniques can be used to introduce one or more foreign DNA moieties into a suitable host cell. The term refers to stable and transient uptake of genetic material, including uptake of peptide- or antibody-linked DNA.

[0129] "Vectors" are capable of transferring nucleic acid sequences to target cells (e.g., viral vectors, non-viral vectors, microparticle vectors, and liposomes). In general, "vector constructs," "expression vectors," and "gene transfer vectors" refer to any nucleic acid construct that is capable of directing the expression of a nucleic acid of interest and capable of transferring a nucleic acid sequence to a target cell. Thus, the term includes cloning and expression vehicles as well as plasmids and viral vectors.

[0130] The terms "variant", "analog" and "mutant protein" refer to biologically active derivatives of a reference molecule that retain the desired activity, such as directed Cas9 endonuclease activity. In general, the terms "variant" and "analog" refer to compounds having a native polypeptide sequence and structure, with one or more amino acid additions, substitutions (generally conservative in nature) and / or deletions relative to the native molecule, as long as the modification does not destroy the biological activity and is "substantially homologous" to the reference molecule as defined below. Generally, the amino acid sequence of such an analog will have a high degree of sequence homology with the reference sequence, such as when the two sequences are aligned, the amino acid sequence homology is greater than 50%, generally greater than 60%-70%, and even more specifically 80%-85% or higher, such as at least 90%-95% or higher. Typically, analogs will include the same number of amino acids, but will also include substitutions, as explained herein. The term "mutant protein" also includes polypeptides having one or more amino acid-like molecules (including but not limited to compounds containing only amino and / or imino molecules), polypeptides containing one or more amino acid analogs (including, for example, non-natural amino acids, etc.), polypeptides with substitution linkages and other modifications known in the art (naturally occurring and non-naturally occurring (such as synthetic), cyclization, branched molecules, etc.). The term also includes molecules containing one or more N-substituted glycine residues ("peptoids") and other synthetic amino acids or peptides. (For descriptions of peptoids, see, for example, U.S. Pat. Nos. 5,831,005; 5,877,278; and 5,977,301; Nguyen et al., Chem. Biol. (2000) 7:463-473; and Simon et al., Proc. Natl. Acad. Sci. USA (1992) 89:9367–9371). Methods for preparing polypeptide analogs and mutant proteins are known in the art and are further described below.

[0131] As explained above, analogs generally include substitutions that are conservative in nature, i.e., substitutions that occur within a family of amino acids whose side chains are related. In particular, amino acids are generally divided into four families: (1) acidic - aspartic acid and glutamic acid; (2) basic - lysine, arginine, histidine; (3) nonpolar - alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan; and (4) uncharged polar - glycine, asparagine, glutamine, cysteine, serine, threonine, and tyrosine. Phenylalanine, tryptophan, and tyrosine are sometimes classified as aromatic amino acids. For example, it is reasonable to predict that substitutions of leucine with isoleucine or valine, aspartic acid with glutamic acid, threonine with serine, or similar conservative substitutions of amino acids with structurally related amino acids alone will not have a significant effect on biological activity. For example, the polypeptide of interest may include up to about 5-10 conservative or non-conservative amino acid substitutions, or even up to about 15-25 conservative or non-conservative amino acid substitutions, or any integer between 5-25, as long as the desired molecular function remains intact. Those skilled in the art can easily determine the regions of the molecule of interest that are tolerant to change by reference to the Hopp / Woods and Kyte-Doolittle diagrams well known in the art.

[0132] "Gene transfer" or "gene delivery" refers to a method or system for reliably inserting a DNA or RNA of interest into a host cell. Such methods can result in transient expression of non-integrated transferred DNA, extrachromosomal replication and expression of transferred replicons (such as episomes), or integration of transferred genetic material into host cell genomic DNA. Gene delivery expression vectors include, but are not limited to, vectors derived from bacterial plasmid vectors, viral vectors, non-viral vectors, adenoviruses, retroviruses, alphaviruses, poxviruses, and vaccinia viruses.

[0133] The term "derived from" is used herein to identify the original source of a molecule, but is not intended to limit the method by which the molecule is prepared, which may be, for example, chemically synthesized or recombinantly prepared.

[0134] A polynucleotide "derived from" a specified sequence refers to a polynucleotide sequence comprising a contiguous sequence of approximately at least about 6 nucleotides, preferably at least about 8 nucleotides, more preferably at least about 10-12 nucleotides, and even more preferably at least about 15-20 nucleotides, corresponding to a region of a specified nucleotide sequence, i.e., identical or complementary. A derived polynucleotide is not necessarily physically derived from a nucleotide sequence of interest, but may be produced in any manner, including but not limited to chemical synthesis, replication, reverse transcription, or transcription, based on information provided by the base sequence in the region from which the polynucleotide is obtained. Thus, it may represent a sense or antisense orientation of the original polynucleotide.

[0135] The term "subject" includes vertebrates and invertebrates, including, but not limited to, mammals, including humans and non-human mammals such as non-human primates, including chimpanzees and other apes and monkey species; laboratory animals such as mice, rats, rabbits, hamsters, guinea pigs, and chinchillas; livestock such as dogs and cats; farm animals such as sheep, goats, pigs, horses, and cattle; and birds such as poultry, wild birds, and game birds, including chickens, turkeys and other quail, ducks, geese, etc. In some cases, the methods of the present disclosure find use in experimental animals, veterinary applications, and in the development of animal models for disease, including, but not limited to, rodents, including mice, rats, and hamsters; primates, and transgenic animals.

[0136] II. Modes for carrying out the invention

[0137] Before describing the present disclosure in detail, it is to be understood that the present disclosure is not limited to specific formulations or process parameters, as these may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments of the present disclosure only and is not intended to be limiting.

[0138] Although several methods and materials similar or equivalent to those described herein can be used to practice the present disclosure, the preferred materials and methods are described herein.

[0139] The present disclosure is based on the development of a method for producing genetic engineering clones in large-scale parallel using RNA-guided nucleases and genomic barcodes. Specifically, high-throughput multiplex genome editing is achieved by promoting precise genome editing in the desired target chromosome locus through homology-mediated repair. The integration of guide RNA and donor DNA sequences as genomic barcodes in a single chromosome locus allows identification, separation and large-scale parallel verification of individual variants from a transformant library. Strains can be arrayed according to their precise genetic modification, as specified by the incorporation of donor DNA into heterologous or natural genes. The inventors have demonstrated that their system provides high editing efficiency in yeast cells and enables simultaneous editing at more than one genomic position (Example 1). The inventors have also developed a codon editing method outside the typical guide RNA recognition region (making complete saturation mutations of protein-coding genes feasible) and a marker-based internal cloning method (removing the background caused by oligonucleotide synthesis errors and incomplete vector backbone cutting). Additionally, homology-directed repair (HDR) in metazoan cells can be enhanced by using CRISPR-interference (CRISPRi), RNA interference (RNAi), or chemical-based inhibition of nonhomologous end joining (NHEJ) and active donor recruitment. The collection of genome-modified strains generated by the methods described herein can be arranged according to their precise genetic modification, as specified by the incorporation of barcoded donor DNA into heterologous or native genes.

[0140] To further understand the present disclosure, a more detailed discussion is provided below regarding multiplexed genome editing, with barcoding and strain validation by these methods.

[0141] A. Multiplex Genome Editing

[0142] As explained above, the methods of the present disclosure provide multiplex genome editing, and guide RNA-donor DNA expression cassette barcoding for cell genome modification is compiled to facilitate verification of individual variants from a transformant library. Multiplex editing is achieved as follows: cells are transfected with a variety of recombinant polynucleotides, each of which contains a genome editing cassette comprising a guide RNA encoding polynucleotide and a donor polynucleotide, the guide RNA being able to hybridize at the genomic target locus to be modified, and the donor polynucleotide comprising a sequence to be edited that is to be integrated into the genomic target locus by homology-mediated repair (HDR). Each genome editing cassette includes a different guide RNA-donor polynucleotide combination, so that a variety of recombinant polynucleotides containing it can generate a variety of different desired edits at one or more genomic target loci. After transfecting the cells with the recombinant polynucleotides, the cells are cultured under conditions suitable for the transcription of the guide RNA from each genome editing cassette. Introducing an RNA-guided nuclease into a cell can form a complex with the transcribed guide RNA, wherein the guide RNA directs the complex to one or more genomic target loci in the cell, where the RNA-guided nuclease generates a double-strand break in the genomic DNA, causing the donor polynucleotide to be integrated into the genomic target locus by HDR to generate a variety of genetically modified cells. In certain embodiments, the method further comprises additional rounds of genetically modified cell genome editing by repeating the steps with different genome editing boxes. The genetically modified cells can be inoculated in an ordered array in a culture medium suitable for their growth to generate cloned arranged colonies.

[0143] A set of genome editing boxes can be designed to produce mutations at multiple sites within a single gene or multiple sites in different genes or anywhere in the genome, including non-coding regions. Such mutations may include insertions, deletions, or substitutions. Each donor polynucleotide comprises a sequence comprising different desired edits to the genome, which can be used to modify a specific target locus in a cell, wherein the donor polynucleotide is integrated into the genome at the target locus by site-directed homologous recombination. For example, the donor polynucleotide can be used to introduce a desired edit into the genome for the purpose of repairing, modifying, replacing, deleting, weakening, or inactivating the target gene.

[0144] In the donor polynucleotide, the sequence flank containing the desired editing is a pair of homology arms, which is responsible for the target locus to be edited in the donor polynucleotide targeting cell. The donor polynucleotide generally includes a 5' homology arm hybridized with the 5' genomic target sequence and a 3' homology arm hybridized with the 3' genomic target sequence. The homology arms are referred to herein as 5' and 3' (i.e., upstream and downstream) homology arms, which relate to the relative position of the homology arms and the nucleotide sequence desired to be edited in the donor polynucleotide. 5' and 3' homology arms hybridize with the target locus of the DNA to be modified, which are respectively referred to herein as "5' target sequence" and "3' target sequence".

[0145] The homology arms must have sufficient complementarity to hybridize to the target sequence, thereby mediating homologous recombination of the donor polynucleotide with the genomic DNA at the target locus. For example, the homology arms can include nucleotide sequences having at least about 80-100% sequence identity to the corresponding genomic target sequence, including any percentage of identity within this range, such as at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity, wherein the nucleotide sequence containing the desired edit is integrated into the genomic DNA by HDR at the genomic target locus recognized by the 5' and 3' homology arms (i.e., having sufficient complementarity for hybridization).

[0146] In certain embodiments, the corresponding homologous nucleotide sequences in the genomic target sequence (ie, the "5' target sequence" and the "3' target sequence") are flanked by specific sites for cutting and / or specific sites for introducing desired edits. The distance between the specific cutting site and the homologous nucleotide sequence (such as each homology arm) can be hundreds of nucleotides. In some embodiments, the distance between the homology arm and the cutting site is 200 nucleotides or less (such as 0, 10, 20, 30, 50, 75, 100, 125, 150, 175 and 200 nucleotides). In most cases, a smaller distance can produce a higher gene targeting rate. In a preferred embodiment, the donor polynucleotide is substantially identical to the target genomic sequence over its entire length, except for the sequence changes to be introduced into the portion of the genome, which covers the specific cutting site and the genomic target sequence to be changed.

[0147] The homology arms can be any length, such as 10 nucleotides or more, 50 nucleotides or more, 100 nucleotides or more, 250 nucleotides or more, 300 nucleotides or more, 350 nucleotides or more, 400 nucleotides or more, 450 nucleotides or more, 500 nucleotides or more, 1000 nucleotides (1 kb) or more, 5000 nucleotides (5 kb) or more, 10000 nucleotides (10 kb) or more, etc. In some cases, the 5' and 3' homology arms are substantially identical in length to each other, such as one homology arm can be 30% shorter or less than the other homology arm, 20% shorter or less than the other homology arm, 10% shorter or less than the other homology arm, 5% shorter or less than the other homology arm, 2% shorter or less than the other homology arm, or only a few nucleotides shorter than the other homology arm. In other cases, the 5' and 3' homology arms are substantially different lengths from each other, such as one homology arm can be 40% or more shorter, 50% or more shorter, sometimes 60% or more shorter, 70% or more shorter, 80% or more shorter, 90% or more shorter, or 95% or more shorter than the other homology arm.

[0148] In certain embodiments, cells containing the modified genome are identified in vitro or in vivo by including a selectable marker expression cassette in the vector. The selectable marker confers an identifiable change to the cell, allowing for positive selection of genetically modified cells that have integrated the donor polynucleotide into the genome. For example, nutritional markers (i.e., genes that confer the ability to grow in nutrient-deficient media) such as cytosine deaminase (Fcy1) in Saccharomyces cerevisiae cerevisiae) on medium containing cytosine as the sole nitrogen source (5-fluorocytosine is toxic to cells producing cytosine deaminase and can be used for counter-selection), imidazole glycerol phosphate dehydratase (HIS3) on medium lacking histidine in Saccharomyces cerevisiae, phosphoribosyl anthranilate isomerase (TRP1) on medium lacking tryptophan (5-fluoroanthranilic acid is toxic to cells producing phosphoribosyl anthranilate isomerase and can be used for counter-selection), guanosine 5'-phosphate decarboxylase (URA3) on medium lacking uracil or uridine (5 -fluoroorotic acid is toxic to cells that produce guanosine 5'-phosphate decarboxylase and can be used for counter-selection); fluorescent or bioluminescent markers (such as mCherry, Dronpa, mOrange, mPlum, Venus, YPet, green fluorescent protein (GFP), enhanced green fluorescent protein (EGFP), phycoerythrin or luciferase); cell surface markers; expression of reporter genes (such as GFP, dsRed, GUS, lacZ, CAT); or drug selection markers such as genes that confer resistance to neomycin, puromycin, hygromycin, DHFR, GPT, bleomycin or histidinol can be used to identify cells. Alternatively, enzymes such as herpes simplex virus thymidine kinase (tk) or chloramphenicol acetyltransferase (CAT) can be used. Any selective marker can be used as long as it can be expressed after integration of the donor polynucleotide by HDR to allow identification of genetically modified cells. More examples of selective markers are well known to those skilled in the art.

[0149] In certain embodiments, the selective marker expression cassette encodes 2 or more selective markers. Selective markers can be used in conjunction with, for example, nutritional markers, or cell surface markers can be used together with fluorescent markers, or drug resistance genes can be used together with suicide genes. In certain embodiments, the donor polynucleotide is provided by a polycistronic vector to allow for the combined expression of multiple selective markers. Polycistronic vectors can include IRES or viral 2A peptides to allow for the expression of more than one selective marker from a single vector, as further described below.

[0150] In diploid cells, genome editing as described herein can cause genetic modification of 1 allele or 2 alleles in genomic DNA. In certain embodiments, at least one selection marker for positive selection is a fluorescent marker, wherein the fluorescence intensity can be measured to determine whether the genetically modified cell includes monoallelic editing or biallelic editing.

[0151] In other embodiments, negative selection markers are used to identify cells without a selection marker expression cassette (i.e., the sequence encoding the positive selection marker is destroyed or deleted). For example, the genome editing cassette is integrated into the vector and can be detected by destroying the selection marker gene. The suicide marker can be included as a negative selection marker to promote negative selection of cells. The suicide gene can be used to select killer cells by inducing apoptosis in genetically modified cells or converting non-toxic drugs into toxic compounds. Examples include suicide genes encoding thymidine kinase, cytosine deaminase, intracellular antibodies, telomerase, caspase and DNase. In certain embodiments, the suicide gene is used in conjunction with one or more other selection markers, such as those described above for positive selection of cells. In addition, the suicide gene can be used for genetically modified cells, such as by allowing it to be arbitrarily destroyed to improve its safety. See, e.g., Jones et al. (2014) Front. Pharmacol. 5:254, Mitsui et al. (2017) Mol. Ther. Methods Clin. Dev. 5:51-58, Greco et al. (2015) Front. Pharmacol. 6:95; incorporated herein by reference.

[0152] Genome editing can be performed on a single cell or cell group of interest, and can be implemented on any cell type, including any cell from a prokaryotic, eukaryotic or archaeal organism, including bacteria, archaea, fungi, protists, plants and animals. Cells from tissues, organs and biopsies, as well as recombinant cells, genetically modified cells, cells from in vitro cultured cell lines and artificial cells (such as nanoparticles, liposomes, polymer vesicles or microcapsules encapsulating nucleic acids) can be used to implement the present disclosure. The disclosed method can also be applied to editing nucleic acids in cell fragments, cell components or organelles containing nucleic acids (such as mitochondria in animal and plant cells, plastids (such as chloroplasts) in plant cells and algae). Before or after the implementation of the genome editing described herein, the cells can be cultured or amplified. In one embodiment, the cell is a yeast cell.

[0153] RNA-guided nucleases can target specific genomic sequences (i.e., genomic target sequences to be modified) by changing their guide RNA sequences. Target-specific guide RNAs include nucleotide sequences complementary to genomic target sequences, thereby regulating the binding of nuclease-gRNA complexes by hybridization at target sites. For example, gRNA can be designed with sequences complementary to minor allele sequences so that nuclease-gRNA complexes target mutation sites. Mutations may include insertions, deletions, or substitutions. For example, mutations may include single nucleotide changes, gene fusions, transpositions, inversions, duplications, frameshifts, missense, nonsense, or other mutations associated with disease phenotypes of interest. Targeting minor alleles may be common genetic variants or rare genetic variants. In certain embodiments, the gRNA is designed to selectively bind to minor alleles distinguished by single base pairs, for example, for allowing nuclease-gRNA complexes to bind to single nucleotide polymorphisms (SNPs). In particular, gRNAs may be designed to target disease-related mutations of interest for genome editing to remove mutations from genes. Alternatively, the gRNA can be designed with a sequence complementary to the sequence of the major or wild-type allele so that the nuclease-gRNA complex targets the allele for genome editing purposes, thereby introducing mutations such as insertions, deletions or substitutions into the gene within the cell's genomic DNA. For example, such genetically modified cells can be used to change phenotypes, confer new properties, or generate disease models for drug screening.

[0154] In certain embodiments, the RNA-guided nuclease for genome editing is a clustered regularly interspaced short palindromic repeats (CRISPR) system Cas nuclease. Any RNA-guided Cas nuclease that can catalyze site-directed DNA cleavage to allow integration of a donor polynucleotide via the HDR mechanism can be used for genome editing, including CRISPR system type I, type II, or type III Cas nucleases. Examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5e (CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9 (Csn1or Csb1, Csb2, Csb3, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, and Cu1966, as well as homologs or modified versions thereof.

[0155] In certain embodiments, a type II CRISPR system Cas9 endonuclease is used. Cas9 nucleases from any species or biologically active fragments, variants, analogs or derivatives thereof retain Cas9 endonuclease activity (i.e., catalyzing DNA site-directed cleavage to produce double-strand breaks) and can be used to perform genome modifications as described herein. Cas9 does not need to be physically derived from an organism, but can be synthesized or recombinantly generated. Cas9 sequences from some bacterial populations are well known in the art and listed in the National Center for Biotechnology Information (NCBI) database. See, for example, NCBI's Cas9 entry, from: Streptococcus pyogenes (WP_002989955, WP_038434062, WP_011528583); Campylobacter jejuni (WP_022552435, YP_002344900), Campylobacter coli (WP_060786116); Campylobacter fetus (WP_059434633); Corynebacterium ulcerans (NC_01568 3. NC_017317); Corynebacterium diphtheriae (NC_016782, NC_016786); Enterococcus faecalis (WP_033919308); Spiroplasma yrphidicola (NC_021284); Prevotella intermedia (NC_017861); Spiroplasma taiwanensis (NC_021846); Streptococcus dolphinii (NC_021314); Belliella baltica (NC_018010); Campylobacter contortus I (NC_018721); Streptococcus thermophilus (YP_820832), Streptococcus mutans (WP_061046374, WP_024786433); Listeria monocytogenes (NP_472073); Listeria monocytogenes (WP_061665472); Legionella pneumophila (WP_062726656); Staphylococcus aureus (WP_ 001573634); Francisella tularensis (WP_032729892, WP_014548420), Enterococcus faecalis (WP_033919308); Lactobacillus rhamnosus (WP_048482595, WP_032965177); and Neisseria meningitidis (WP_061704949, YP_002342100); all sequences of which (as entered through the filing date of this application) are incorporated herein by reference. Any of these sequences or variants thereof include sequences having at least about 70-100% sequence identity, including any percentage of identity within this range, such as 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99% sequence identity, and the sequences or variants thereof can be used for genome editing as described herein.See also Fonfara et al. (2014) Nucleic Acids Res. 42(4):2577-90; Kapitonov et al. (2015) J. Bacteriol. 198(5):797-807, Shmakov et al. (2015) Mol. Cell. 60(3):385-397 and Chylinski et al. (2014) Nucleic Acids Res. 42(10):6091-6105); for sequence comparisons of Cas9 and discussion of genetic diversity and phylogenetic analysis.

[0156] CRISPR-Cas systems occur naturally in bacteria and archaea, where they play a role in RNA-mediated adaptive immunity against foreign DNA. Bacterial type II CRISPR systems use the endonuclease Cas9, which forms a complex with a guide RNA (RNA) that specifically hybridizes to a complementary genomic target sequence, where the Cas9 endonuclease catalyzes cleavage to generate double-strand breaks. Targeting Cas9 also typically relies on the presence of a 5′ protospacer adjacent motif (PAM) in the DNA at or near the gRNA binding site.

[0157] The genomic target site generally includes a nucleotide sequence complementary to the gRNA, and may also include a protospacer sequence adjacent motif (PAM). In certain embodiments, the target site includes 20-30 base pairs in addition to the 3 base pair PAM. Typically, the first nucleotide of the PAM can be any nucleotide, and the other 2 nucleotides depend on the selected specific Cas9 protein. Exemplary PAM sequences are known to those skilled in the art, including but not limited to NNG, NGN, NAG and NGG, where N represents any nucleotide. In certain embodiments, the allele targeted by the gRNA includes a mutation that produces a PAM in the allele, wherein the PAM promotes the binding of the Cas9-gRNA complex to the allele.

[0158] In certain embodiments, the gRNA is 5-50 nucleotides, 10-30 nucleotides, 15-25 nucleotides, 18-22 nucleotides, or 19-21 nucleotides in length, or any length between the indicated ranges, including, for example, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides in length. The guide RNA can be a single guide RNA containing crRNA and tracrRNA sequences in a single RNA molecule, or the guide RNA can include 2 RNA molecules with crRNA and tracrRNA sequences located in different RNA molecules.

[0159] In another embodiment, a CRISPR nuclease (Cpf1) from Prevotella and Francisella 1 can be used. Cpf1 is another type II CRISPR / Cas system RNA-guided nuclease, which has similarities to Cas9 and can be used similarly. Unlike Cas9, Cpf1 does not require tracrRNA and depends only on the crRNA in its RNA, which provides the advantage that a shorter guide RNA than Cas9 can be used with Cpf1 for targeting. Cpf1 can cut DNA or RNA. The PAM site recognized by Cpf1 has the sequence 5'-YTN-3' (where "Y" is pyrimidine and "N" is any nucleobase) or 5'-TTN-3', which is opposite to the G-rich PAM site recognized by Cas9. Cpf1 cutting DNA can generate double-strand breaks with sticky ends, with 4 or 5 nucleotide overhangs. For discussion of Cpf1, see, e.g., Ledford et al. (2015) Nature. 526(7571):17-17, Zetsche et al. (2015) Cell. 163(3):759-771, Murovec et al. (2017) Plant Biotechnol. J. 15(8):917-926, Zhang et al. (2017) Front. Plant Sci. 8:177, Fernandes et al. (2016) Postepy Biochem. 62(3):315-326; incorporated herein by reference.

[0160] C2c1 is another type of RNA-guided nuclease that can be used in the type II CRISPR / Cas system. C2c1 is similar to Cas9 and depends on crRNA and tracrRNA to guide to the target site. C2c1 is described in, for example, Shmakov et al. (2015) Mol Cell. 60 (3): 385-397, Zhang et al. (2017) Front Plant Sci. 8: 177; incorporated herein by reference.

[0161] In another embodiment, an engineered RNA-guided FokI nuclease can be used. The RNA-guided FokI nuclease includes a fusion of an inactivated Cas9 (dCas9) and a FokI endonuclease (FokI-dCas9), wherein the dCas9 portion imparts guide RNA-dependent targeting to FokI. The description of the engineered RNA-guided FokI nuclease is described in, for example, Havlicek et al. (2017) Mol. Ther. 25 (2): 342-355, Pan et al. (2016) Sci Rep. 6: 35794, Tsai et al. (2014) Nat Biotechnol. 32 (6): 569-576; incorporated herein by reference.

[0162] RNA-guided nucleases can be provided in the form of proteins, such as nucleases complexed with gRNA, or provided by nucleic acids encoding RNA-guided nucleases, such as RNA (e.g., messenger RNA) or DNA (expression vectors). Codon usage can be optimized to improve the production of RNA-guided nucleases in specific cells or organisms. For example, nucleic acids encoding RNA-guided nucleases can be modified to replace codons with a higher frequency of use than naturally occurring polynucleotide sequences in yeast cells, bacterial cells, human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, or any other host cell of interest. When nucleic acids encoding RNA-guided nucleases are introduced into cells, the protein can be expressed transiently, conditionally, or constitutively in the cells.

[0163] Donor polynucleotides and gRNAs are readily synthesized by standard techniques, such as solid phase synthesis via phosphoramidite chemistry, as disclosed in U.S. Pat. Nos. 4,458,066 and 4,415,732, incorporated herein by reference; Beaucage et al., Tetrahedron (1992) 48 :2223-2311; and Applied Biosystems User Bulletin No. 13 (April 1, 1987). Other chemical synthesis methods include, for example, Narang et al., Meth. Enzymol. (1979) 68 :90 described phosphate triester method and Brown et al., Meth. Enzymol. (1979) 68: 109 disclosed phosphodiester method. In view of the short length of gRNA (usually about 20 nucleotides in length) and donor polynucleotide (usually about 100-150 nucleotides), gRNA-donor polynucleotide cassettes can be generated by standard oligonucleotide synthesis techniques and then connected to a vector. In addition, gRNA-donor polynucleotide cassette libraries for thousands of genomic targets can be easily formed using highly parallel array-based oligonucleotide library synthesis methods (see, for example, Cleary et al. (2004) Nature Methods 1: 241-248, Svensen et al. (2011) PLoS One 6 (9): e24906).

[0164] In addition, the linker sequence can be added to the oligonucleotide to facilitate high-throughput amplification or sequencing. For example, a pair of linker sequences can be added to the 5' and 3' ends of the oligonucleotide to allow simultaneous amplification or sequencing of multiple oligonucleotides by the same set of primers. In addition, restriction sites can be incorporated into the oligonucleotide to facilitate cloning of the oligonucleotide into a vector. For example, the oligonucleotide containing the gRNA-donor polynucleotide box can be designed with a common 5' restriction site and a common 3' restriction site to facilitate connection to the genome modification vector. Restrictive digestion of each oligonucleotide is selectively implemented at the common 5' restriction site and the common 3' restriction site to generate a restriction fragment that can be cloned into a vector (such as a plasmid or a viral vector), and then the cell is transformed with a vector containing the gRNA-donor polynucleotide box.

[0165] The polynucleotide encoding the gRNA-donor polynucleotide box can be amplified, for example, before being connected to the genome modification vector or before sequencing after barcoding. Any method for amplifying oligonucleotides can be used, including but not limited to polymerase chain reaction (PCR), isothermal amplification, nucleic acid sequence-dependent amplification (NASBA), transcription-mediated amplification (TMA), strand displacement amplification (SDA) and ligase chain reaction (LCR). In one embodiment, the genome editing box includes common 5' and 3' priming sites to allow the gRNA-donor polynucleotide sequence to be amplified in parallel with a set of universal primers. In another embodiment, a set of selective primers is used to selectively amplify a subset of gRNA-donor polynucleotides from the collected mixture.

[0166] The cells transformed with the recombinant polynucleotide containing the genome editing cassette can be prokaryotic cells or eukaryotic cells, and are preferably designed for efficient incorporation of the gRNA-donor polynucleotide library by transformation. Methods for introducing nucleic acids into host cells are well known in the art. Common methods of transformation include chemically induced transformation, usually with divalent cations (such as CaCl2), and electroporation. See, for example, Sambrook et al. (2001) Molecular Cloning, a laboratory manual, 3 rdedition, Cold Spring Harbor Laboratories, New York, Davis et al. (1995) Basic Methods in Molecular Biology, 2 nd edition, McGraw-Hill, and Chu et al. (1981) Gene 13:197; incorporated herein by reference in its entirety.

[0167] Normal random diffusion of donor DNA with DNA breaks is rate-limiting for homologous repair. Active donor recruitment can be used to increase the frequency of genetic modification of cells through HDR. The active donor recruitment method includes: a) introducing a fusion protein into a cell, the protein comprising a protein that selectively binds to DNA breaks connected to a polypeptide containing a nucleic acid binding domain; and b) introducing a donor polynucleotide into the cell, comprising i) a nucleotide sequence that is complementary enough to hybridize with a sequence adjacent to the DNA break, and ii) a nucleotide sequence comprising a binding site recognized by the nucleic acid binding domain of the fusion protein, wherein the nucleic acid binding domain selectively binds to the binding site on the donor polynucleotide to generate a complex between the donor polynucleotide and the fusion protein, thereby recruiting the donor polynucleotide to the DNA break and promoting HDR.

[0168] DNA breaks can be generated by site-specific nucleases, such as, but not limited to, Cas nucleases (such as Cas9, Cpf1 or C2c1), engineered RNA-guided FokI nucleases, zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), restriction endonucleases, meganucleases, homing endonucleases. Any site-specific nuclease that selectively cleaves sequences at the donor polynucleotide target integration site can be used.

[0169] The DNA break can be a single-strand (nick) or double-strand DNA break. If the DNA break is a single-strand DNA break, the fusion protein used includes a protein that selectively binds to single-strand DNA breaks, and if the DNA break is a double-strand DNA break, the fusion protein used includes a protein that selectively binds to double-strand DNA breaks.

[0170] In the fusion, the protein that selectively binds to the DNA break can be, for example, an RNA-guided nuclease, such as a Cas nuclease (eg, Cas9 or Cpf1) or an engineered RNA-guided FokI nuclease.

[0171] The donor polynucleotide can be single-stranded or double-stranded and can be composed of RNA or DNA. A donor polynucleotide containing DNA can be generated from a donor polynucleotide containing RNA, if desired, by reverse transcription using a reverse transcriptase. Depending on the type of nucleic acid binding domain in the fusion protein, the donor polynucleotide can include, for example, corresponding binding sites, including RNA sequences recognized by the RNA binding domain or DNA sequences recognized by the DNA binding domain. For example, a fusion protein can be constructed with a LexA DNA binding domain that is to be matched with a corresponding LexA binding site of the donor polynucleotide. In another example, a fusion protein can be constructed with a FKH1 DNA binding domain that is to be matched with a corresponding FKH1 binding site of the donor polynucleotide.

[0172] The DNA binding domain can be any protein or domain from a protein that binds to a known DNA sequence. Non-limiting examples include LexA, Gal4, zinc finger proteins, TALEs, and transcription factors. Examples of each of these proteins are well known in the art.

[0173] In another embodiment, the fusion protein may further include a FHA phosphothreonine binding domain, wherein the donor polynucleotide is selectively recruited to a DNA break with a protein containing a phosphorylated threonine residue located sufficiently close to the DNA break for the FHA phosphothreonine binding domain to bind to the phosphorylated threonine residue. The FHA phosphothreonine binding domain can be combined with any DNA binding domain (such as a fusion of FHK1-LexA) for donor recruitment.

[0174] Without being bound by theory, it is contemplated that donor recruitment proteins may include polypeptide domains from any protein recruited to DNA breaks, such as double-stranded DNA breaks. Non-limiting examples include proteins that bind to regions of DNA damage and / or DNA repair proteins. Phospho-Ser / Thr-binding domains appear as key regulators of cell cycle progression and DNA damage signaling. Such domains include 14-3-3 proteins, WW domains, Polo box domains (in PLK1), WD40 repeats (including the E3 ligase SCF βTrCPIn some embodiments, the donor recruitment system includes BRCT domains (including those in BRCA1), BRCT domains (including those in BRCA1), and FHA domains (such as in CHK2 and MDC1). These domains all have the potential to be used in donor recruitment systems. The FHA domain is well conserved in bacteria and therefore also has donor recruitment utility in bacteria as well as eukaryotes. Proteins or genes encoding these proteins are provided in Tables 1-5 without limitation. Additional genes / proteins are known in the art and can be found, for example, by searching public gene or protein databases that are known to play a role in DNA repair or DNA damage binding (such as gene ontology term analysis). It is expected that proteins from species (such as eukaryotic proteins, proteins from yeast, mammalian cells, including human proteins, and / or proteins from fungi) can be used. In an embodiment, the donor recruitment protein includes a polypeptide sequence from a DNA breakage recruitment protein from the same kingdom, phylum, class, order, family, genus and / or species as the cell to be genetically modified.

[0175] Table 1. Human proteins recruited to DNA breaks

[0176]

[0177]

[0178] Table 2. Mammalian FOX genes

[0179] Foxa1 Foxd2 Foxg2 Foxj3 Foxn3 Foxp3 Foxa2 Foxd3 Foxg1 Foxk1 Foxn4 Foxp4 Foxa3 Foxd4 Foxh1 Foxk2 Foxo1 Foxq1 Foxb1 Foxe1 Foxi1 Foxl1 Foxo3 Foxr1 Foxb2 Foxe3 Foxi2 Foxl2 Foxo4 Foxr2 Foxc1 Foxf1 Foxi3 Foxm1 Foxo6 Foxs1 Foxc2 Foxf2 Foxj1 Foxn1 Foxp1 Foxd1 Foxg1 Foxj2 Foxn2 Foxp2

[0180] Table 3. Human DNA damage binding genes

[0181] MUTYH MSH3 ERCC4 PCNA XRCC6 REV1 HMGB2 RAD1 APEX1 saga_people ERCC2 DDB1 BRCA1 NBN DCLRE1B ERCC3 RPA1 ddb1-ddb2_people XRCC5 BLM TDG POLK POLB FANCG POT1 tftc_people WRN NEIL1 XRCC1 GTF2H3 RBBP8 RPA4 CREBBP MSH2-MSH6_People EP300 POLQ DCLRE1A XPA H2AFX AUNIP OGG1 RPS3 RAD18 MSH6 RPA3 APTX CUL4B ERCC1 Q6ZNB5 UNG MSH5 RPA2 DDB2 TP53BP1 RAD23A RAD23B FEN1 POLD1 M0R2N6 MPG CRY2 HMGB1 POLI PNKP NEIL3 MSH2 POLH E9PQ18 RECQL4 NEIL2 MSH4 DCLRE1C XPC

[0182] Table 4: Human DNA repair

[0183]

[0184]

[0185]

[0186] Table 5: Yeast DNA repair

[0187]

[0188]

[0189] In an embodiment, the donor recruitment protein comprises a polypeptide sequence derived from a protein that is recruited to a DNA break, such as a double-stranded DNA break. In an embodiment, the polypeptide sequence is the portion of the protein that is recruited to the DNA break, particularly the portion of the protein that is responsible for recruitment to the DNA break. In an embodiment, the donor recruitment protein comprises a phospho-Ser / Thr-binding domain. In an embodiment, the phospho-Ser / Thr-binding domain is a 14-3-3 domain, a WW domain, a Polo-box domain (in PLK1), a WD40 repeat (including the E3 ligase SCF βTrCP In an embodiment, the donor recruitment protein comprises a polypeptide sequence derived from any of the proteins listed in Tables 1-5.

[0190] In certain embodiments, inhibitors of non-homologous end joining (NHEJ) pathways are used to further increase the frequency of genetic modification of cells by HDR. Examples of NHEJ pathway inhibitors include any compound (agent) that inhibits or blocks the expression or activity of any protein component in the NHEJ pathway. The protein components of the NHEJ pathway include, but are not limited to, Ku70, Ku86, DNA protein kinase (DNA-PK), Rad50, MRE11, NBS1, DNA ligase IV, and XRCC4. An exemplary inhibitor is wortmannin that inhibits at least one protein component (such as DNA-PK) in the NHEJ pathway. Another exemplary inhibitor is Scr7 (5,6-bis ((E)-benzylideneamino)-2-thiopyrimidine-4-ol), which inhibits DSB connection (Maruyama et al. (2015) Nat. Biotechnol. 33 (5): 538-542, Lin et al. (2016) Sci. Rep. 6: 34531). RNA interference or CRISPR interference can also be used to block the expression of protein components of the NHEJ pathway (such as DNA-PK or DNA ligase IV). For example, small interfering RNA (siRNA), hairpin RNA and other RNA or RNA: DNA species can be cut or separated in vivo to form siRNA, which can be used to inhibit the NHEJ pathway through RNA interference. Alternatively, inactivated Cas9 (dCas9) and single guide RNA (sgRNA) are complementary to the promoter or exon sequence of NHEJ pathway genes and can be used for transcriptional inhibition through CRISPR interference. Alternatively, HDR enhancers such as RS-1 can be used to increase the frequency of HDR in cells (Song et al. (2016) Nat. Commun. 7: 10548).

[0191] Barcoding is accomplished by integrating the genome editing cassette at a separate chromosomal locus (i.e., a chromosomal barcode locus) in each transfected cell that is different from the edited target locus. The genome editing cassette itself can be used as a barcode to identify genome editing of cells. Integration at the chromosomal barcode locus can avoid problems associated with plasmid instability in barcode retention.

[0192] In certain embodiments, the integration of the genome editing cassette at the chromosome barcode locus is performed with HDR. The recombinant polynucleotide can be designed with a pair of universal homology arms flanking the genome editing cassette, which can hybridize with complementary sequences at the chromosome barcode locus. In addition, each recombinant polynucleotide also includes a second guide RNA that can hybridize at the chromosome barcode locus. The formation of a complex between this second guide RNA and the RNA-guided nuclease can guide the RNA-guided nuclease to the chromosome barcode locus, wherein the RNA-guided nuclease generates a double-strand break at the chromosome barcode locus, and the genome editing cassette is integrated into the chromosome barcode by HDR.

[0193] In other embodiments, the genome editing cassette is integrated at the chromosome barcode using a site-specific recombinase system. Exemplary site-specific recombinase systems that can be used for this purpose include Cre-loxP site-specific recombinase system, Flp-FRT site-specific recombinase system, PhiC31-att site-specific recombinase system, and Dre-rox site-specific recombinase system. These and other site-specific recombinase systems that can be used to implement the present disclosure are described, for example, Wirth et al. (2007) Curr. Opin. Biotechnol. 18 (5): 411-419; Branda et al. (2004) Dev. Cell 6 (1): 7-28; Birling et al. (2009) Methods Mol. Biol. 561: 245-263; Bucholtz et al. (2008) J. Vis. Exp. May 29(15)pii:718; Nern et al. (2011) Proc. Natl. Acad. Sci. USA 108(34):14198-14203; Smith et al. (2010) Biochem. Soc. Trans. 38(2):388-394; Turan et al. (2011) FASEB J. 25(12):4088-4107; Garcia-Otin et al. (2006) Front. Biosci. 11:1108-1136; Gaj et al. (2014) Biotechnol Bioeng. 111(1):1-15; Krappmann (2014) Appl. Microbiol. Biotechnol. 98(5):1971-1982; Kolb et al. (2002) Cloning Stem Cells 4(1):65-80; and Lopatniuk et al. (2015) J. Appl. Genet. 56(4):547-550; incorporated herein by reference in their entirety.

[0194] A recombination target site for a site-specific recombinase is added to the chromosome barcode locus to allow integration by site-specific recombination. In addition, the recombinant polynucleotide is designed with a matching recombination target site for the site-specific recombinase, so that site-specific recombination between the recombination target site on the recombinant polynucleotide and the recombination target site on the chromosome barcode locus causes the genome editing cassette to be integrated at the chromosome barcode locus.

[0195] Alternatively or additionally, a unique barcode can be used to identify each target RNA-donor polynucleotide pair used for multiplex genome editing. Such barcodes can be inserted into the chromosomal barcode locus at each round of genome editing to identify the genome editing round number and the target RNA and / or donor polynucleotide used for genetic modification of the cell.

[0196] The barcode may include one or more nucleotide sequences for identifying the nucleic acid or cell associated with the barcode. The barcode may be 3-1000 or more nucleotides in length, preferably 10-250 nucleotides in length, more preferably 10-30 nucleotides in length, including any length within these ranges, such as 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides in length.

[0197] In some embodiments, the barcode is also used to identify the location of the cell, colony or sample from which the nucleic acid originated (i.e., position barcode), such as the colony location in a cell array or the well location in a multi-well plate. In particular, the barcode can be used to identify the location of genetically modified cells in a cell array.

[0198] In certain embodiments, barcoded cells are used for high-throughput positional barcoding of genetically modified cells, wherein the barcode sequence is used to identify the colony from which each gRNA and donor polynucleotide originate. The use of such barcodes allows gRNAs and donor polynucleotides from different cells to be collected into a single reaction mixture for sequencing, while still being able to trace a specific gRNA-donor polynucleotide combination to the colony from which it originated. Exemplary yeast barcoded cells are described in Smith et al. (2017) Mol. Syst. Biol. 13 (2): 913, which is incorporated herein by reference in its entirety.

[0199] In certain embodiments, genetically modified cells containing a library of gRNA-donor polynucleotide cassettes are initially seeded at separate locations in an ordered array. Barcoded cells are seeded in a matching array, and the gRNA-donor polynucleotide cassette from each genetically modified cell is introduced into each corresponding barcoded cell. For example, this can be accomplished by joining genetically modified cells to barcoded cells.

[0200] Example 1 describes the use of yeast, Saccharomyces cerevisiae, for this purpose. Saccharomyces cerevisiae exists in diploid and haploid forms. Joining occurs only between haploid forms of yeast of different joining types, which can be a or α joining types. The alleles (MATa or MATα) of the MAT locus determine the mating type. Diploid cells are produced from hybridization of MATa and MATα yeast strains. Therefore, haploid genetically modified yeast cells containing gRNA-donor polynucleotide boxes can be hybridized with haploid barcode yeast cells to generate diploid yeast cells, including gRNA-donor polynucleotide boxes and barcode sequences on different nucleic acids. For example, the genetically modified yeast cells of strain MATα can be joined with the barcode yeast cells of strain MATa. Alternatively, the genetically modified yeast cells of strain MATa can be joined with the barcode yeast cells of strain MATα.

[0201] The gRNA-donor polynucleotide box is translocated to the adjacent position of the barcode sequence to tag the box with a barcode, which can be accomplished with any suitable site-specific recombinase system. The site-specific recombinase catalyzes the DNA exchange reaction between two recombination target sites. A "recombination target site" is a region of a nucleic acid molecule, typically 30-50 nucleotides in length, comprising a binding site or sequence-specific motif recognized by a site-specific recombinase. After binding to the target site, the site-specific recombinase catalyzes the recombination of a specific DNA sequence at the target site. The relative orientation of the target site determines the outcome of the recombination, which can cause excision, insertion, inversion, translocation, or box exchange. If the recombination target site is on a different DNA molecule, translocation occurs. Site-specific recombinase systems typically include tyrosine recombinases or serine recombinases, but other types of site-specific recombinases may also be used in conjunction with their specific recombination target sites. Exemplary site-specific recombinase systems include Cre-loxP, Flp-FRT, PhiC31-att, and Dre-rox site-specific recombinase systems.

[0202] The recombination target site of the site-specific recombinase can incorporate the gRNA-donor polynucleotide cassette in several ways. For example, a polynucleotide containing a gRNA-donor cassette can be amplified with primers containing a recombination target site, which can undergo recombination with a barcode cell recombination target site. Alternatively, the gRNA-donor polynucleotide cassette can be integrated into the genome or plasmid of a host cell adjacent to the locus at the recombination target site, which can undergo recombination with a barcode cell recombination target site to produce a barcode-gRNA-donor polynucleotide fusion sequence. In addition, a selectable marker can be used to select clones that successfully undergo site-specific recombination.

[0203] In some cases, a cell population can be enriched for those containing a genetic modification by isolating the genetically modified cells from the remaining population. Isolating genetically modified cells typically relies on the expression of a selectable marker co-integrated with the desired editing at the target locus. After the donor polynucleotide is integrated by HDR, positive selection is performed to isolate cells from the population, such as to produce an enriched cell population containing a genetic modification.

[0204] Cell separation can be accomplished by any convenient separation technique appropriate for the selectable marker used, including but not limited to flow cytometry, fluorescence activated cell sorting (FACS), magnetic activated cell sorting (MACS), elutriation, immunopurification, and affinity chromatography. For example, if a fluorescent marker is used, cells can be isolated by fluorescence activated cell sorting (FACS), while if a cell surface marker is used, cells can be separated from a heterogeneous population by affinity separation techniques, such as MACS, affinity chromatography, "elutriation" with an affinity reagent attached to a solid matrix, immunopurification with a cell surface marker-specific antibody, or other convenient technique.

[0205] In certain embodiments, positive selection or negative selection of genetically modified cells is performed with a binding agent that specifically binds to a selection marker on the cell (e.g., produced from a selection marker expression cassette incorporated into a donor polynucleotide). Examples of binding agents include, but are not limited to, antibodies, antibody mimics, and aptamers. In some embodiments, the binding agent binds to the selection marker with high affinity. The binding agent can be fixed to a solid support to facilitate separation of the genetically modified cells from liquid culture. Exemplary solid supports include magnetic beads, non-magnetic beads, slides, gels, membranes, and microtiter plate wells.

[0206] In certain embodiments, the binding agent includes an antibody that specifically binds to a selectable marker on a cell. Any type of antibody can be used, including polyclonal and monoclonal antibodies, hybrid antibodies, altered antibodies, chimeric antibodies and humanized antibodies, as well as hybrid (chimeric) antibody molecules (see, e.g., Winter et al. (1991) Nature 349:293-299; and U.S. Pat. No. 4,816,567); F(ab')2 and F(ab) fragments; F vmolecules (non-covalent heterodimers, see, e.g., Inbar et al. (1972) Proc Natl Acad Sci USA 69:2659-2662; and Ehrlich et al. (1980) Biochem 19:4091-4096); single-chain Fv molecules (sFv) (see, e.g., Huston et al. (1988) Proc Natl Acad Sci USA 85:5879-5883); nanobodies or single domain antibodies (sdAbs) (see, e.g., Wang et al. (2016) Int J Nanomedicine 11:3287-3303, Vincke et al. (2012) Methods Mol Biol 911:15-26; dimeric and trimeric antibody fragment constructs; minibodies (see, e.g., Pack et al. (1992) Biochem 31:1579-1584; Cumber et al. (1992) J Immunology 149B:120-126); humanized antibody molecules (see, e.g., Riechmann et al. (1988) Nature 332:323-327; Verhoeyan et al. (1988) Science 239:1534-1536; and British Patent Publication No. GB ​​2,276,169, published September 21, 1994); and any functional fragments obtained from such molecules, wherein these fragments retain the specific binding properties of the parent antibody molecule (i.e., specific binding to a selectable marker on a cell).

[0207] In other embodiments, the binding agent includes an aptamer that specifically binds to a selection marker on a cell. Any type of aptamer can be used, including DNA, RNA, heteronucleic acid (XNA) or peptide aptamers that specifically bind to a target antibody isotype. For example, such aptamers can be identified by screening a combinatorial library. Nucleic acid aptamers (such as DNA or RNA aptamers) that selectively bind to a target antibody isotype can be generated as follows: repeated rounds of in vitro ligand selection or systematic evolution are completed by exponential enrichment (SELEX). Peptide aptamers that bind to a selection marker on a cell can be isolated from a combinatorial library and improved by directed mutagenesis or repeated rounds of mutation and selection. For descriptions of methods of generating aptamers, see, for example, Aptamers: Tools for Nanotherapy and MolecularImaging (RNVeedu ed., Pan Stanford, 2016), Nucleic Acid and Peptide Aptamers: Methods and Protocols (Methods in Molecular Biology, G. Mayer ed., Humana Press, 2009), Nucleic Acid Aptamers: Selection, Characterization, and Application (Methods in Molecular Biology, G. Mayer ed., Humana Press, 2016), AptamersSelected by Cell-SELEX for Theranostics (W. Tan, al.(2002)Nucleic Acids Res. 30(20):e108, Kenan et al. (1999) Methods Mol Biol. 118:217-231; Platella et al. (2016) Biochim. Biophys. Acta Nov 16pii:S0304-4165(16)30447-0, and Lyu et al. (2016) Theranostics 6(9):1440-1452; the entire texts are incorporated herein by reference.

[0208] In other embodiments, the binding agent comprises an antibody mimetic. Any type of antibody mimetic can be used, including but not limited to affibody molecules (Nygren (2008) FEBS J. 275 (11): 2668-2676), affilin (Ebersbach et al. (2007) J. Mol. Biol. 372 (1): 172-185), affimer (Johnson et al. (2012) Anal. Chem. 84 (15): 6553-6560), affitin (Krehenbrink et al. (2008) J. Mol. Biol. 383 (5): 1058-1068), alphabody (Desmet et al. (2014) Nature Communications 5: 5237), anticalin (Skerra (2008) FEBS J. 275(11):2677-2683), avimer (Silverman et al. (2005) Nat. Biotechnol. 23(12):1556-1561), darpin (Stumpp et al. (2008) Drug Discov. Today 13(15-16):695-701), fynomer (Grabulovski et al. (2007) J. Biol. Chem. 282(5):3196-3204) and monobody (Koide et al. (2007) Methods Mol. Biol. 352:95-109).

[0209] In positive selection, cells carrying the selection marker are collected, while in negative selection, cells carrying the selection marker are removed from the cell population. For example, in positive selection, a surface marker specific binding agent can be fixed to a solid support (such as a column or magnetic beads) and used to collect cells of interest on the solid support. Uninterested cells do not bind to the solid support (such as flow through the column or are not attached to the magnetic beads). In negative selection, a binding agent is used to deplete uninterested cell populations. Interested cells are those that are not bound to the binding agent (such as flow through the column or retained after removing the magnetic beads).

[0210] Dead cells are selected by using dyes that preferentially stain dead cells, such as propidium iodide. Any technique that does not unduly impair the viability of the genetically modified cells may be used.

[0211] Highly enriched compositions with required genetically modified cells can be generated in this way."Highly enriched" means that genetically modified cells account for 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, or 95% or more, or 98% or more of the cell composition. In other words, the composition can be the basic pure composition of genetically modified cells.

[0212] Genetically modified cells generated by the methods described herein can be used immediately. Alternatively, cells can be frozen at liquid nitrogen temperatures and stored for long periods of time before thawing and use. In this case, cells can be frozen in 10% DMSO, 50% serum, 40% buffered medium, or some other such solution, as is commonly used in the art to store cells at these freezing temperatures, and thawed in a manner generally known in the art for thawing frozen cultured cells.

[0213] The method steps of barcoding using RNA-guided nucleases, genome modification vectors containing guide RNA encoding expression cassettes and donor polynucleotides, and barcoded cells as described herein can be repeated to provide any desired number of DNA modifications encoded with barcodes.

[0214] Provided herein are methods for multiplexing genetically engineered cells, the methods comprising: (a) transfecting a plurality of cells with a plurality of different recombinant polynucleotides, each recombinant polynucleotide comprising a genome editing cassette comprising a first nucleic acid sequence encoding a first guide RNA (gRNA) that can hybridize at a genomic target locus to be modified and a donor polynucleotide, thereby forming a gRNA-donor polynucleotide combination, wherein each recombinant polynucleotide comprises a different genome editing cassette comprising a different gRNA-donor polynucleotide combination, and allowing each cell to express the first nucleic acid sequence, thereby forming a gRNA; and (b) introducing an RNA-guided nuclease into each of the plurality of cells, wherein the RNA-guided nuclease in each cell forms a complex with the gRNA, thereby forming a gRNA-RNA-guided nuclease complex, and allowing the gRNA-RNA-guided nuclease complex to modify the genomic target locus by integrating the donor polynucleotide into the genomic target locus, thereby generating a plurality of genetically engineered cells.

[0215] On the other hand, a method for multiplexing genetically engineered cells is provided, the method comprising: (a) transfecting a plurality of cells with a plurality of different recombinant polynucleotides, each recombinant polynucleotide comprising a unique polynucleotide barcode and a genome editing cassette, wherein the editing cassette comprises a first nucleic acid sequence encoding a first guide RNA (gRNA) that can hybridize at a genomic target locus to be modified and a donor polynucleotide, thereby forming a gRNA-donor polynucleotide combination, wherein each recombinant polynucleotide comprises a different genome editing cassette comprising a different gRNA-donor polynucleotide combination, and allowing each cell to express the first nucleic acid sequence, thereby forming a gRNA; and (b) introducing an RNA-guided nuclease into each of the plurality of cells, wherein the RNA-guided nuclease in each cell forms a complex with the gRNA, thereby forming a gRNA-RNA-guided nuclease complex, and allowing the gRNA-RNA-guided nuclease complex to modify the genomic target locus by integrating the donor polynucleotide into the genomic target locus, thereby generating a plurality of genetically engineered cells.

[0216] In an embodiment, the method further comprises sequence verification and arranging a plurality of genetically modified cells, the method comprising: (c) placing a plurality of genetically modified cells in an ordered array in a culture medium suitable for the growth of genetically modified cells; (d) culturing a plurality of genetically modified cells under certain conditions, wherein each genetically modified cell generates a clonal colony in the ordered array; (e) introducing a genome editing cassette from the ordered array colony into a barcoded cell, wherein the barcoded cell comprises a nucleic acid containing a site-specific recombinase recombination target site and a barcode sequence for identifying the colony position corresponding to the genome editing cassette in the ordered array; (f) transporting the genome editing cassette to the barcoded cell; The method further comprises repeating (e) to (h) for all colonies in the ordered array to identify the guide RNA and donor polynucleotide sequence of the genome editing cassette for each colony in the ordered array.

[0217] In an embodiment, each recombinant polynucleotide further comprises a second nucleic acid sequence encoding an RNA-guided nuclease. In an embodiment, the RNA-guided nuclease is provided by a vector or a second nucleic acid sequence integrated into the cell genome. In an embodiment, the genome editing cassette and RNA-guided nuclease are provided by a single vector or separate vectors.

[0218] In embodiments, the method further comprises identifying the presence of the donor polynucleotide in at least one of the plurality of genetically engineered cells. In embodiments, identifying the presence of the donor polynucleotide comprises identifying a barcode.

[0219] In an embodiment, the barcode is inserted into the genome of the plurality of genetically engineered cells at a chromosomal barcode locus.

[0220] In an embodiment, the RNA-guided nuclease is provided by a second nucleic acid sequence that is integrated into the chromosomal barcode locus, and wherein the barcode is inserted into the chromosomal barcode locus to remove the second nucleic acid sequence from the chromosomal barcode locus.

[0221] In an embodiment, the chromosomal barcode locus further comprises a promoter operably linked to any genome editing cassette first nucleic acid sequence that is integrated at the chromosomal barcode locus.

[0222] In an embodiment, each recombinant polynucleotide is provided by a vector. In an embodiment, the vector includes a promoter operably connected to a gRNA encoding polynucleotide. In an embodiment, the promoter is a constitutive or inducible promoter. In an embodiment, the vector is a plasmid or a viral vector. In an embodiment, the vector is a high copy number vector.

[0223] In embodiments, the RNA-guided nuclease is a Cas nuclease or an engineered RNA-guided FokI nuclease. In embodiments, the Cas nuclease is Cas9 or Cpf1.

[0224] In an embodiment, each recombinant polynucleotide further comprises a second nucleic acid sequence encoding a second guide RNA (guide X) that can hybridize with the recombinant polynucleotide, wherein the guide X forms a complex with the nuclease in each cell, thereby the guide X-nuclease complex cuts the recombinant polynucleotide. In an embodiment, the recombinant polynucleotide is a plasmid vector and the guide X-nuclease complex linearizes the plasmid vector. In an embodiment, the guide X-nuclease complex integrates at least a portion of the recombinant polynucleotide into a chromosomal barcode locus. In an embodiment, the nuclease is an RNA-guided nuclease. In an embodiment, the nuclease is a second RNA-guided nuclease introduced into a cell. In an embodiment, the second RNA-guided nuclease is a Cas nuclease or an engineered RNA-guided FokI nuclease. In an embodiment, the nuclease is selected from a large range of nucleases, FokI-nucleases, CRISPR-associated nucleases, zinc finger nucleases (ZFNs), and transcription activator-like effector nucleases (TALENs).

[0225] In embodiments, the donor polynucleotide is donor DNA.

[0226] In embodiments, each recombinant polynucleotide further comprises a DNA binding sequence known to bind a DNA binding domain.

[0227] In an embodiment, the method further comprises introducing into the cell a donor recruitment protein comprising a DNA binding domain and a DNA break site localization domain, wherein the localization domain selectively recruits the donor recruitment protein to the DNA break.

[0228] In an embodiment, the chromosomal barcode locus comprises a polynucleotide encoding an RNA-guided nuclease, a nuclease and / or a donor recruitment protein; and wherein the barcode is inserted into the chromosomal barcode locus to remove the polynucleotide encoding the RNA-guided nuclease, a nuclease and / or a donor recruitment protein from the chromosomal barcode locus.

[0229] In embodiments, each donor polynucleotide introduces a different mutation into the genomic DNA. In embodiments, the mutation is selected from the group consisting of an insertion, a deletion, and a substitution.

[0230] In embodiments, at least one donor polynucleotide introduces a mutation in the genomic DNA that inactivates a gene.

[0231] In embodiments, at least one donor polynucleotide removes a mutation from a gene in the genomic DNA.

[0232] In an embodiment, the plurality of recombinant polynucleotides can generate mutations at multiple sites in a single gene or non-coding region. In an embodiment, the plurality of recombinant polynucleotides can generate mutations at multiple sites in different genes or non-coding regions.

[0233] In embodiments, the method further comprises using a selectable marker that selects clones that have undergone successful integration of the donor polynucleotide at the genomic target locus or successful integration of the genome editing cassette at the chromosomal barcode locus.

[0234] In embodiments, the cell is a yeast cell. In embodiments, the yeast cell is a haploid yeast cell. In embodiments, the yeast cell is a diploid yeast cell.

[0235] In embodiments, the method further comprises inhibiting non-homologous end joining (NHEJ).

[0236] In embodiments, the genetically modified cell is a haploid yeast cell and the barcoded cell is a haploid yeast cell capable of conjugating with the genetically modified cell.

[0237] In embodiments, introducing a genome editing cassette from an ordered array colony into a barcoded cell comprises hybridizing a clone from the colony with the barcoded cell to generate a diploid yeast cell. In embodiments, the genetically modified cell is of strain MATa and the barcoded yeast cell is of strain MATa. In embodiments, the genetically modified cell is of strain MATa and the barcoded yeast cell is of strain MATa.

[0238] In embodiments, the genome editing cassette is flanked by restriction sites recognized by a meganuclease.In embodiments, the recombinase system in the barcoded cell uses a meganuclease to generate DNA double-strand breaks.

[0239] In an embodiment, the recombinase system in the barcoded cell is a Cre-loxP site-specific recombinase system, a Flp-FRT site-specific recombinase system, a PhiC31-att site-specific recombinase system or a Dre-rox site-specific recombinase system.

[0240] Another aspect provides an ordered array of colonies comprising clones of genetically modified cells generated by the methods described herein, wherein the colonies are indexed according to the verified sequences of their guide RNA and donor polynucleotides.

[0241] In another aspect, a method for locating a donor polynucleotide to a genomic target locus in a cell is provided, the method comprising: (a) transfecting a cell with a recombinant polynucleotide, the recombinant polynucleotide comprising a genomic editing cassette containing a donor polynucleotide and a DNA binding sequence known to bind a DNA binding domain; (b) introducing a nuclease into the cell, wherein the nuclease recognizes and causes a double-stranded DNA break at the genomic target locus; (c) introducing a donor recruitment protein into the cell, the donor recruitment protein comprising a DNA binding domain and a DNA break site localization domain and allowing the donor recruitment protein to selectively recruit DNA breaks, thereby locating the donor polynucleotide to the genomic target locus. In an embodiment, the DNA break is a double-stranded break.

[0242] In embodiments, the donor polynucleotide is positioned to the genomic target by loading a DNA repair enzyme onto the donor DNA. In embodiments, the donor polynucleotide is positioned to the genomic target locus by donor recruitment proteins interacting with one or more reagents (such as DNA repair enzymes, DNA break binding proteins and / or reagents generated or recruited therein at the DNA break) at the genomic target locus.

[0243] In embodiments, the donor recruitment protein is a fusion protein.

[0244] In embodiments, the DNA binding domain comprises a polypeptide sequence from a DNA binding protein. In embodiments, the DNA binding protein is selected from LexA, Gal4DBD, zinc finger proteins, TALEs, and transcription factors. In embodiments, the DNA binding protein is streptavidin, and wherein the biotin is conjugated to a donor polynucleotide. The DNA binding protein may include any protein known to bind DNA at a known DNA sequence.

[0245] In an embodiment, the DNA break site localization domain comprises a polypeptide sequence from a protein that binds to a DNA break site, such as a double-stranded DNA break site, or a region near a DNA break site caused by DNA breakage. In an embodiment, the protein that binds to a DNA break site or a region near a DNA break site caused by DNA breakage is a protein that participates in DNA repair. In an embodiment, the protein that participates in DNA repair is selected from a DNA break binding protein, a FOX transcription factor, and a protein from Table 1, Table 2, Table 3, Table 4, or Table 5.

[0246] In an embodiment, the nuclease is selected from a meganuclease, a FokI-nuclease, a CRISPR-associated nuclease, a zinc finger nuclease (ZFN), and a transcription activator-like effector nuclease (TALEN).

[0247] In an embodiment, the nuclease is an RNA-guided nuclease.

[0248] In embodiments, the nuclease modifies the genomic target locus by integrating the donor polynucleotide into the genomic target locus, thereby producing a genetically engineered cell.

[0249] In embodiments, the genetically modified cells are genetically engineered therapeutic cells. The genetically engineered therapeutic cells are genetically engineered immune cells. In embodiments, the genetically engineered immune cells are cancer-targeted T cells or natural killer cells.

[0250] On the other hand, a gene editing vector is provided, including a genome editing box, containing (i) a barcode, (ii) a first nucleic acid sequence encoding a first guide RNA (gRNA) that can hybridize at a target locus of the genome to be modified, and (iii) a donor polynucleotide, thereby forming a barcode-gRNA-donor polynucleotide combination.

[0251] On the other hand, a gene editing vector is provided, including a genome editing box, containing (i) a first nucleic acid sequence encoding a first guide RNA (gRNA) that can hybridize at a target locus of the genome to be modified, and (ii) a donor polynucleotide, thereby forming a gRNA-donor polynucleotide combination.

[0252] On the other hand, a gene editing vector library is provided, each gene editing vector comprising a genome editing box containing (i) a barcode, (ii) a first nucleic acid sequence encoding a first guide RNA (gRNA) that can hybridize at a target locus of the genome to be modified, and (iii) a donor polynucleotide, thereby forming a barcode-gRNA-donor polynucleotide combination; wherein each recombinant polynucleotide comprises a different genome editing box containing a different barcode-gRNA-donor polynucleotide combination.

[0253] On the other hand, a gene editing vector library is provided, each gene editing vector comprising a genome editing box containing (i) a first nucleic acid sequence encoding a first guide RNA (gRNA) that can hybridize at a target locus of the genome to be modified, and (ii) a donor polynucleotide, thereby forming a gRNA-donor polynucleotide combination; wherein each recombinant polynucleotide comprises a different genome editing box containing a different gRNA-donor polynucleotide combination.

[0254] In an embodiment, each vector further comprises a polynucleotide encoding a second guide RNA (Guide X) that can hybridize to the vector. In an embodiment, the Guide X can hybridize to a chromosomal barcode locus.

[0255] In embodiments, each vector further comprises a DNA binding sequence known to bind a DNA binding moiety.

[0256] In an embodiment, each vector further comprises a polynucleotide encoding an RNA-guided nuclease.

[0257] On the other hand, a gene editing vector is provided, comprising a donor polynucleotide and a first nucleic acid sequence, the first nucleic acid sequence encoding a first guide RNA (guide X) that can hybridize with the vector at the target site, so that when the cell expresses guide X, guide X hybridizes with the vector and forms a DNA break at the target site. In an embodiment, the vector includes a second nucleic acid sequence encoding a second guide RNA (gRNA) that can hybridize at the target locus of the genome to be modified. In an embodiment, the vector includes a DNA binding sequence that is known to bind to a DNA binding domain. In an embodiment, the vector includes a polynucleotide encoding a nuclease. In an embodiment, the nuclease is selected from a large range of nucleases, FokI-nucleases, CRISPR-associated nucleases, zinc finger nucleases (ZFNs), and transcription activator-like effector nucleases (TALENs).

[0258] On the other hand, a kit is provided, comprising: (a) a gene editing vector described herein, including embodiments thereof; and (b) a nuclease or a polynucleotide encoding a nuclease.

[0259] On the other hand, a kit is provided, comprising: (a) a gene editing vector described herein, including embodiments thereof; and (b) reagents for genetically modifying cells.

[0260] In an embodiment, each recombinant polynucleotide further comprises a second nucleic acid sequence encoding an RNA-guided nuclease.

[0261] On the other hand, a composition comprising a target cell, a nuclease, and a gene editing vector described herein is provided. In an embodiment, the composition includes a donor recruitment protein, the donor recruitment protein comprising a DNA binding portion and a DNA break site positioning portion that selectively recruits the donor recruitment protein to the DNA break site. In an embodiment, the target cell is a cell from a subject. In an embodiment, the subject suffers from cancer.

[0262] In embodiments, the target cell is an immune cell. In embodiments, the immune cell is a T cell.

[0263] In embodiments, the donor polynucleotide encodes a therapeutic agent. In embodiments, the therapeutic agent is a chimeric antigen receptor or a T cell receptor.

[0264] In embodiments, the subject suffers from a disease that can be treated by incorporating the donor DNA into the genome of the cell.

[0265] In embodiments, the cell is a human cell. In embodiments, the subject is a human.

[0266] B. Nucleic Acids Encoding Donor Polynucleotides, Guide RNAs, and RNA-Guided Nucleases

[0267] In certain embodiments, the gRNA-donor polynucleotide box and / or RNA-guided nuclease are expressed in vivo from a vector. A "vector" is a combination of elements (matter) that can be used to deliver a nucleic acid of interest to the interior of a cell. The gRNA-donor polynucleotide box and RNA-guided nucleic acid can be introduced into cells using a single vector or separate vectors. The ability of the construct to generate donor polynucleotides, guide RNAs, and RNA-guided nucleases (such as Cas9) and genetically modified cells can be determined empirically (e.g., see Example 1, describing the use of nutritional markers such as FCY1 and HIS3 in detecting genetically modified yeast cells).

[0268] Many vectors are known in the art, including but not limited to linear polynucleotides, polynucleotides associated with ions or amphoteric compounds, plasmids and viruses. Therefore, the term "vector" includes autonomously replicating plasmids or viruses. Examples of viral vectors include but are not limited to adenoviral vectors, adeno-associated viral vectors, retroviral vectors, lentiviral vectors, etc. Expression constructs can replicate in living cells or can be synthesized. For the purposes of this application, the terms "expression construct", "expression vector" and "vector" are used interchangeably to demonstrate the application of the present disclosure in a general and illustrative sense, and are not intended to limit the present disclosure.

[0269] In certain embodiments, the nucleic acid encoding the polynucleotide of interest is under the transcriptional control of a promoter. "Promoter" refers to a DNA sequence recognized by a cell synthesis machine or an introduced synthesis machine, which is required for the initiation of transcription of a specific gene. The term promoter is used herein to refer to a group of transcriptional control molecules that are clustered near the start site of RNA polymerase I, II or III. Typical promoters for mammalian cell expression include SV40 early promoter, CMV promoter such as CMV immediate early promoter (see, e.g., U.S. Patent Nos. 5,168,062 and 5,385,839, incorporated herein by reference in their entirety), mouse mammary tumor virus LTR promoter, adenovirus major late promoter (Ad MLP) and herpes simplex virus promoter, etc. Other non-viral promoters, such as promoters derived from mouse metallothionein genes, have also been found to be useful for mammalian expression. These and other promoters can be obtained from commercially available plasmids using techniques well known in the art. See, e.g., Sambrook et al., supra. Enhancer elements can be used in conjunction with promoters to increase construct expression levels. Examples include the SV40 early gene enhancer, as described in Dijkema et al., EMBO J. (1985) 4 :761, an enhancer / promoter derived from the Rous sarcoma virus long terminal repeat (LTR), as described in Gorman et al., Proc. Natl. Acad. Sci. USA (1982b) 79 :6777 and elements derived from human CMV as described in Boshart et al., Cell (1985)41 :521, for example, elements incorporating the CMV intron A sequence.

[0270] In one embodiment, an expression vector for expressing a donor polynucleotide, gRNA, or RNA-guided nuclease (such as Cas9) includes a promoter "operably" linked to a donor polynucleotide, gRNA, or RNA-guided nuclease encoding polynucleotide. As used herein, the phrase "operably linked" or "under transcriptional control" refers to a promoter that is in the correct position and orientation relative to a polynucleotide to control transcription initiation by an RNA polymerase and expression of a donor polynucleotide, gRNA, or RNA-guided nuclease.

[0271] Typically, a transcription terminator / polyadenylation signal is also present in the expression construct. Examples of such sequences include, but are not limited to, those derived from SV40, as described in Sambrook et al., supra, and the bovine growth hormone terminator sequence (see, e.g., U.S. Patent No. 5,122,458). In addition, 5'-UTR sequences can be placed near the coding sequence to increase its expression. Such sequences include UTRs containing internal ribosome entry sites (IRES).

[0272] Inclusion of an IRES can allow translation of one or more open reading frames from the vector. The IRES element attracts the eukaryotic ribosomal translation initiation complex and promotes the start of translation. See, e.g., Kaufman et al., Nuc. Acids Res. (1991) 19 :4485-4490; Gurtu et al., Biochem. Biophys. Res. Comm. (1996) 229 :295-298; Rees et al., BioTechniques (1996) 20 :102-110; Kobayashi et al., BioTechniques (1996) 21 :399-402; and Mosser et al., BioTechniques (1997 22 150-161. A large number of IRES sequences are known and include sequences derived from a wide variety of viruses, such as the leader sequence from picornaviruses, such as the encephalomyocarditis virus (EMCV) UTR (Jang et al. J. Virol. (1989) 63 :1651-1660), polio leader sequence, hepatitis A virus leader sequence, hepatitis C virus IRES, human rhinovirus type 2 IRES (Dobrikova et al., Proc. Natl. Acad. Sci. (2003) 100(25): 15125-15130), IRES element from foot-and-mouth disease virus (Ramesh et al., Nucl. Acid Res. (1996) 24 :2697-2700), Giardia virus IRES (Garlapati et al., J. Biol. Chem. (2004) 279 (5): 3389-3397) etc. Various non-viral IRES sequences also find use herein, including but not limited to IRES sequences from yeast, as well as human angiotensin II type 1 receptor (Martin et al., Mol. Cell Endocrinol. (2003) 212: 51-61), fibroblast growth factor IRES (FGF-1 IRES and FGF-2 IRES, Martineau et al. (2004) Mol. Cell. Biol. 24 (17): 7622-7635), vascular endothelial growth factor IRES (Baranick et al. (2008) Proc. Natl. Acad. Sci. USA 105 (12): 4733-4738, Stein et al. (1998) Mol. Cell. Biol. 18 (6): 3112-3119, Bert et al. (2006) RNA 12(6):1074-1083) and insulin-like growth factor 2 IRES (Pedersen et al. (2002) Biochem. J. 363(Pt 1):37-44). These elements are readily commercially available, for example, as plasmids sold by Clontech (Mountain View, CA), Invivogen (San Diego, CA), Addgene (Cambridge, MA) and GeneCopoeia (Rockville, MD). See also IRESite: The database of experimentally verified IRES structures (iresite.org). IRES sequences can be incorporated into vectors, for example, for expression of multiple selection markers or RNA-guided nucleases (such as Cas9) and one or more selection markers from an expression cassette.

[0273] Alternatively, polynucleotides encoding viral T2A peptides can be used to allow multiple protein products (such as Cas9, one or more selection markers) to be generated from a single vector. The 2A linker peptide is inserted between the coding sequences in the polycistronic construct. The 2A peptide self-cleaves, allowing co-expressed proteins from the polycistronic construct to be generated at equimolar levels. 2A peptides from a variety of viruses can be used, including but not limited to 2A peptides derived from foot-and-mouth disease virus, horse-type rhinitis virus, Thosea asigna virus, and porcine Teschovirus type 1. See, for example, Kim et al. (2011) PLoS One 6 (4): e18556; Trichas et al. (2008) BMC Biol. 6: 40, Provost et al. (2007) Genesis 45 (10): 625-629; Furler et al. (2001) Gene Ther. 8 (11): 864-873; incorporated herein by reference in their entirety.

[0274] In certain embodiments, the expression construct includes a plasmid suitable for transforming yeast cells. Yeast expression plasmids generally include yeast-specific origin of replication (ORI) and nutritional selection markers (such as HIS3, URA3, LYS2, LEU2, TRP1, MET15, ura4+, leu1+, ade6+), antibiotic selection markers (such as kanamycin resistance), fluorescent markers (such as mCherry) or other markers for selecting transformed yeast cells. Yeast plasmids may also include components to allow shuttling between bacterial hosts (such as Escherichia coli (E.coli)) and yeast cells. Several different types of yeast plasmids are available, including yeast integrating plasmids (YIp), which lack an ORI and integrate into the host chromosome by homologous recombination; yeast replicating plasmids (YRp), which contain an autonomously replicating sequence (ARS) and are capable of independent replication; yeast centromeric plasmids (YCp), which are low-copy vectors containing a partial ARS and a partial centromeric sequence (CEN); and yeast episomal plasmids (YEp), which are high-copy number plasmids that include a fragment from the 2μ circle (a native yeast plasmid) and allow stable propagation of 50 or more copies per cell.

[0275] Alternatively, bacterial plasmid vectors can be used to transform bacterial hosts. Many bacterial expression vectors are known to those skilled in the art, and selecting an appropriate vector is a matter of choice. Bacterial expression vectors include, but are not limited to, pACYC177, pASK75, pBAD, pBADM, pBAT, pCal, pET, pETM, pGAT, pGEX, pHAT, pKK223, pMal, pProEx, pQE, and pZA31 vectors. See, for example, Sambrook et al., supra.

[0276] In other embodiments, the expression construct comprises a virus or an engineered construct derived from a viral genome. Some virus-based systems have been developed to transfer genes to mammalian cells. These include adenovirus, retrovirus (γ-retrovirus and lentivirus), poxvirus, adeno-associated virus, baculovirus and herpes simplex virus (see, e.g., Warnock et al. (2011) Methods Mol. Biol. 737: 1-25; Walther et al. (2000) Drugs 60 (2): 249-271; and Lundstrom (2003) Trends Biotechnol. 21 (3): 117-122; incorporated herein by reference in its entirety). The ability of certain viruses to enter cells through receptor-mediated endocytosis, integrate into the host cell genome and stably and efficiently express viral genes makes them attractive candidates for transferring foreign genes to mammalian cells.

[0277] For example, retroviruses provide a convenient platform for gene delivery systems. Selected sequences can be inserted into vectors and packaged into retroviral particles using techniques known in the art. The recombinant virus can then be isolated and delivered to a subject's cells in vivo or ex vivo. Several retroviral systems have been described (U.S. Pat. No. 5,219,740; Miller and Rosman (1989) BioTechniques 7:980-990; Miller, AD (1990) Human Gene Therapy 1:5-14; Scarpa et al. (1991) Virology 180:849-852; Burns et al. (1993) Proc. Natl. Acad. Sci. USA 90:8033-8037; Boris-Lawrie and Temin (1993) Cur. Opin. Genet. Develop. 3:102-109; and Ferry et al. (2011) Curr. Pharm. Des. 17(24):2516-2527). Lentiviruses are a class of retroviruses that are specifically used for delivering polynucleotides to mammalian cells because they can infect both dividing and non-dividing cells (see, e.g., Lois et al. (2002) Science 295:868-872; Durand et al. (2011) Viruses 3(2):132-159; incorporated herein by reference).

[0278] Some adenoviral vectors have also been described. Unlike retroviruses that integrate into the host genome, adenoviruses remain extrachromosomal, thereby minimizing the risk associated with insertional mutagenesis (Haj-Ahmad and Graham, J. Virol. (1986) 57:267-274; Bett et al., J. Virol. (1993) 67:5911-5921; Mittereder et al., Human Gene Therapy (1994) 5:717-729; Seth et al., J. Virol. (1994) 68:933-940; Barr et al., Gene Therapy (1994) 1:51-58; Berkner, KL BioTechniques (1988) 6:616-629; and Rich et al., Human Gene Therapy (1993) 4:461-476). In addition, a variety of adeno-associated virus (AAV) vector systems have been developed for gene delivery. AAV vectors can be readily constructed using techniques well known in the art. See, e.g., U.S. Pat. Nos. 5,173,414 and 5,139,941; International Publication Nos. WO 92 / 01070 (published Jan. 23, 1992) and WO 93 / 03769 (published Mar. 4, 1993); Lebkowski et al., Molec. Cell. Biol. (1988) 8:3988-3996; Vincent et al., Vaccines 90 (1990) (Cold Spring Harbor Laboratory Press); Carter, BJ Current Opinion in Biotechnology (1992) 3:533-539; Muzyczka, N. Current Topics in Microbiol. and Immunol. (1992) 158:97-129; Kotin, RM Human Gene Therapy (1994) 5:793-801; Shelling and Smith, Gene Therapy (1994) 1: 165-169; and Zhou et al., J. Exp. Med. (1994) 179: 1867-1875.

[0279] Another vector system that can be used to deliver the polynucleotides of the present disclosure is the enterally administered recombinant poxvirus vaccine described by Small, Jr., PA, et al. (US Pat. No. 5,676,950, issued Oct. 14, 1997, incorporated herein by reference).

[0280] Additional viral vectors found to be useful for delivering nucleic acid molecules of interest include those derived from the poxvirus family, including vaccinia virus and fowlpox virus. For example, vaccinia virus recombinants expressing nucleic acid molecules of interest (such as donor polynucleotides, gRNA or RNA-guided nucleases) can be constructed as follows. DNA encoding a specific nucleic acid sequence is first inserted into a suitable vector so that it is adjacent to a vaccinia promoter and flanked by vaccinia DNA sequences, such as sequences encoding thymidine kinase (TK). This vector is then used to transfect cells that are simultaneously infected with vaccinia. Homologous recombination is used to insert a vaccinia promoter plus a gene encoding a sequence of interest into the viral genome. The resulting TK-recombinants can be selected as follows: cells are cultured in the presence of 5-bromodeoxyuridine, from which tolerant viral plaques are picked out.

[0281] Alternatively, avian pox viruses such as fowlpox virus and canarypox virus can also be used to deliver nucleic acid molecules of interest. The use of avian pox vectors is particularly desirable in humans and other mammalian species because members of the genus Avipox productively replicate only in susceptible avian species and are therefore not infectious in mammalian cells. Methods for producing recombinant avian pox viruses are known in the art and employ genetic recombination, as described above with respect to the production of vaccinia viruses. See, for example, WO 91 / 12882; WO 89 / 03429; and WO 92 / 03545.

[0282] Molecular conjugate vectors such as the adenovirus chimeric vectors described by Michael et al., J. Biol. Chem. (1993) 268: 6866-6869 and Wagner et al., Proc. Natl. Acad. Sci. USA (1992) 89: 6099-6103 can also be used for gene delivery.

[0283] Members of the genus Alphavirus, such as, but not limited to, vectors derived from Sindbis virus (SIN), Semliki Forest virus (SFV), and Venezuelan equine encephalitis virus (VEE), also find use as viral vectors to deliver the polynucleotides of the present disclosure. For a description of Sindbis virus-derived vectors useful in practicing the present methods, see Dubensky et al. (1996) J. Virol. 70:508-519; and International Publication Nos. WO 95 / 07995, WO 96 / 17072; and Dubensky, Jr., TW, et al., U.S. Pat. No. 5,843,723, issued Dec. 1, 1998, and Dubensky, Jr., TW, U.S. Pat. No. 5,789,245, issued Aug. 4, 1998, both of which are incorporated herein by reference. Particularly preferred are chimeric alphavirus vectors composed of sequences derived from Sindbis virus and Venezuelan equine encephalitis virus. See, e.g., Perri et al. (2003) J. Virol. 77: 10394-10403 and International Publication Nos. WO 02 / 099035, WO 02 / 080982, WO 01 / 81609 and WO 00 / 61772; incorporated herein by reference in their entireties.

[0284] The infection / transfection system based on cowpox can be conveniently used to provide inducible transient expression of polynucleotides of interest (such as gRNA-donor polynucleotide boxes, polynucleotides encoding RNA-guided nucleases) in host cells. In this system, cells are first infected in vitro with cowpox virus recombinants encoding bacteriophage T7 RNA polymerase. This polymerase exhibits strong specificity in that it only transcribes templates carrying T7 promoters. After infection, cells are transfected with polynucleotides of interest driven by T7 promoters. The polymerase expressed by the cytoplasm of the cowpox virus recombinant transcribes the transfected DNA into RNA. The method provides high-level, transient, cytoplasmic generation of large amounts of RNA. See, for example, Elroy-Stein and Moss, Proc. Natl. Acad. Sci. USA (1990) 87: 6743-6747; Fuerst et al., Proc. Natl. Acad. Sci. USA (1986) 83: 8122-8126.

[0285] As an alternative to infecting with cowpox or fowlpox virus recombinants or delivering nucleic acids with other viruses, an amplification system that causes high-level expression after introducing into host cells can be used. In particular, the T7 RNA polymerase promoter in front of the T7 RNA polymerase coding region can be transformed. From this template, RNA translation will produce T7 RNA polymerase, which then transcribes more templates. Concomitantly, there is a cDNA expressed under the control of the T7 promoter. Therefore, some T7 RNA polymerases generated from the amplification template RNA can cause the desired gene to be transcribed. Because the initial amplification requires some T7 RNA polymerases, the T7 RNA polymerase can be introduced into cells together with the template to initiate a transcriptional reaction. Polymerase can be introduced as a protein or on a plasmid encoding RNA polymerase. For further discussion of the T7 system and its use for transforming cells, see, e.g., International Publication No. WO 94 / 26911; Studier and Moffatt, J. Mol. Biol. (1986) 189: 113-130; Deng and Wolff, Gene (1994) 143: 245-249; Gao et al., Biochem. Biophys. Res. Commun. (1994) 200: 1201-1206; Gao and Huang, Nuc. Acids Res. (1993) 21: 2867-2872; Chen et al., Nuc. Acids Res. (1994) 22: 2114-2120; and U.S. Pat. No. 5,135,855.

[0286] Insect cell expression systems such as baculovirus systems can also be used and are known to those skilled in the art and described, for example, in Baculovirus and Insect Cell Expression Protocols (Methods in Molecular Biology, DW Murhammer ed., Humana Press, 2 nd (Springer, 1992). Materials and methods for baculovirus / insect cell expression systems are commercially available in kit form, particularly from Thermo Fisher Scientific (Waltham, MA) and Clontech (Mountain View, CA).

[0287] Plant expression systems can also be used to transform plant cells. Typically, such systems use virus-based vectors to transfect plant cells with heterologous genes. These systems are described, for example, in Porta et al., Mol. Biotech. (1996).5 :209-221; and Hackland et al., Arch. Virol. (1994) 139 :1-22.

[0288] To achieve expression of a sense or antisense gene construct, the expression construct must be delivered into the cell. This delivery can be accomplished in vitro, such as in laboratory procedures for transforming cell lines, or in vivo or ex vivo, such as for treating certain disease states. One delivery mechanism is via viral infection, wherein the expression construct is encapsulated in infectious viral particles.

[0289] The present disclosure also contemplates several non-viral methods for transferring expression constructs into cultured mammalian cells. These include the use of calcium phosphate precipitation, DEAE-dextran, electroporation, direct microinjection, DNA-loaded liposomes, liposome-DNA complexes, cell sonication, gene bombardment with high-speed microprojectiles, and receptor-mediated transfection (see, e.g., Graham and Van Der Eb (1973) Virology 52:456-467; Chen and Okayama (1987) Mol. Cell Biol. 7:2745-2752; Rippe et al. (1990) Mol. Cell Biol. 10:689-695; Gopal (1985) Mol. Cell Biol. 5:1188-1190; Tur-Kaspa et al. (1986) Mol. Cell. Biol. 6:716-718; Potter et al. (1984) Proc. Natl. Acad. Sci. USA 81:7161-7165); Harland and Weintraub (1985) J. Cell Biol. 101: 1094-1099); Nicolau and Sene (1982) Biochim. Biophys. Acta 721: 185-190; Fraley et al. (1979) Proc. Natl. Acad. Sci. USA 76:3348-3352; Fechheimer et al. (1987) ProcNatl.Acad.Sci.USA 84:8463-8467; Yang et al. (1990) Proc.Natl.Acad.Sci.USA 87:9568-9572; Wu and Wu (1987) J. Biol. Chem. 262:4429-4432; Wu and Wu (1988) Biochemistry 27:887-892; incorporated herein by reference). Some of these techniques can be successfully adapted for in vivo or ex vivo applications.

[0290] Once the expression construct is delivered into the cell, the nucleic acid encoding the gene of interest can be located at different sites and expressed at the site. In certain embodiments, the nucleic acid encoding the gene can be stably integrated into the cell genome. This integration can be in homologous position and direction (gene replacement) through homologous recombination, or can be integrated at random non-specific positions (gene enhancement). In certain embodiments, the nucleic acid can be stably maintained in the cell as a separate free segment of DNA. This nucleic acid segment or "episome" encodes a sequence that is sufficient to allow maintenance and replication, independently or synchronously with the host cell cycle. How the expression construct is delivered to the cell and where the nucleic acid remains in the cell depends on the expression construct type used.

[0291] In another embodiment of the present disclosure, the expression construct may simply consist of naked recombinant DNA or a plasmid. Construct transfer may be performed by any of the above methods, which physically or chemically permeate the cell membrane. This is particularly useful for in vitro transfer, but it may also be used in vivo. Dubensky et al. (Proc. Natl. Acad. Sci. USA (1984) 81: 7529-7533) successfully injected polyomavirus DNA into the liver and spleen of adult and newborn mice in the form of calcium phosphate precipitates, showing active viral replication and acute infection. Benvenisty and Neshif (Proc. Natl. Acad. Sci. USA (1986) 83: 9551-9555) also demonstrated that direct intraperitoneal injection of calcium phosphate precipitated plasmids caused transfected gene expression. It is envisioned that DNA encoding a gene of interest may also be transferred in vivo in a similar manner and express the gene product.

[0292] In another embodiment, naked DNA expression constructs can be transferred into cells by particle bombardment. This method depends on the ability to accelerate DNA-coated microprojectiles to high speeds, allowing them to pierce cell membranes and enter cells without killing them (Klein et al. (1987) Nature 327:70-73). Several devices have been developed for accelerating small particles. One such device depends on high voltage discharge to generate an electric current, which in turn provides power (Yang et al. (1990) Proc. Natl. Acad. Sci. USA 87:9568-9572). Microprojectiles can be composed of biologically inert materials such as tungsten or gold beads.

[0293] In another embodiment, the expression construct can be delivered by liposomes. Liposomes are porous structures characterized by a phospholipid bilayer membrane and an internal aqueous medium. Multilamellar liposomes have multiple lipid layers separated by an aqueous medium. When phospholipids are suspended in an excess aqueous solution, they form spontaneously. The lipid components undergo self-rearrangement before the closed structure forms, trapping water and dissolved solutes between the lipid bilayers (Ghosh and Bachhawat (1991), Liver Diseases, Targeted Diagnosis and Therapy Using Specific Receptors and Ligands, Wu et al. (eds.), Marcel Dekker, NY, 87-104). It is also contemplated to use liposome-DNA complexes.

[0294] In certain embodiments of the present disclosure, the liposomes can be complexed with hemagglutinating virus (HVJ). This has been shown to facilitate cell membrane fusion and promote cell entry of liposome-encapsulated DNA (Kaneda et al. (1989) Science 243: 375-378). In other embodiments, the liposomes can be complexed or used in combination with nuclear non-histone chromosomal proteins (HMG-I) (Kato et al. (1991) J. Biol. Chem. 266 (6): 3361-3364). In other embodiments, the liposomes can be complexed or used in combination with HVJ and HMG-I. Since such expression constructs have been successfully used to transfer and express nucleic acids in vitro and in vivo, they are suitable for the present disclosure. When bacterial promoters are used in DNA constructs, it is also necessary to incorporate suitable bacterial polymerases into the liposomes.

[0295] Other expression constructs that can be used to deliver nucleic acids to cells are receptor-mediated delivery vehicles. These utilize selective uptake of macromolecules by receptor-mediated endocytosis in almost all eukaryotic cells. Due to the cell type-specific distribution of multiple receptors, delivery can be highly specific (Wu and Wu (1993) Adv. Drug Delivery Rev. 12: 159-167).

[0296] Receptor-mediated gene targeting vectors generally consist of two components: a cell receptor-specific ligand and a DNA binder. Several ligands are used for receptor-mediated gene transfer. The most widely identified ligands are asialoglycoprotein (ASOR) and transferrin (see, e.g., Wu and Wu (1987), supra; Wagner et al. (1990) Proc. Natl. Acad. Sci. USA 87 (9): 3410-3414). Recently, synthetic pseudoglycoproteins that recognize the same receptor as ASOR have been used as gene delivery vectors (Ferkol et al. (1993) FASEB J. 7: 1081-1091; Perales et al. (1994) Proc. Natl. Acad. Sci. USA 91 (9): 4086-4090), and epidermal growth factor (EGF) is used to deliver genes to squamous cell carcinoma cells (Myers, EPO 0273085).

[0297] In other embodiments, the delivery vehicle may include a ligand and a liposome. For example, Nicolau et al. (Methods Enzymol. (1987) 149: 157-176) used lactosylceramide, a galactose-terminal asialoganglioside, into liposomes and observed an increase in the uptake of insulin genes by hepatocytes. Thus, nucleic acids encoding specific genes may also be specifically delivered to cells by any number of receptor-ligand systems, with or without liposomes. Similarly, antibodies to surface antigens can similarly be used as targeting moieties.

[0298] In a specific example, the recombinant polynucleotide encoding the gRNA-donor polynucleotide box or the RNA-guided nuclease can be co-administered with a cationic lipid. Examples of cationic lipids include, but are not limited to, lipofectin, DOTMA, DOPE, and DOTAP. WO / 0071096 discloses specific incorporation by reference, describing different preparations that can be effectively used for gene therapy, such as DOTAP: cholesterol or cholesterol-derived preparations. Other publications also discuss different lipid or liposome preparations, including nanoparticles and methods of administration; these include, but are not limited to, U.S. Patent Publications 20030203865, 20020150626, 20030032615, and 20040048787, which are specifically incorporated by reference to disclose the preparations and other related aspects of nucleic acid delivery. Methods for forming particles are also disclosed in US Pat. Nos. 5,844,107, 5,877,302, 6,008,336, 6,077,835, 5,972,901, 6,200,801, and 5,972,900, which are incorporated by reference for these aspects.

[0299] In certain embodiments, gene transfer can be more easily performed under ex vivo conditions. Ex vivo gene transfer refers to isolating cells from a subject, delivering nucleic acid to the cells in vitro, and then returning the modified cells to the subject. This can involve collecting a biological sample, including cells from the subject. For example, blood can be obtained by venipuncture, and solid tissue samples can be obtained by surgical techniques according to methods well known in the art.

[0300] Typically, but not always, the subject receiving the cell (i.e., the recipient) is also the subject from which the cell is harvested or obtained, with the advantage that the donated cells are autologous. However, the cells can be obtained from another subject (i.e., the donor), from a cell culture of the donor, or from an established cell culture system. The cells may be obtained from the same or different species as the subject being treated, but preferably the same species, more preferably having the same immune profile as the subject. For example, such cells can be obtained from a biological sample, including cells from a close relative or a matched donor, and then transfected with a nucleic acid (e.g., encoding a donor polynucleotide, gRNA, or RNA-guided nuclease), and administered to a subject in need of genome modification, for example, for the treatment of a disease or disorder.

[0301] C. Sequencing of the Barcoded gRNA-Donor Polynucleotide Cassette

[0302] Any high throughput technology for sequencing can be used to implement the present disclosure. DNA sequencing techniques include dideoxy sequencing reactions (Sanger method), using labeled terminators or primers and gel separation in plates or capillaries, sequencing by synthesis using reversible terminator labeled nucleotides, pyrophosphate sequencing, 454 sequencing, sequencing by synthesis followed by ligation using allele-specific hybridization labeled clone libraries, real-time monitoring of labeled nucleotide incorporation during the polymerization step, polymerase clone sequencing, SOLID sequencing, etc.

[0303] Certain high-throughput methods of sequencing include a step in which individual molecules are spatially separated on a solid surface where they are sequenced in parallel. Such solid surfaces may include non-porous surfaces (such as Solexa sequencing, e.g., Bentley et al., Nature, 456:53-59 (2008), or whole genome sequencing, e.g., Drmanac et al., Science, 327:78-81 (2010)), well arrays, which may include bead- or particle-bound templates (such as with 454, e.g., Margulies et al., Nature, 437:376-380 (2005), or Ion Torrent sequencing, U.S. Patent Publication Nos. 2010 / 0137143 or 2010 / 0304982), micromachined membranes (such as with SMRT sequencing, e.g., Eid et al., Science, 323:133-138 (2009)), or bead arrays (such as with SOLiD sequencing or polymerase cloning sequencing, e.g., Kim et al., Science, 316:1481-1414 (2007)). Such methods may include amplifying the isolated molecules, either before or after their spatial segregation on a solid surface. Existing amplification may include emulsion-based amplification such as emulsion PCR, or rolling circle amplification.

[0304] Of particular interest is sequencing on the Illumina MiSeq, NextSeq and HiSeq platforms, which use reversible terminator sequencing by synthesis technology (see, e.g., Shen et al. (2012) BMC Bioinformatics 13:160; Junemann et al. (2013) Nat. Biotechnol. 31(4):294-296; Glenn (2011) Mol. Ecol. Resour. 11(5):759-769; Thudi et al. (2012) Brief Funct. Genomics 11(1):3-11; incorporated herein by reference).

[0305] These sequencing methods can thus be used to sequence barcoded gRNA-donor polynucleotide cassettes to associate their sequences with adjacent (shorter) barcodes, identifying their corresponding colonies in an ordered array. Short DNA barcodes can also be used to multiplex sequence ordered array samples. Thus, clones containing any desired gRNA-donor polynucleotide cassettes can then be picked out of an ordered array of cells (e.g., using an automated robotic device or manually).

[0306] D. Test kit

[0307] The above reagents can be provided in a kit, including a recombinant polynucleotide encoding a gRNA-donor polynucleotide box, an RNA-guided nucleic acid, a barcode cell, a culture medium suitable for cell growth, and a site-specific recombinase system, with suitable instructions and other necessary reagents for genome modification and barcoding as described herein. The kit may also include cells for genome modification, reagents for positive selection and negative selection of cells, and a transfection agent. The kit typically includes a gRNA-donor polynucleotide box, an RNA-guided nucleic acid, a barcode cell, a culture medium suitable for cell growth, and a site-specific recombinase system, and other required reagents in a separate container. Instructions for completing the genome modification and barcoding as described herein (such as written, CD-ROM, DVD, Blu-ray, flash drive, digital download, etc.) are typically included in the kit. Depending on the specific experiment used, the kit can also include other packaging reagents and materials (i.e., wash buffer, etc.). The genome editing and barcoding described herein can be implemented with these kits.

[0308] E. Application

[0309] The genome editing and barcoding methods disclosed herein have multiple applications in basic research and development and regenerative medicine. The method can be used to introduce mutations (such as insertions, deletions or substitutions) into any gene in the genomic DNA of a cell. For example, the methods described herein can be used to inactivate cell genes to determine the effects of gene knockout or to study the effects of known pathogenic mutations. Such genetically modified cells can be used as disease models for drug screening. Alternatively, the methods described herein can be used to remove mutations, such as pathogenic mutations, from genes in the genomic DNA of a cell. In particular, the genome editing described herein can be used to develop cell lines with desired characteristics, such as adding reporter genes at desired sites, or improving efficacy, controllable safety and / or survival.

[0310] In particular, the disclosed method is used to generate a collection of arrayed strains with known genetic modifications for a variety of purposes, including but not limited to protein engineering, DNA variant generation, strain engineering, metabolic engineering, or drug screening. Mutated strains can be arranged in an array according to their known gRNA and donor polynucleotide sequences, and the position depends on, for example, the targeted chromosomal locus or the modified gene. In addition, the strain can be phenotyped to determine the impact of a specific mutation. The arrayed strains can be grown in a culture plate or a liquid culture medium. For example, the strain can be parsed into an array containing a culture plate or a separate tube containing a culture medium. Subsequently, any colony or colony combination with a genetic modification of interest can be selected from the arrayed strains to inoculate a liquid culture medium and grow in batches. Several rounds of genome modification can be performed to optimize desired properties, such as increasing biomass, improving growth under different conditions, or optimizing metabolic production of different compounds.

[0311] In certain embodiments, methods described herein are used to generate arrays of genetically modified yeast strains. For example, such arrays of yeast strains can be used to optimize the production of bread, beer, wine, biofuels, animal-free production of antibodies, enzymes and other proteins, and other yeast-based technologies. Genetically modified yeast strains are also found to be useful for drug screening, metabolic production of compounds, vaccine production, pathogen detection, and generation of DNA and protein variants.

[0312] III. Experiment

[0313] The following are examples of specific embodiments to accomplish the present disclosure. The examples are provided for illustrative purposes only and are not intended to limit the scope of the present disclosure in any way.

[0314] Efforts have been made to ensure accuracy with respect to numbers used (eg amounts, temperature, etc.) but some experimental errors and deviations should, of course, be allowed for.

[0315] Example 1. Scarless genome editing via two-step homology-mediated repair

[0316] introduction

[0317] We previously described a cost-effective method, called recombinase-mediated indexing (REDI), which involves the integration of multiplex libraries into yeast, site-specific recombination to index the library DNA, and next-generation sequencing to identify desired clones. REDI was originally developed to generate high-quality DNA libraries, circumventing the high synthesis error rates and inability to obtain individual oligonucleotides associated with array-synthesized oligonucleotides. It has also been used to rapidly generate CRISPRi panels to transcriptionally repress essential yeast open reading frames (ORFs).

[0318] Here we extend this technology for massively parallel production of genetically engineered clones. Our approach involves large-scale, highly efficient genome editing using a plasmid system that facilitates integration of gRNA and donor dummies as genomic barcodes, allowing identification, isolation, and massively parallel validation of individual variants from libraries of transformants. Importantly, we also outline key strategies to enhance HR in metazoan cells, including CRISPR-interference (CRISPRi), RNA interference (RNAi), or chemical-based inhibition of NHEJ, combined with active donor recruitment.

[0319] result

[0320] We previously described an inexpensive, high-throughput, yeast-based method for profiling validated sequences from complex mixtures, termed recombinase-mediated indexing or REDI 17. Based on the REDI system, a dual editing-barcoding system is now described, involving CRISPR / Cas9-mediated editing of the target genomic locus, using a high-copy plasmid to carry donor DNA, followed by SceI-mediated capture of the target worker cassette into the REDI locus. The integration of gRNA and donor DNA sequences as barcodes enables (1) strain isolation by REDI and (2) robust phenotyping after competitive growth. The high-copy (2-micron) nature of the guide-donor plasmid enables efficient repair. The guide RNA-donor DNA cassette is integrated into the REDI locus, generating precisely one barcode molecule per cell, thereby circumventing noise that may arise from copy number variation and plasmid loss, which is a characteristic of the vector-based barcode and confounds phenotyping accuracy.

[0321] To allow for the parallel generation of many genetically engineered variants, we used gRNA / donor DNA pairs synthesized on the same oligonucleotide molecule. Therefore, an internal cloning step was used to maintain the oligonucleotide length below the limits of array-based synthesis and to avoid erroneous incorporation of DNA synthesis into the constant portion of the guide RNA sequence (Fig. S1). The internal cloning step inserted this DNA as a sequence-perfect insert (see Fig. 9 ,method).

[0322] Efficient isolation of edited clones after transformation with libraries encoding thousands to millions of genomic modifications requires a system with optimal editing efficiency. Therefore, we systematically evaluated multiple parameters of the CRISPR / Cas9 editing system in yeast, including the promoters for Cas9 and guide RNA expression. We found that the tRNA-HDV promoter for guide RNA expression yielded the best editing efficiency. We also examined the importance of Cas9 expression levels, as this parameter was found to be important in previous yeast studies using linear donor DNA. 13 . We generated a construct targeting the yeast ADE2 locus that generated yeast with a characteristic red color when mutated. The donor DNA was designed to incorporate a frameshift mutation at the ADE2 locus and to tolerate recognition by its companion guide RNA (see Methods). This construct was co-transformed into yeast with a Cas9 expression construct, or transformed into yeast pre-expressing Cas9. When transformed into yeast pre-expressing Cas9, nearly all clones incorporated the desired changes encoded by the donor DNA, as shown by the predominance of red colonies ( Figure 2A , top right). Sequencing at the ADE2 locus also verified that the desired changes had been incorporated into 6 independent clones ( Figure 2B Importantly, these experiments showed that cell death occurs in the absence of donor DNA, rather than survival via the error-prone NHEJ pathway that dominates in most other systems ( Figure 2A, upper left). Thus, expression of Cas9 under a strong constitutive promoter results in cells that are strictly dependent on the plasmid-borne donor DNA for survival, and only clones that survive the transformation accurately incorporate the changes mediated by the donor DNA.

[0323] We demonstrate that transformation of a plasmid carrying Cas9 into cells pre-expressing guide RNA can result in similarly high levels of editing efficiency and survival. The improved survival may be attributed to providing sufficient time for the guide-donor plasmid to accumulate high copy numbers, resulting in enhanced DNA break repair. Additionally, we tested an inducible promoter (Gal1 promoter) for Cas9 and found that it provided equally efficient editing.

[0324] We next sought to demonstrate that genomic barcode integration at the REDI locus could be readily achieved following target editing. This was accomplished in 2 different ways. ADE2 edited cells were moved to galactose medium to induce SceI expression and cleavage of the SceI sites flanking the FCY1 counter-selectable marker at the REDI locus. This high-throughput genomic integration approach is an extension of our previous approach described for transformation oligonucleotide integration. 17 The guide RNA-donor DNA cassette from the plasmid was efficiently incorporated into all clones tested ( Figure 2C Alternatively, we used gRNAs targeting the SceI site or the counter-selectable FCY1 gene, along with CRISPR for genome editing to integrate the editing cassette into the REDI locus, while using CRISPR cleavage. Therefore, these clones were hybridized to the REDI barcoded strain, and paired-end Illumina sequencing was then used to identify and isolate these clones from a highly complex library of variants.

[0325] To establish the scalability of our approach, gRNA-donor DNA was designed and purchased (Agilent Technologies) to allow single amino acid saturation mutagenesis of a heterologous ORF (mCherry). To make complete saturation of the ORF feasible and to ensure that cells incorporate the desired changes, we designed a novel synonymous codon spreading strategy to enable editing at sites outside the guide recognition region (Figure 10). We selected some guide-donors isolated from the library to verify their functionality. Unexpectedly, one guide-donor plasmid caused high toxicity and low survival ( Figure 3 , right). This guide RNA targets the initial methionine codon (ATG) and the adjacent TPI1 promoter sequence. The same guide RNA target sequence is also present in the native yeast TPI1 gene. Therefore, the construct is expected to induce double-strand breaks at 2 locations in the yeast genome. Despite containing homology to the TPI1 promoter, the donor DNA lacks any homology to the start of the TPI1 ORF, indicating that target cleavage site repair requires sufficient homology on both sides of the dsDNA break. This is consistent with our data demonstrating the toxicity of pre-expressed Cas9 and gRNA in the absence of donor DNA ( Figure 2A), this result suggests that gRNAs with strong off-target effects may cause cell death after transformation if there is no donor DNA to repair these breaks. Importantly, we expect that these guide-donor sequences are not captured by the REDI separation protocol and therefore do not result in false positives or negatives. This indicates that the editing system we describe has extremely high fidelity, emphasizing its utility for exploring the genome-wide effects of natural and artificial variants.

[0326] We noted that transformation of cells with plasmids with functional gRNA and pre-expression of Cas9 resulted in significantly fewer colonies (~10-fold) relative to plasmids containing non-functional guide RNAs, indicating that ~90% of cells transformed with guide-donors were unable to complete homology repair despite the presence of the donor on the plasmid in the nucleus. We infer that Rad51-mediated homology search for donor DNA may be rate-limiting for cell survival in our system.

[0327] To test this hypothesis, we developed a system to actively recruit donors to dsDNA break sites (Figure 4). Note that less than ~0.01% of transformants survived from Cas9-gRNA expression in the absence of donor DNA, and less than ~10% survived in the presence of donor DNA. All survivors incorporated sequence changes specified by the donor DNA, indicating that the vast majority of survivors use homologous recombination to repair dsDNA breaks. In addition, the presence of non-functional gRNA sequences in the combined editing experiments caused a significant bottleneck for editing cells and enriched for gRNA libraries that did not produce any genome modifications. This is an important issue to address because the typical array synthesis error of 1 / 200 is expected to result in ~10% of 20-mer guide sequences containing at least one error [(1-1 / 200)^20~0.1].

[0328] To increase the fraction of cells that survive the editing process and reduce bottlenecks, a system for active donor recruitment was implemented, reasoning that random diffusion of donor DNA to the cut site is rate-limiting for homologous repair. We co-expressed the LexA DNA binding domain (DBD) fused to Fkh1 (Fkh1 binds to the HML recombination enhancer and regulates donor preference during mating type switching, which is achieved by recruiting DNA with Fkh1 that binds to donor DNA (Saccharomyces Genome Database, Li et al. (2012) PLoS Genet. 8(4): e1002630) and Cas9, and transformed the guide-donor plasmid containing the LexA binding site (Figure 4). We also designed a system that directly fused LexA to Cas9 to ensure that the donor is present in parallel with dsDNA cutting. This resulted in a significant increase in survival and the efficiency of homologous recombination-mediated precise editing ( Figure 5AWe are currently testing Cas9-LexA DBD fusions, which are expected to produce similar increases in editing efficiency and should be generally applicable to all models in which RGNs can be introduced.

[0329] The 2-micron plasmid requires multiple cell generations to accumulate to its highest level in the nucleus. Therefore, we tested whether pre-expressing the guide-donor would have the same effect as pre-expressing Cas9 when transformed with the opposite plasmid. Under the same transformation conditions, it was unexpectedly found that transformation of the Cas9 plasmid into cells carrying the guide-donor resulted in significantly higher numbers of edited colonies with similar or higher editing efficiencies ( Figure 5B ). Using inducible Cas9 produced similar improved results. In addition, we realized that if we included the cleavage site on the guide-donor plasmid in addition to the genome, we greatly improved the editing efficiency to the point where we had extremely high survival and simultaneous editing at both the Ade2 and REDI loci.

[0330] Finally, we are currently testing direct repair of the genomic integration cassette using the SceI meganuclease to cleave a DNA landing pad with a counter-selectable marker flanked by SceI sites within a region containing a promoter and terminator for guide RNA expression, which is then flanked by LexA-Fkh1 binding sites ( Fig.11 ). We speculate that this may result in similar levels of editing efficiency and allow direct genomic integration from an amplified oligonucleotide library followed by induction of Cas9 expression. This further has the advantage of ensuring only one edit per cell.

[0331] A major advantage of our system is that it utilizes REDI strain profiling technology and a platform for high-throughput precision editing, with guide-donor integration and active donor recruitment to RGN dsDNA break sites. This technology was previously applied to purify oligonucleotides (U.S. Patent Application Publication No. 20160122748, incorporated herein by reference in its entirety), but we have modified the technology to allow the formation of a functional strain collection. In particular, this enables us to profile individual edited strains, verify gRNA and donor sequences, and allow the isolation of perfect sequence guide-donors and equimolar collection variant strains (Figure 6). Another key aspect of our technology is that REDI-mediated strain profiling and rearrangement into arrays allow unambiguous confirmation of edited loci (Figure 7). This is impossible with any existing strategy that uses multiplex editing, and makes it feasible to test validation strains in separate wells for non-growth-based phenotyping, which is particularly important in many functional genomic applications (such as improving strains for the production of compounds, proteins or enzyme activity, analyzing protein locations, and verifying edited strains by multiplex whole genome sequencing). Our platform promises to revolutionize high-throughput genome editing, enabling more efficient, precise, and validated editing than any currently available technology or model system. The full workflow of our platform is detailed in Figure 8 .

[0332] To improve phenotyping in batch cultures, a system has also been developed to barcode editing cassettes. In this system, each editing cassette is associated with a random barcode ( Fig.12 These associations were then confirmed by paired-end sequencing of the barcoded guides and donors. The small barcodes could then be sequenced as proxies for the editing cassettes for phenotyping experiments, reducing phenotyping costs and making internal editing replicas feasible ( Fig.13 ).

[0333] discuss

[0334] High-throughput genetic engineering using RGNs coupled with arrayed synthetic oligonucleotides encoding guide RNA and donor DNA has great potential for a variety of applications 1 Current progress in this field has been rapid but is limited to generating large mutant libraries in libraries, which are not suitable for many phenotyping methods. By combining REDI with a novel high-throughput Cas9-based genome editing system using arrayed diffractive oligonucleotides, we have addressed this key limitation. Our method provides a simple mechanism to rapidly generate arrayed libraries of yeast variants. It can be applied to generate mutations anywhere in the yeast genome or in heterologous genes and pathways expressed in yeast hosts, and is particularly valuable for engineering strains for high-value chemical synthesis.

[0335] method

[0336] Oligonucleotide libraries were purchased from Agilent or Twist Biosciences. The basic oligonucleotide design is a sequence containing ~20nt specific sequence for CRISPR nucleases such as Cas9 or Cpf1, and a donor sequence containing the desired mutation (Figure 1). In addition, we can add additional synonymous mutations to be able to obtain amino acid changes outside the cleavage site without the need for PAM mutations (Figure 10). On either side of these modified sequences are ~30-90nt homologous sequences that match the genomic target.

[0337] The oligonucleotides are PCR amplified with primers that add additional sequences and subsequently ligated or assembled via Gibson Assembly into a plasmid containing a promoter to express the gRNA flanked by homologous sequences for the REDI integration locus ( Fig. 9 ). In addition, the method we developed allows for internal cloning of gRNA constant parts as well as selectable markers such as His3 or KanR2, which can select only cassettes that successfully incorporate the gRNA constant part and reduce background caused by synthesis errors or cloning errors (Figures 10 and 12).

[0338] The box encodes 2 edits, one modifies the genome and one integrates the box into the REDI locus (Figures 1 and 13). Cas9 and gRNA are expressed from different plasmids (Figure 1). In different iterations of the method, the gRNA plasmid or the Cas9 plasmid is transformed into the host (yeast), followed by the second transformation of the other plasmid (Figures 1 and 5). Both are expressed under a constitutive promoter. Alternatively, we can express under an inducible promoter, such as a galactose-inducible promoter or a tetracycline-inducible promoter. One of the two plasmids also contains SceI or other site-specific nuclease genes, under the control of an inducible or constitutive promoter. By selecting two plasmids, we ensure that the genome is edited by Cas9. The SceI gene can then be induced to integrate the gRNA-donor box at the REDI locus (Figure 6), in which anti-selective markers such as Fcy1 are deleted. Alternatively, we can achieve this with a second constant guide RNA, recruiting Cas9 to cut the REDI barcode locus that lacks Fcy1 and integrates the gRNA-donor-barcode box ( Fig.13 ), in a manner similar to SceI meganuclease cleavage. Selection can then be performed for successful integration of this cassette, which serves as a barcode for its encoded edit. These barcodes allow for the profiling of edited strains with REDI and for combined competitive growth experiments. When we perform REDI, only cassettes that perfectly encode the desired edit and have no other undesired edits are selected, making our method highly specific.

[0339] In addition to Cas9, our plasmids may also include enzymes such as Fkh1-LexA or Cas9-LexA to direct donor DNA to the DNA double-strand break site (Figure 4). This can greatly increase the editing survival rate and homologous recombination efficiency. After editing, the resulting edited cells can be analyzed in a manner similar to our previously reported REDI method (Figure 6). This additionally allows us to parse the edited cells into sub-libraries and verify that we have indeed prepared the desired edits, which is by sequencing specific regions where all edits for a single plate are expected to occur (Figure 7). If there is no edit at this position, we infer that the strain represents an edit that has not been generated and is removed from the collection.

[0340] Our integrated gRNA-donor cassette and / or its associated barcode can be used to track edited cells before or after REDI and editing confirmation. This allows high-throughput batch culture phenotyping. Additionally, we can phenotype strains on array plates by methods such as microscopy.

[0341] Example 2. Gene Editing Using the Cpf1-Donor System Produces Efficient Editing

[0342] When used in a similar manner as described in Example 1, the Cpf1 guide-donor system resulted in highly efficient (>99%) editing and editing was enhanced ˜10-fold with Cpf1, with the extent of donor recruitment similar to that of Cas9.

[0343] Fig.14A and 14B Provide data. Fig.14A Cell colonies pre-expressing Cpf1 are shown, transformed with a Cpf1 guide-donor plasmid targeting the ADE2 gene (the guide has a Cpf1 scaffold). The donor DNA encodes a mutation that causes a frameshift. Fig. 14B Shown are the % red colonies (ratio of red:white colonies) when the Cpf1 guide-donor and non-editing plasmid were mixed at a 17:3 ratio and transformed into cells expressing Cpf1 without (left) or with (right) LexA-FHA.

[0344] Example 3: Plasmid Spike-In experiments demonstrated that LexA-FHA and linearized vectors improve HDR efficiency and editing survival.

[0345] The plasmid for editing ADE2 ORF was mixed with the non-editing plasmid at 85% (17:3) and transformed into a cell line carrying Cas9 ( Fig.16 , top panel) or Cas9 and LexA-FHA ( Fig.16 Using the same strain for each transformation allowed for direct comparison of total colonies per row.

[0346] Fig.16Provide data. The y-axis indicates the total number of colonies observed in each transformation, and the x-axis indicates the percentage of red colonies, which represents the survival of the editing ADE2 process. The shape of each point corresponds to the restriction enzyme used to linearize the plasmid in vitro before transformation. 5 different columns correspond to different forms of spike-in mixtures. The first number corresponds to the number of genomic loci cuts by ADE2 editing plasmids (2 indicates cutting at ADE2 and chromosome barcode loci, and 1 indicates cutting only at ADE2), and the second number corresponds to the number of genomic loci cuts by non-editing plasmids (1 indicates cutting at the guide X recognition site (in this case, SceI site) of the chromosome barcode locus, and 0 indicates that there is no guide RNA on the non-editing plasmid). For example, 2v1 corresponds to a mixture, in which the ADE2 editing plasmid cuts the genome at the ADE2ORF and chromosome barcode locus, and the non-editing plasmid only cuts the chromosome barcode locus. In addition, the plasmid includes a SceI site, in which case the plasmid is cut by a SceI guide RNA that also targets the chromosome barcode locus, or does not include a SceI site, in which case the plasmid remains intact even if the SceI guide is expressed. The plasmid cut with SceI gRNA is linearized in vivo. Since these mixtures are prepared separately (although quantified by 85% mass), the most effective comparison of % red colonies can be prepared separately in each column. Editing toxicity in the absence of LexA-FHA or plasmid linearization results in little survival (samples use dotted circles, no enzyme-no reps 1 and 2). The maximum transformation survival occurs with LexA-FHA, without plasmid linearization (samples use dotted circles, no enzyme-no reps1 and 2).

[0347] These data show that plasmid linearization before transformation or the use of targeted fusion proteins (such as exA-FHA) can greatly improve editing efficiency relative to non-edited plasmids. In addition, this method does not require plasmid transformation, and it is also compatible with linear donor molecules because the barcode is captured at the barcode locus. In addition, in vivo linearization of the plasmid increases the ratio of properly edited cells to those with non-edited guides. Amplification of non-edited guide vectors is reduced, using linear donors, self-cutting donor plasmids or LexA-FHA. The total number of colonies that can be obtained is important for preparing complex libraries and is highest in the presence of donor recruitment proteins such as LexA-FHA.

[0348] Example 4. Donor DNA recruitment in human cells.

[0349] The donor recruitment technique described herein can also be used in mammalian cells. Applying the same concepts used in yeast, a protein that recruits to DNA double-strand breaks, TP53BP1, was selected. The normal role of TP53BP1 in cells is to bind double-strand breaks and promoter non-homologous end joining (NHEJ). The subdomain amino acids 1221-1718 of this protein was shown to act in a dominant negative manner for NHEJ (dn53BP1) (Xie et al., 2007). We hypothesized that this protein would be recruited to the break and, when fused to the LexA DNA binding domain, could be used to bring donor DNA to the break site, when the donor DNA contains a LexA site. In addition, since it can inhibit NHEJ, it may increase the rate of homology-directed repair (HDR), regardless of whether a LexA site is provided.

[0350] To test this, 2 versions of plasmids were generated that expressed the NLS, dn53BP1, fused to a C-terminal LexA DNA binding domain. One version expressed a gRNA for CACNA1D and the other for the gene PPP1R12C. The gRNAs for these sites were previously identified (Wang et al., 2018). A second plasmid expressing Cas9 and a third plasmid containing donor sequences for either CACNA1D or PPP1R12C (~300 nt of homology flanking either side of the XbaI site that would be introduced, with a small DNA segment missing that included the gRNA PAM sequence) were used. There were 2 versions of each donor plasmid, one with 4 LexA sites and one without a LexA site. Plasmids were built using Gibson Assembly. Both Cas9 and dn53BP1-LexA were expressed from the EF1α promoter.

[0351] Each of the 3 plasmids (25ng) was transiently transfected into Hek293 cells, inoculated one day before reaching a density of 10,000 cells per well in a 96-well plate, using X-tremeGENE 9 transfection reagent (Sigma Aldrich). Each group of conditions was tested in triplicate. Cells were grown for 72 hours after transfection, then harvested by removing the culture medium and washing the cells with water. Half of the cells were moved to a 96-well PCR plate, precipitated, and then DNA was extracted with 100 μl LucigenQuickExtract DNA extraction solution per sample.

[0352] The QuickExtract solution was then diluted as follows: 5 μl of QuickExtract was added to 20 μl of water. 2 μl was used to add to the PCR. 14 rounds of PCR were performed in a 25 μl Q5PCR mixture with internal primers that bind to the gene target of interest, and Read1 and Read2TruSeq primers (Illumina) were also added. One primer binds so far away from the editing site that it cannot be found in the homology region of the donor DNA provided, thereby amplifying only genomic DNA (rather than donor DNA). Another primer binds 32 or 33 nt away from the DNA sequence to be introduced by homology-mediated repair (HDR). This primer is used with Read1. After the first 14 rounds, an additional 25 μl Q5PCR mixture is added with primers for P5 and P7 adapters (Illumina) and sequencing indexes.

[0353] Samples were sequenced on an Illumina MiSeq to observe the distribution of edits at the cleavage site. Since the Read1 primer is closer to the cleavage site, Read1 was analyzed to determine the ratio of HDR and NHEJ. NHEJ was defined as a sequence containing an insertion or deletion within the gRNA recognition sequence or the target gene PAM sequence. HDR was defined as a sequence that mapped to the donor sequence.

[0354] result

[0355] Fig.17 Shown is the efficiency of HDR in the presence of the donor recruitment protein dn53BP1-LexA, with or without LexA sites. Targeting 2 independent genes (CACNA1D (CAC) and PPP1R12C (PPP)). The first panel shows the NHEJ ratio at the cleavage site. The second panel shows the total HDR ratio at the cleavage site, and the third panel shows the ratio of HDR to NHEJ in the cell.

[0356] It was found that the dn53BP1-LexA fusion can promote the HDR rate at the gRNA cut site when the LexA DNA site is present on the donor plasmid. When there is no LexA site on the donor plasmid, no increase in HDR is observed, but the NHEJ rate is similar. This suggests that DNA repair can generally be improved by using a fusion protein that is recruited to the break and contains a domain that binds the donor DNA and directs it to the break site.

[0357] References

[0358] 1. Garst AD, Bassalo MC, Pines G, Lynch SA, Halweg-Edwards AL, Liu R, et al. Genome-wide mapping of mutations at single-nucleotide resolution for protein, metabolic and genome engineering. Nat Biotechnol [Internet]. December 12, 2016; Available from: http: / / www.ncbi.nlm.nih.gov / pubmed / 27941803

[0359] 2. Jinek M, Chylinski K, Fonfara I, Hauer M, Doudna JA, Charpentier E. A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science. August 2012; 337(6096):816–21.

[0360] 3. Koike-Yusa H, Li Y, Tan E-P, Velasco-Herrera MDC, Yusa K. Genome-wide recessive genetic screening in mammalian cells with a lentiviral CRISPR-guide RNA library. Nat Biotechnol. March 2014; 32(3):267–73.

[0361] 4. Shalem O, Sanjana NE, Hartenian E, Shi X, Scott DA, Mikkelsen TS et al. Genome-scale CRISPR-Cas9 knockout screening in human cells. Science. January 2014; 343(6166):84–7.

[0362] 5. Wang T, Wei JJ, Sabatini DM, Lander ES. Genetic screens in human cells using the CRISPR-Cas9 system. Science. January 2014; 343(6166): 80–4.

[0363] 6. Zhou Y, Zhu S, Cai C, Yuan P, Li C, Huang Y, et al. High-throughput screening of a CRISPR / Cas9 library for functional genomics in human cells. Nature. May 2014; 509(7501): 487–91.

[0364] 7. Gilbert LA, Horlbeck MA, Adamson B, Villalta JE, Chen Y, Whitehead EH, et al. Genome-Scale CRISPR-Mediated Control of Gene Repression and Activation. Cell. October 2014; 159(3): 647–61.

[0365] 8. Konermann S, Brigham MD, Trevino AE, Joung J, Abudayyeh OO, Barcena C, et al. Genome-scale transcriptional activation by an engineered CRISPR-Cas9 complex. Nature. January 2015; 517(7536): 583–8.

[0366] 9. Ronda C, Maury J, T, Jacobsen SAB, Germann SM, Harrison SJ, et al. CrEdit: CRISPR mediated multi-loci gene integration in Saccharomyces cerevisiae. Microb Cell Fact. 2015; 14: 97.

[0367] 10. Ryan OW, Skerker JM, Maurer MJ, Li X, Tsai JC, Poddar S, et al. Selection of chromosomal DNA libraries using a multiplex CRISPR system. Elife. 2014;3.

[0368] 11. T, Bonde I, M, Harrison SJ, Kristensen M, Pedersen LE, et al. Multiplex metabolic pathway engineering using CRISPR / Cas9 in Saccharomyces cerevisiae. Metab Eng. March 2015;28:213–22.

[0369] 12. Bao Z, Xiao H, Liang J, Zhang L, Xiong X, Sun N, et al. Homology-integrated CRISPR-Cas (HI-CRISPR) system for one-step multigene disruption in Saccharomyces cerevisiae. ACS Synth Biol. March 2015;4(5):585–94.

[0370] 13. DiCarlo JE, Norville JE, Mali P, Rios X, Aach J, Church GM. Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic Acids Res. April 2013;41(7):4336–43.

[0371] 14. Ryan OW, Cate JHD. Multiplex engineering of industrial yeast genomes using CRISPRm. Methods Enzymol. 2014;546:473–89.

[0372] 15. Richardson CD, Ray GJ, DeWitt MA, Curie GL, Corn JE. Enhancing homology-directed genome editing by catalytically active and inactive CRISPR-Cas9 using asymmetric donor DNA. Nat Biotechnol. January 2016;

[0373] 16. Chu VT, Weber T, Wefers B, Wurst W, Sander S, Rajewsky K, et al. Increasing the efficiency of homology-directed repair for CRISPR-Cas9-induced precise gene editing in mammalian cells. Nat Biotechnol. May 2015; 33(5):543–8.

[0374] 17. Justin D. Smith, Ulrich Schlecht, Weihong Xu, Sundari Suresh, Joe Horecka, Michael J. Proctor, Raeka S. Aiyar, Richard A. O. Bennett, Angela Chu, Yong Fuga Li, Kevin Roy, Ronald W. Davis, Lars M. Steinmetz, Richard W. Hyman, Sasha F. Levy RPSO. High-throughput Parsing of Complex DNA Libraries for Isolation and Functional Characterization of Clonal, Sequence-verified DNA. Revis Mol Syst Biol.

[0375] 18.Wang,Y.,Liu,K.I.,Sutrisnoh,N.-A.B.,Srinivasan,H.,Zhang,J.,Li,J.,…Tan,M.H.(2018).Systematic evaluation of CRISPR-Cas systems reveals designprinciples for genome editing in human cells.Genome Biology,19(1),62.https: / / doi.org / 10.1186 / s13059-018-1445-x

[0376] 19.Xie,A.,Hartlerode,A.,Stucki,M.,Odate,S.,Puget,N.,Kwok,A.,…Scully,R.(2007).Distinct roles of chromatin-associated proteins MDC1 and 53BP1 inmammalian double-strand break repair.Molecular Cell,28(6),1045–1057.https: / / doi.org / 10.1016 / j.molcel.2007.12.005

[0377] Embodiment

[0378] Embodiment 1. A method for multiplex genetic modification and barcoding of cells, the method comprising: a) providing a plurality of recombinant polynucleotides, wherein each recombinant polynucleotide comprises a genome editing cassette comprising a guide RNA (gRNA) encoding polynucleotide and a donor polynucleotide, wherein the gRNA can hybridize at a genomic target locus to be modified, and the donor polynucleotide comprises a 5' homology arm that hybridizes to a 5' genomic target sequence and a 3' homology arm that hybridizes to a 3' genomic target sequence, wherein the homology arms flank a nucleotide sequence containing a desired edit to be integrated into the genomic target locus, wherein each recombinant polynucleotide comprises a different genome editing cassette comprising a different guide RNA-donor polynucleotide combination, such that the plurality of recombinant polynucleotides can produce a plurality of different desired edits at one or more genomic target loci; and (b) using the plurality of recombinant polynucleotides c) culturing the transfected cells under conditions suitable for transcription, wherein guide RNA is produced from each genome editing cassette; d) introducing an RNA-guided nuclease into the cells, wherein the RNA-guided nuclease forms a complex with the guide RNA produced in the cells, and the guide RNA directs the complex to one or more genomic target loci, wherein the RNA-guided nuclease generates a double-stranded break in the genomic DNA of the cells at the one or more genomic target loci, and the donor polynucleotide present in each cell is integrated at the genomic target locus identified by its 5' homology arm and 3' homology arm through homology-directed repair (HDR), thereby generating multiple genetically modified cells; and e) barcoding the multiple genetically modified cells by integrating the genome editing cassette present in each genetically modified cell at a chromosomal barcode locus.

[0379] Embodiment 2. A method as described in Embodiment 1, wherein each genome editing box also includes a promoter operably linked to the guide RNA encoding polynucleotide.

[0380] Embodiment 3. A method as described in Embodiment 1, wherein the chromosomal barcode locus further includes a promoter, which is operably linked to a polynucleotide encoding a guide RNA of any genome editing box that is integrated at the chromosomal barcode locus.

[0381] Embodiment 4. A method as described in embodiment 1, wherein each recombinant polynucleotide is provided by a vector.

[0382] Embodiment 5. A method as described in Embodiment 4, wherein the vector comprises a promoter operably linked to the guide RNA encoding polynucleotide.

[0383] Embodiment 6. A method as described in embodiment 5, wherein the promoter is a constitutive or inducible promoter.

[0384] Embodiment 7. The method of embodiment 4 further comprises replication of the vector within the transfected cells.

[0385] Embodiment 8. A method as described in Embodiment 4, wherein the vector is a plasmid or a viral vector.

[0386] Embodiment 9. The method of embodiment 4, wherein the vector is a high copy number vector.

[0387] Embodiment 10. A method as described in embodiment 1, wherein the RNA-guided nuclease is provided by a vector or a recombinant polynucleotide integrated into the cell genome.

[0388] Embodiment 11. A method as described in embodiment 10, wherein the genome editing box and RNA-guided nuclease are provided by a single vector or separate vectors.

[0389] Embodiment 12. A method as described in Embodiment 1, wherein the genome editing box also includes a tRNA sequence at the 5' end of the guide RNA encoding nucleotide sequence.

[0390] Embodiment 13. A method as described in Embodiment 1, wherein the genome editing box further includes a nucleotide sequence encoding a hepatitis delta virus (HDV) ribozyme at the 5' end of the nucleotide sequence encoding the guide RNA.

[0391] Embodiment 14. A method as described in embodiment 1, wherein the RNA-guided nuclease is a Cas nuclease or an engineered RNA-guided FokI nuclease.

[0392] Embodiment 15. A method as described in Embodiment 14, wherein the Cas nuclease is Cas9 or Cpf1.

[0393] Embodiment 16. A method as described in embodiment 1, wherein each of the donor polynucleotides introduces a different mutation into the genomic DNA.

[0394] Embodiment 17. A method as described in embodiment 16, wherein the mutation is selected from the group consisting of insertion, deletion and substitution.

[0395] Embodiment 18. A method as described in embodiment 16, wherein the at least one donor polynucleotide introduces a mutation inactivating a gene in the genomic DNA.

[0396] Embodiment 19. The method of embodiment 1, wherein the at least one donor polynucleotide removes a mutation from a gene in the genomic DNA.

[0397] Embodiment 20. A method as described in embodiment 1, wherein the multiple recombinant polynucleotides are capable of generating mutations at multiple sites within a single gene or non-coding region.

[0398] Embodiment 21. A method as described in embodiment 1, wherein the multiple recombinant polynucleotides are capable of generating mutations at multiple sites in different genes or non-coding regions.

[0399] Embodiment 22. A method as described in Embodiment 1, wherein the integration of the genome editing cassette present in each genetically modified cell at the chromosomal barcode locus is performed using HDR.

[0400] Embodiment 23. A method as described in Embodiment 22, wherein each recombinant polynucleotide also includes a pair of universal homology arms flanking the genome editing box, which can hybridize to complementary sequences at the chromosome barcode locus to allow the genome editing box to be integrated at the chromosome barcode locus through HDR.

[0401] Embodiment 24. A method as described in Embodiment 23, wherein each of the recombinant polynucleotides further includes a second guide RNA that can hybridize at the chromosomal barcode locus.

[0402] Embodiment 25. A method as described in embodiment 24, wherein the RNA-guided nuclease further forms a complex with a second guide RNA, and the second guide RNA guides the complex to the chromosome barcode locus, wherein the RNA-guided nuclease forms a double-stranded break at the chromosome barcode locus, and the genome editing box is integrated into the chromosome barcode locus through HDR.

[0403] Embodiment 26. A method as described in Embodiment 1, wherein the integration of the genome editing cassette present in each genetically modified cell at the chromosomal barcode locus is carried out using a site-specific recombinase system.

[0404] Embodiment 27. A method as described in Embodiment 26, wherein the site-specific recombinase system comprises a Cre-loxP site-specific recombinase system, a Flp-FRT site-specific recombinase system, a PhiC31-att site-specific recombinase system or a Dre-rox site-specific recombinase system.

[0405] Embodiment 28. A method as described in Embodiment 27, wherein the chromosomal barcode locus further includes a first recombination target site for a site-specific recombinase and the recombinant polynucleotide further includes a second recombination target site for the site-specific recombinase, and the site-specific recombination between the first recombination target site and the second site-specific recombination site causes the integration of the genomic editing box at the chromosomal barcode locus.

[0406] Embodiment 29. The method of embodiment 1, further comprising using a selectable marker to select clones that have undergone successful integration of the donor polynucleotide at the genomic target locus or successful integration of the genome editing cassette at the chromosomal barcode locus.

[0407] Embodiment 30. The method of embodiment 1, wherein the cell is a yeast cell.

[0408] Embodiment 31. The method of embodiment 1, wherein the yeast cell is a haploid yeast cell.

[0409] Embodiment 32. A method as described in Embodiment 1, wherein each recombinant polynucleotide further comprises a pair of restriction sites flanking the genome editing box.

[0410] Embodiment 33. A method as described in embodiment 32, wherein the restriction site is recognized by a meganuclease that produces DNA double-strand breaks.

[0411] Embodiment 34. A method as described in Embodiment 33, wherein the expression of the meganuclease is controlled by an inducible promoter.

[0412] Embodiment 35. A method as described in Embodiment 34, wherein the meganuclease is SceI.

[0413] Embodiment 36. The method of embodiment 1, further comprising performing additional rounds of genetic modification and genomic barcoding on the genetically modified cells by repeating steps (a)-(e) using different genome editing cassettes.

[0414] Embodiment 37. A method as described in embodiment 1, wherein each genome editing box also includes a unique barcode sequence for identifying the guide RNA and donor polynucleotide encoded by each genome editing box.

[0415] Embodiment 38. A method as described in Embodiment 37, further comprising sequencing each genome editing box.

[0416] Embodiment 39. A method as described in Embodiment 38, wherein the sequencing is performed before transfecting the cells.

[0417] Embodiment 40. A method as described in Embodiment 37, further comprising deleting the polynucleotide encoding the guide RNA and the donor polynucleotide at the chromosomal barcode locus where each genome editing box is integrated, while retaining the unique barcode at the chromosomal barcode locus.

[0418] Embodiment 41. A method as described in Embodiment 40, further comprising sequencing the barcode of the chromosomal barcode locus of at least one genetically modified cell to identify the genome editing box used to genetically modify the cell.

[0419] Embodiment 42. The method of embodiment 1, further comprising inhibiting non-homologous end joining (NHEJ).

[0420] Embodiment 43. A method as described in Embodiment 42, wherein the inhibition comprises contacting the cell with a small molecule inhibitor selected from wortmannin and Scr7.

[0421] Embodiment 44. A method as described in Embodiment 42, wherein the inhibition comprises using RNA interference or CRISPR interference to inhibit the expression of NHEJ pathway protein components.

[0422] Embodiment 45. The method of embodiment 1, further comprising using an HDR enhancer or active donor recruitment to increase the HDR frequency of cells.

[0423] Embodiment 46. The method of embodiment 1, further comprising using a selectable marker to select clones that have undergone successful integration of the donor polynucleotide at one or more genomic target loci via HDR.

[0424] Embodiment 47. The method of embodiment 1, further comprising phenotyping at least one genetically modified cell.

[0425] Embodiment 48. The method of embodiment 1, further comprising performing whole genome sequencing on at least one genetically modified cell.

[0426] Embodiment 49. The method as described in Embodiment 1 further includes sequence verification of the plurality of genetically modified cells and arranging them side by side into an array, the method comprising: a) placing a plurality of genetically modified cells in an ordered array in a culture medium suitable for the growth of genetically modified cells; b) culturing a plurality of genetically modified cells under certain conditions so that each genetically modified cell generates a clonal colony in the ordered array; c) introducing a genome editing cassette from the ordered array colony into the barcoded cells, wherein the barcoded cells include a nucleic acid comprising a recombination target site for a site-specific recombinase and a barcode sequence for identifying the colony position in the ordered array corresponding to the genome editing cassette; d) using a site-specific The recombinase system transposes the genome editing cassette to a position adjacent to the barcode sequence of the barcoded cell, wherein site-specific recombination with the barcoded cell recombination target site can produce a nucleic acid including the barcode sequence linked to the genome editing cassette; e) sequencing the nucleic acid containing the barcode sequence of the barcoded cell linked to the genome editing cassette to identify the guide RNA sequence and the donor polynucleotide sequence from the genome editing cassette in the colony, wherein the barcode sequence of the barcoded cell is used to identify the position of the colony in the ordered array from which the genome editing cassette originated; and f) picking the clone containing the genome editing cassette in the ordered array colony identified by the barcode of the barcoded cell.

[0427] Embodiment 50. The method of embodiment 49, wherein the genetically modified cell is a haploid yeast cell and the barcoded cell is a haploid yeast cell capable of conjugating with the genetically modified cell.

[0428] Embodiment 51. A method as described in embodiment 50, wherein the introduction of the genome editing cassette from the ordered array colony into the barcoded cell comprises joining the clone from the colony with the barcoded cell to generate a diploid yeast cell.

[0429] Embodiment 52. The method of embodiment 51, wherein the genetically modified cell is of strain MATα and the barcoded yeast cell is of strain MATa.

[0430] Embodiment 53. The method of embodiment 51, wherein the genetically modified cell is of strain MATa and the barcoded yeast cell is of strain MATα.

[0431] Embodiment 54. A method as described in Embodiment 49, wherein the genome editing box is flanked by restriction sites recognized by a large range of nucleases.

[0432] Embodiment 55. The method of embodiment 54, wherein the recombinase system in the barcoded cell uses a meganuclease to generate DNA double-strand breaks.

[0433] Embodiment 56. A method as described in embodiment 49, wherein the recombinase system in the barcoded cell is a Cre-loxP site-specific recombinase system, a Flp-FRT site-specific recombinase system, a PhiC31-att site-specific recombinase system or a Dre-rox site-specific recombinase system.

[0434] Embodiment 57. A method as described in Embodiment 49, wherein the method further comprises repeating c)-f) with all colonies in the ordered array to identify the sequence of the genome editing box guide RNA and the donor polynucleotide for each colony in the ordered array.

[0435] Embodiment 58. An ordered array of colonies comprising genetically modified cell clones produced by the method of embodiment 49, wherein the colonies are indexed according to the verified sequences of their guide RNA and donor polynucleotides.

[0436] Embodiment 59. A method for promoting homology-directed repair (HDR) by actively recruiting donors to DNA breaks, the method comprising: a) introducing into a cell a fusion protein comprising a protein that selectively binds to DNA breaks linked to a polypeptide comprising a nucleic acid binding domain; and b) introducing into the cell a donor polynucleotide comprising i) a nucleotide sequence that is sufficiently complementary to hybridize to a sequence adjacent to the DNA break and ii) a nucleotide sequence comprising a binding site recognized by the nucleic acid binding domain of the fusion protein, wherein the nucleic acid binding domain selectively binds to the binding site on the donor polynucleotide to produce a complex of the donor polynucleotide and the fusion protein, thereby recruiting the donor polynucleotide to the DNA break and promoting HDR.

[0437] Embodiment 60. A method as described in Embodiment 59, wherein the protein recruited to the DNA break is an RNA-guided nuclease.

[0438] Embodiment 61. A method as described in Embodiment 59, wherein the RNA-guided nuclease is a Cas nuclease or an engineered RNA-guided FokI nuclease.

[0439] Embodiment 62. A method as described in Embodiment 61, wherein the Cas nuclease is Cas9 or Cpf1.

[0440] Embodiment 63. A method as described in Embodiment 59, wherein the DNA break is a single-stranded or double-stranded DNA break.

[0441] Embodiment 64. A method as described in Embodiment 63, wherein the fusion protein includes a protein that selectively binds to single-stranded DNA breaks or double-stranded DNA breaks.

[0442] Embodiment 65. A method as described in Embodiment 59, wherein the donor polynucleotide is single-stranded or double-stranded.

[0443] Embodiment 66. A method as described in Embodiment 59, wherein the nucleic acid binding domain is an RNA binding domain and the binding site comprises an RNA sequence recognized by the RNA binding domain.

[0444] Embodiment 67. A method as described in Embodiment 59, wherein the nucleic acid binding domain is a DNA binding domain and the binding site comprises a DNA sequence recognized by the DNA binding domain.

[0445] Embodiment 68. A method as described in Embodiment 67, wherein the DNA binding domain is a LexA DNA binding domain and the binding site is a LexA binding site.

[0446] Embodiment 69. A method as described in embodiment 67, wherein the DNA binding domain is a forkhead homolog 1 (FKH1) DNA binding domain and the binding site is a FKH1 binding site.

[0447] Embodiment 70. A method as described in Embodiment 59, wherein the polypeptide containing a nucleic acid binding domain also includes a forkhead-associated (FHA) phosphothreonine binding domain, wherein the donor polynucleotide is selectively recruited to a DNA break having a protein that contains a phosphorylated threonine residue that is located sufficiently close to the DNA break for the FHA phosphothreonine binding domain to bind to the phosphorylated threonine residue.

[0448] Embodiment 71. A method as described in Embodiment 59, wherein the nucleic acid binding domain-containing polypeptide comprises a LexA DNA binding domain connected to a FHA phosphothreonine binding domain.

[0449] Embodiment 72. A method as described in Embodiment 59, wherein the donor polynucleotide is provided by a recombinant polynucleotide, which recombinant polynucleotide includes a promoter operably linked to the donor polynucleotide.

[0450] Embodiment 73. A method as described in Embodiment 59, wherein the fusion protein is provided by a recombinant polynucleotide, which recombinant polynucleotide includes a promoter operably linked to the fusion protein encoding polynucleotide.

[0451] Embodiment 74. A method as described in Embodiment 59, wherein the donor polynucleotide and fusion protein are provided by a single vector or separate vectors.

[0452] Embodiment 75. The method of embodiment 74, wherein at least one vector is a viral vector or a plasmid.

[0453] Embodiment 76. A method as described in Embodiment 50, wherein the donor polynucleotide is RNA or DNA.

[0454] Embodiment 77. A method as described in Embodiment 76, wherein the method further comprises reverse transcribing the RNA-containing donor polynucleotide with a reverse transcriptase to produce a DNA-containing donor polynucleotide.

[0455] Embodiment 78. A method as described in Embodiment 59, wherein the DNA breaks are produced by site-specific nucleases.

[0456] Embodiment 79. A method as described in Embodiment 78, wherein the site-specific nuclease is selected from Cas nuclease, engineered RNA-guided FokI nuclease, large-range nuclease, zinc finger nuclease (ZFN) and transcription activator-like effector nuclease (TALEN).

[0457] Embodiment 80. A kit for multiplex genetic modification and barcoding of cells, the kit comprising: a) a plurality of recombinant polynucleotides, wherein each recombinant polynucleotide comprises a genome editing cassette comprising a guide RNA (gRNA) encoding polynucleotide and a donor polynucleotide, wherein the gRNA is capable of hybridizing at a genomic target locus to be modified, the donor polynucleotide comprising a 5' homology arm that hybridizes to a 5' genomic target sequence and a 3' homology arm that hybridizes to a 3' genomic target sequence, wherein the homology arms flank a nucleotide sequence containing a desired edit to be integrated into the genomic target locus, wherein each recombinant polynucleotide comprises a different genome editing cassette comprising a different guide RNA-donor polynucleotide combination, such that the plurality of recombinant polynucleotides are capable of generating a plurality of different desired edits at one or more genomic target loci; and b) an RNA-guided nuclease; and c) a cell comprising a chromosomal barcode locus, wherein the barcode locus comprises an integration site for at least one recombinant polynucleotide genome editing cassette.

[0458] Embodiment 81. A kit as described in Embodiment 80, wherein each recombinant polynucleotide further comprises a pair of universal homology arms flanking the genome editing box, which can hybridize complementary sequences at the integration site of the chromosomal barcode locus to allow integration of the genome editing box at the chromosomal barcode locus by homology-directed repair (HDR).

[0459] Embodiment 82. A kit as described in Embodiment 81, wherein each of the recombinant polynucleotides further comprises a second guide RNA capable of hybridizing at the chromosomal barcode locus.

[0460] Embodiment 83. The kit of embodiment 80, further comprising a site-specific recombinase system.

[0461] Embodiment 84. A kit as described in embodiment 83, wherein the site-specific recombinase system is a Cre-loxP site-specific recombinase system, a Flp-FRT site-specific recombinase system, a PhiC31-att site-specific recombinase system or a Dre-rox site-specific recombinase system.

[0462] Embodiment 85. A kit as described in Embodiment 83, wherein the chromosomal barcode locus further includes a first recombination target site for a site-specific recombinase and the recombinant polynucleotide further includes a second recombination target site for the site-specific recombinase, so that site-specific recombination can occur between the first recombination target site and the second site-specific recombination site to allow the genome editing box to be integrated at the chromosomal barcode locus.

[0463] Embodiment 86. A kit as described in Embodiment 80, wherein the RNA-guided nuclease is a Cas nuclease or an engineered RNA-guided FokI nuclease.

[0464] Embodiment 87. A kit as described in Embodiment 86, wherein the Cas nuclease is Cas9 or Cpf1.

[0465] Embodiment 88. A kit as described in Embodiment 80, wherein the kit further comprises a fusion protein comprising a polypeptide containing a nucleic acid binding domain linked to a protein that selectively binds to DNA fragments generated by RNA-guided nucleases.

[0466] Embodiment 89. A kit as described in embodiment 88, wherein the donor polynucleotide also includes a nucleotide sequence that is sufficiently complementary to the sequence adjacent to the hybridized DNA break, and a nucleotide sequence that contains a binding site recognized by the nucleic acid binding domain of the fusion protein.

[0467] Embodiment 90. The kit of embodiment 89, wherein the nucleic acid binding domain is a LexA DNA binding domain and the binding site is a LexA binding site, or the nucleic acid binding domain is a forkhead homolog 1 (FKH1) DNA binding domain and the binding site is a FKH1 binding site.

[0468] Embodiment 91. A kit as described in Embodiment 90, wherein the polypeptide containing a nucleic acid binding domain also includes a forkhead-associated (FHA) phosphothreonine binding domain.

[0469] Embodiment 92. The kit of Embodiment 91, wherein the nucleic acid binding domain-containing polypeptide comprises a LexA DNA binding domain linked to a FHA phosphothreonine binding domain.

[0470] While preferred embodiments of the present disclosure have been illustrated and described, it will be appreciated that various changes can be made therein without departing from the spirit and scope of the disclosure.

Claims

1. A method for localizing a donor polynucleotide to a genomic target locus in a target cell, the method comprising: (a) transfecting a target cell with a first recombinant polynucleotide, wherein the first recombinant polynucleotide comprises a genome editing cassette, wherein the genome editing cassette comprises: (i) a promoter operably linked to a nucleic acid sequence encoding a guide RNA (gRNA) that can hybridize to a target locus in the genome to be modified, and (ii) a donor polynucleotide comprising a sequence to be integrated into the genomic target locus; (b) transfecting the target cell with a second recombinant polynucleotide, wherein the second recombinant polynucleotide encodes the first polypeptide and encodes the second polypeptide, wherein (i) the first polypeptide is an RNA-guided nuclease, wherein when complexed with the gRNA, the RNA-guided nuclease recognizes the genomic target locus and generates a DNA break site at the genomic target locus; and (ii) the second polypeptide is a donor recruitment protein, wherein the donor recruitment protein comprises: (1) a DNA binding domain that binds to the donor polynucleotide, and (2) a DNA break site localization domain that binds to a region near the DNA break at the genomic target locus; wherein when bound to both the region near the DNA break and the donor polynucleotide, the donor recruitment protein selectively recruits the donor polynucleotide to the DNA break site, thereby positioning the donor polynucleotide to the genomic target locus, wherein the DNA binding domain comprises a LexA DNA binding domain, and the DNA break site localization domain comprises a forkhead-associated (FHA) phosphothreonine binding domain or comprises a TP53BP1 domain, The method described is a non-therapeutic method.

2. The method of claim 1, wherein the DNA breaks are double-stranded DNA breaks.

3. The method of claim 1 or 2, wherein the donor recruitment protein is a fusion protein.

4. The method of claim 1 or 2, wherein the DNA break site localization domain comprises a polypeptide sequence from a protein that binds to a DNA break site, or to a region near a DNA break site caused by DNA breakage.

5. The method according to claim 4, wherein the protein that binds to the DNA break site or to the region near the DNA break site caused by DNA breakage is a protein involved in DNA repair.

6. The method of claim 1 or 2, wherein the donor polynucleotide is further integrated into a genomic target locus, thereby generating a genetically engineered cell.

7. The method of claim 6, wherein the genetically engineered cells are genetically engineered therapeutic cells.

8. The method of claim 7, wherein the genetically engineered therapeutic cells are genetically engineered immune cells.

9. The method of claim 8, wherein the genetically engineered immune cells are cancer-targeting T cells or natural killer cells.

10. The method of claim 1 or 2, wherein the DNA break site localization domain binds to genomic DNA or a protein, and the protein is specifically associated with the genomic DNA.

11. The method of claim 1, wherein the donor recruitment protein comprises a LexA DNA binding domain linked to a DNA break site localization domain, the DNA break site localization domain comprising a FHA phosphothreonine binding domain.

12. A composition comprising a target cell, a first polypeptide or a polynucleotide encoding the first polypeptide, a gene editing vector, and a second polypeptide or a polynucleotide encoding the second polypeptide, The gene editing vector comprises a recombinant polynucleotide, the recombinant polynucleotide comprises a genome editing cassette, and the genome editing cassette comprises: (i) a promoter operably linked to a nucleic acid sequence encoding a guide RNA (gRNA) that can hybridize to a target locus in the genome to be modified, and (ii) a donor polynucleotide comprising a sequence to be integrated into the genomic target locus; wherein the first polypeptide is an RNA-guided nuclease, wherein when complexed with the gRNA, the RNA-guided nuclease recognizes the genomic target locus and generates a DNA break site at the genomic target locus; Wherein the second polypeptide is a donor recruitment protein, wherein the donor recruitment protein comprises: (1) a DNA binding domain that binds to the donor polynucleotide, and (2) a DNA break site localization domain that binds to a region near the DNA break obtained at the genomic target locus; wherein when bound to both the region near the DNA break and the donor polynucleotide, the donor recruitment protein selectively recruits the donor polynucleotide to the DNA break site, thereby positioning the donor polynucleotide to the genomic target locus, The DNA binding domain comprises a LexA DNA binding domain, and the DNA break site localization domain comprises a forkhead-associated (FHA) phosphothreonine binding domain or comprises a TP53BP1 domain.

13. The composition of claim 12, wherein the target cell is a cell from a subject.

14. The composition of claim 13, wherein the subject has cancer.

15. The composition of any one of claims 12-14, wherein the target cell is an immune cell.

16. The composition of claim 15, wherein the immune cell is a T cell.

17. The composition of any one of claims 12-14, wherein the donor polynucleotide encodes a therapeutic agent.

18. The composition of claim 17, wherein the therapeutic agent is a chimeric antigen receptor or a T cell receptor.

19. The composition of claim 13, wherein the subject suffers from a disease that can be treated by integrating the donor DNA into the cellular genome.

20. The composition of any one of claims 12-14, wherein the cell is a human cell.

21. The composition of any one of claims 12-14, wherein the DNA break site localization domain binds to a DNA or a protein, the DNA comprising the DNA break site, and the protein specifically associates with the DNA comprising the DNA break site.

22. The composition of claim 12, wherein the donor recruitment protein comprises a LexA DNA binding domain linked to a DNA break site localization domain, the DNA break site localization domain comprising a FHA phosphothreonine binding domain.

23. A kit comprising: (a) Gene editing vectors; (b) a first polypeptide or a polynucleotide encoding the first polypeptide, and (c) a second polypeptide or a polynucleotide encoding the second polypeptide, The gene editing vector comprises a recombinant polynucleotide, the recombinant polynucleotide comprises a genome editing cassette, and the genome editing cassette comprises: (i) a promoter operably linked to a nucleic acid sequence encoding a guide RNA (gRNA) that can hybridize to a target locus in the genome to be modified, and (ii) a donor polynucleotide comprising a sequence to be integrated into the genomic target locus; wherein the first polypeptide is an RNA-guided nuclease, wherein when complexed with the gRNA, the RNA-guided nuclease recognizes the genomic target locus and generates a DNA break site at the genomic target locus; Wherein the second polypeptide is a donor recruitment protein, wherein the donor recruitment protein comprises: (1) a DNA binding domain that binds to the donor polynucleotide, and (2) a DNA break site localization domain that binds to a region near the DNA break at the genomic target locus; wherein when bound to both the region near the DNA break and the donor polynucleotide, the donor recruitment protein selectively recruits the donor polynucleotide to the DNA break site, thereby positioning the donor polynucleotide to the genomic target locus, The DNA binding domain comprises a LexA DNA binding domain, and the DNA break site localization domain comprises a forkhead-associated (FHA) phosphothreonine binding domain or comprises a TP53BP1 domain.

24. The kit of claim 23, further comprising (d) a cell engineered to express a nuclease.

25. The kit of claim 23, wherein the donor recruitment protein comprises a LexA DNA binding domain linked to a DNA break site localization domain, the DNA break site localization domain comprising a FHA phosphothreonine binding domain.

26. The method of claim 1 or 2, the composition of any one of claims 12-14, or the kit of any one of claims 23-25, wherein more than one donor recruitment protein binds to the DNA break or the vicinity of the DNA break simultaneously.

Citation Information

Patent Citations

  • Lipid-protein-sugar particles for delivery of nucleic acids

    US20020150626A1

  • Lipid-mediated polynucleotide administration to deliver a biologically active peptide and to induce a cellular immune response

    US20030032615A1

  • Lipid-comprising drug delivery complexes and methods for their production

    US20030203865A1

  • Lyophilizable and enhanced compacted nucleic acids

    US20040048787A1

  • Methods and apparatus for measuring analytes

    US20100137143A1