Method for introducing large chromosomes, modified chromosomes, and organisms using the same

JP2024533683A5Pending Publication Date: 2025-10-01IMMUNOCAN BIOTECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024518502
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-09-24
Filing Date
2022-09-23
Publication Date
2025-10-01

AI Technical Summary

Technical Problem

The manipulation of large fragments of genes and chromosomes is a significant challenge in gene editing, particularly due to the limitations of existing delivery vectors in carrying large chromosome fragments and the inefficiency of replacing large interchromosomal sequences.

Method used

The use of Massive Fragment Across Species In situ Replacement Technology (MASIRT) for introducing large sequence fragments between chromosomes, involving the generation of engineered chromosomes through homology-directed repair using CRISPR/Cas endonucleases and guide nucleic acids, enabling the insertion of template sequences into target chromosomes.

Benefits of technology

This method allows for efficient replacement of large chromosomal sequences, such as human genes in model organisms, reducing time and cost while increasing the efficiency of creating animals with humanized genes, and facilitating the study of diseases and disorders.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure relates to methods for introducing large sequence fragments between chromosomes and generating chromosomal rearrangements using the double-strand break repair pathway and homology directed repair. The disclosure further relates to the chromosomes produced by these methods, and to cells and transgenic animals containing these chromosomes.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a method for introducing a large chromosome, a modified chromosome, and an organism using the same.

[0002] INCORPORATION BY REFERENCE OF SEQUENCE LISTING This application contains a Sequence Listing which has been submitted in ASCII format via EFS-WEB and is hereby incorporated by reference in its entirety.

[0003] background Manipulation of large fragments of genes and chromosomes is a powerful tool for basic and translational research, as well as for the development of therapeutics. Human genes range in size from a few hundred bases to at least 2,300 kilobases (KB), and human chromosomes range in size from 38 megabase pairs (MB) to nearly 250 MB. Thus, to effectively study large genes, regions spanning multiple genes, and parts of chromosomes, large sequence fragments must be manipulated. However, the manipulation of large fragments is one of the most important challenges in the field of gene editing. The present disclosure provides a method for manipulating large sequences. Summary of the Invention

[0004] overview The present disclosure provides a method of generating an engineered chromosome, comprising: (a) providing a cell comprising a target chromosome comprising a target sequence and a template chromosome comprising a template sequence; (b) contacting the cell with (i) a first nucleic acid molecule comprising a 5' homology arm comprising, from 5' to 3', a 5' homology arm comprising a nucleotide sequence upstream of the 5' end of the target sequence, at least a first marker, and a 3' homology arm comprising a nucleotide sequence upstream of the 5' end of the template sequence, and (ii) a second nucleic acid molecule comprising a 5' homology arm comprising, from 5' to 3', a 5' homology arm comprising a nucleotide sequence downstream of the 3' end of the template sequence, at least a second marker, and a 3' homology arm comprising a nucleotide sequence downstream of the 3' end of the target sequence; (c) generating double-stranded breaks on either side of the target sequence and at the 5' and 3' ends of the template sequence, whereby the template sequence and the first and second markers are inserted into the target chromosome; and (d) selecting a cell or a plurality of cells expressing the first and second markers.

[0005] In some embodiments, the first marker is located at the 5' end of the template sequence and the second marker is located at the 3' end of the template sequence after insertion of the template sequence.

[0006] In some embodiments, the length of the 5' and 3' homology arms of the first and second nucleic acid molecules is between about 20 base pairs (bp) and 2,000 base pairs (bp), between about 50 bp and 1,500 bp, between about 100 bp and 1,400 bp, between about 150 bp and 1,300 bp, between about 200 bp and 1,200 bp, between about 300 bp and 1,100 bp, between about 400 bp and 1,000 bp, or between about 500 bp and 900 bp, or between about 600 bp and 800 bp. In some embodiments, the length of the 5' and 3' homology arms of the first and second nucleic acid molecules is between about 400 bp and 1,500 bp, between about 500 bp and 1,300 bp, or between about 600 bp and 1,000 bp. In some embodiments, the length of the 5' and 3' homology arms of the first and second nucleic acid molecules are between about 600 bp and 1,000 bp.

[0007] In some embodiments, the length of the template sequence is at least 25 kilobase pairs (KB), at least 50 KB, at least 100 KB, at least 200 KB, at least 400 KB, at least 500 KB, at least 600 KB, at least 700 KB, at least 800 KB, at least 900 KB, at least 1 megabase pair (MB), at least 2 MB, at least 3 MB, at least 4 MB, at least 5 MB, at least 6 MB, at least 7 MB, at least 8 MB, at least 9 MB, at least 10 MB, at least 15 MB, at least 20 MB, at least 25 MB, at least 30 MB, at least 40 MB, at least 50 MB, at least 60 MB, at least 70 MB, at least 80 MB, at least 90 MB, at least 100 MB, at least 120 MB, at least 140 MB, at least 160 MB, at least 180 MB, at least 200 MB, at least 220 MB, or at least 250 MB.In some embodiments, the length of the template sequence is between 50KB and 250MB, 50KB and 100MB, 50KB and 50MB, 50KB and 20MB, 50KB and 10MB, 50KB and 5MB, 50KB and 3MB, 50KB and 2MB, 50KB and 1MB, 100KB and 200MB, 100KB and 100MB, 100KB and 50MB, 100KB and 20MB, 100KB and 10MB, 100KB and 5MB, 100KB and 3MB, 100KB and 2MB, 100KB and 1MB, 100KB and 500KB, 200KB and 100MB, 200KB and 50MB, 200KB and 20MB, 200KB and 10MB, 200KB and 5MB, 200KB and 3MB, The file size is between 200KB and 2MB, 200KB and 1MB, 200KB and 500KB, 500KB and 100MB, 500KB and 50MB, 500KB and 20MB, 500KB and 10MB, 500KB and 5MB, 500KB and 3MB, 500KB and 2MB, 500KB and 1MB, 1MB and 100MB, 1MB and 50MB, 1MB and 20MB, 1MB and 10MB, 1MB and 5MB, 1MB and 3MB, 1MB and 2MB, 3MB and 100MB, 3MB and 50MB, 3MB and 20MB, 3MB and 10MB, 3 MB and 5MB, 5MB and 100MB, 5MB and 50MB, 5MB and 20MB, 5MB and 10MB, 10MB and 100MB, 10MB and 50MB, or 10MB and 20MB. In some embodiments, the template sequence is between 200KB and 50MB, between 1MB and 20MB, between 1MB and 10MB, between 1MB and 5MB, between 1MB and 3MB, between 3MB and 20MB, between 3MB and 10MB, between 3MB and 7MB, or between 3MB and 5MB in length.

[0008] In some embodiments, generating the double-stranded break in (c) comprises inducing the double-stranded break using a CRISPR / Cas endonuclease and one or more guide nucleic acids (gNAs), one or more zinc finger nucleases, one or more transcription activator-like effector nucleases (TALENs), or one or more CRE recombinases. In some embodiments, the CRISPR / Cas endonuclease is selected from the group consisting of CasI, CasIB, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, CasX, CasY, Cas12a (Cpf1), Cas12b, Cas13a, CsyI, Csy2, Csy3, CseI, Cse2, CscI, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, CmrI, Cmr3, Cmr4, Cmr5, Cmr6, CsbI, Csb2, Csb3, Csx17, CsxI4, Csx10, Csx16, CsaX, Csx3, In some embodiments, the CRISPR / Cas endonuclease comprises Cas9, Cpf1 (Cas12a), Cas12b, CasX, CasY, C2c1, or C2c3, or a homolog, ortholog, or modified version thereof. In some embodiments, the CRISPR / Cas endonuclease comprises Cas9. In some embodiments, the gNA comprises a single guide RNA (sgRNA).

[0009] In some embodiments, the target chromosome comprises, from 5' to 3', the sequence of the 5' homology arm of a first nucleic acid molecule, the target sequence, and the sequence of the 3' homology arm of a second nucleic acid molecule. In some embodiments, the template chromosome comprises, from 5' to 3', the sequence of the 3' homology arm of a first nucleic acid molecule, the template sequence, and the sequence of the 5' homology arm of a second nucleic acid molecule.

[0010] In some embodiments, the target sequence comprises at least 1 gene, at least 2 genes, at least 3 genes, at least 5 genes, at least 10 genes, at least 20 genes, at least 30 genes, at least 40 genes, at least 50 genes, at least 100 genes, or at least 200 genes. In some embodiments, the target sequence comprises one or more genes that are homologous to one or more genes of the template sequence.

[0011] In some embodiments, the template sequence comprises a naturally occurring sequence. In some embodiments, the template sequence comprises at least 1 gene, at least 2 genes, at least 3 genes, at least 5 genes, at least 10 genes, at least 20 genes, at least 30 genes, at least 40 genes, at least 50 genes, at least 100 genes, or at least 200 genes. In some embodiments, the template sequence comprises one or more modifications to a naturally occurring sequence. In some embodiments, the template sequence comprises an artificial sequence. In some embodiments, the artificial sequence comprises a sequence encoding one or more antibodies or antigen-binding fragments thereof. In some embodiments, the one or more antibodies or antigen-binding fragments thereof comprise scFvs, bispecific antibodies, or multispecific antibodies.

[0012] In some embodiments, the target sequence is deleted by inserting the template sequence. In some embodiments, (a) the target chromosome comprises, from 5' to 3', the sequence of the 5' homology arm of the first nucleic acid molecule, the first sgRNA target sequence, the target sequence, the second sgRNA target sequence, and the sequence of the 3' homology arm of the second nucleic acid molecule; and (b) the template chromosome comprises, from 5' to 3', the third sgRNA target sequence, the sequence of the 3' homology arm of the first nucleic acid molecule, the template sequence, the sequence of the 5' homology arm of the second nucleic acid molecule, and the fourth sgRNA target sequence. In some embodiments, generating the double-stranded break comprises contacting the cell with a CRISPR / Cas endonuclease and the first, second, third, and fourth sgRNA. In some embodiments, the first, second, third, and fourth sgRNA comprise targeting sequences specific for the first, second, third, and fourth sgRNA target sequences.

[0013] In some embodiments, contacting the cell with the CRISPR / Cas endonuclease and the sgRNA comprises transfecting the cell with one or more nucleic acid molecules encoding the CRISPR / Cas endonuclease and the sgRNA.

[0014] In some embodiments, the insertion of the template sequence includes little or no deletion of the sequence of the target sequence. In some embodiments, the insertion of the template sequence disrupts one or more functions of the target sequence. In some embodiments, the insertion of the template sequence disrupts the gene of the target sequence. In some embodiments, (a) the target chromosome includes, from 5' to 3', the sequence of the 5' homology arm of the first nucleic acid molecule, the first sgRNA target sequence, and the sequence of the 3' homology arm of the second nucleic acid molecule; and (b) the template chromosome includes, from 5' to 3', the second sgRNA target sequence, the sequence of the 3' homology arm of the first nucleic acid molecule, the template sequence, the sequence of the 5' homology arm of the second nucleic acid molecule, and the third sgRNA target sequence. In some embodiments, the generation of the double-stranded break includes contacting the cell with a CRISPR / Cas endonuclease and the first, second, and third sgRNA. In some embodiments, the first, second, and third sgRNAs comprise targeting sequences specific for the first, second, and third sgRNA target sequences. In some embodiments, contacting the cell with the CRISPR / Cas endonuclease and the sgRNA comprises transfecting the cell with one or more nucleic acid molecules encoding the CRISPR / Cas endonuclease and the sgRNA.

[0015] In some embodiments, the first or second marker comprises a fluorescent protein operably linked to a promoter capable of expressing the fluorescent protein in the cell. In some embodiments, the fluorescent protein comprises green fluorescent protein (GFP), yellow fluorescent protein (YFP), red fluorescent protein (RFP), cyan fluorescent protein (CFP), blue fluorescent protein (BFP), dsRed, mCherry, or tdTomato. In some embodiments, the fluorescent protein comprises GFP. In some embodiments, the first marker further comprises a selectable marker. In some embodiments, the second marker further comprises a selectable marker. In some embodiments, the selectable marker is selected from the group consisting of dihydrofolate reductase (DHFR), glutamine synthetase (GS), puromycin acetyltransferase, blasticidin deaminase, histidinol dehydrogenase, hygromycin phosphotransferase (hph), bleomycin resistance gene, and aminoglycoside phosphotransferase (neomycin resistance). In some embodiments, the first marker and the second marker are not the same selectable marker. In some embodiments, the first marker comprises GFP and puromycin acetyltransferase operably linked to a promoter capable of expressing GFP in the cell, and the second marker comprises hygromycin phosphotransferase.

[0016] In some embodiments, the method further comprises (e) deleting all or part of the first or second marker after step (d). In some embodiments, deleting the first or second marker comprises inducing the deletion using a gNA comprising a CRISPR / Cas endonuclease and a targeting sequence specific to the sequence encoding the marker.

[0017] In some embodiments, the cells comprise hybrid cells, embryonic hybrid stem (EHS) cells, or zygotes. In some embodiments, the EHS cells are generated by fusing ES cells of any two species selected from the group consisting of mouse, rat, rabbit, guinea pig, hamster, sheep, goat, donkey, cow, horse, camel, chicken, and monkey. In some embodiments, the EHS cells are generated by fusing a human embryonic stem cell with an embryonic stem cell from a non-human species. In some embodiments, the non-human species is mouse, rat, rabbit, guinea pig, hamster, sheep, goat, donkey, cow, horse, camel, chicken, or monkey. In some embodiments, the EHS cells are generated by fusing ES cells from any two different species selected from the group consisting of mouse, rat, rabbit, guinea pig, hamster, sheep, goat, donkey, cow, horse, camel, chicken, and monkey. In some embodiments, the fusion comprises electrofusion, virus-induced fusion, or chemically-induced fusion.

[0018] In some embodiments, the cell comprises a hybrid cell. In some embodiments, generating the hybrid cell comprises: (a) generating a micronucleated human cell; and (b) fusing the micronucleated human cell with a cell from a non-human species, thereby generating a hybrid cell. In some embodiments, the micronucleated human cell is generated by exposing a human cell to colcemid under conditions sufficient to induce micronucleation and recovering the micronucleated cell using centrifugation. In some embodiments, the non-human species is a mouse, rat, rabbit, guinea pig, hamster, sheep, goat, donkey, cow, horse, camel, chicken, or monkey. In some embodiments, the cell from the non-human species is an ES cell and the hybrid cell is an EHS cell.

[0019] In some embodiments, the target sequence comprises a gene encoding an immunoglobulin or a T cell receptor subunit. In some embodiments, the target chromosome comprises mouse chromosome 12 and the template chromosome comprises human chromosome 14. In some embodiments, the target sequence comprises a mouse Igh variable region sequence. In some embodiments, the mouse Igh variable region sequence comprises sequences encoding mouse VH, DH and JH1-6 gene segments and intervening non-coding sequences. In some embodiments, the template sequence comprises a human IGH variable region sequence. In some embodiments, the human IGH variable region sequence comprises sequences encoding human VH, DH and JH1-6 gene segments and intervening non-coding sequences. In some embodiments, the target sequence comprises a mouse Igl variable region sequence. In some embodiments, the target sequence comprises a mouse Igk variable region sequence. In some embodiments, the template sequence comprises a human IGL variable region sequence. In some embodiments, the template sequence comprises a human IGK variable region sequence. In some embodiments, the mouse Igk variable region sequence comprises a mouse V k , and J k1-5 The template sequence includes sequences encoding gene segments, as well as intervening non-coding sequences. In some embodiments, the template sequence includes a human IGK variable region sequence. In some embodiments, the human IGK variable region sequence is a human V k , and J k1-5 A gene segment includes sequences that code for the gene, as well as intervening non-coding sequences.

[0020] In some embodiments, the method further comprises recovering recombinant chromosomes from the cells selected in step (d). In some embodiments, recovering the recombinant chromosomes comprises exposing the cells to colcemid under conditions sufficient to induce micronucleus formation and recovering the micronucleated cells using centrifugation.

[0021] In some embodiments, the first and second nucleic acid molecules are plasmids.

[0022] The present disclosure provides recombinant chromosomes produced by the methods of the present disclosure.

[0023] In some embodiments, the recombinant chromosome is mouse chromosome 12 comprising a sequence of a human IGH variable region in place of the mouse Igh variable region. In some embodiments, the mouse Igh variable region comprises VH, DH and JH1-6 gene segments and intervening non-coding sequences. In some embodiments, the human IGH variable region comprises VH, DH and JH1-6 gene segments and intervening non-coding sequences. In some embodiments, the recombinant chromosome is mouse chromosome 6 comprising a sequence of a human IGK variable region in place of the mouse Igk variable region. In some embodiments, the mouse Igk variable region sequence comprises a sequence of a mouse ... k , and J k1-5 The template sequence includes sequences encoding gene segments, as well as intervening non-coding sequences. In some embodiments, the template sequence includes a human IGK variable region sequence. In some embodiments, the human IGK variable region sequence is a human V k , and J k1-5 A gene segment includes sequences that code for the gene, as well as intervening non-coding sequences.

[0024] The present disclosure provides a cell comprising a recombinant chromosome of the present disclosure.

[0025] In some embodiments, the cell is hybridizable with a mouse ES cell. In some embodiments, the cell is an embryonic stem (ES) cell, an embryonic hybrid stem (EHS) cell, or a zygotic cell. In some embodiments, the EHS cell is a hybrid of a human ES cell and a mouse ES cell. In some embodiments, the ES cell is a mouse ES cell. In some embodiments, the cell is a micronucleated cell.

[0026] The present disclosure provides a method comprising generating mouse embryonic stem cells, comprising: (a) fusing a micronucleated cell comprising a recombinant chromosome produced by a method according to any one of the methods of the present disclosure with a mouse ES cell, wherein (i) the mouse ES cell comprises a chromosome homologous to the recombinant chromosome, the homologous chromosome comprising a first fluorescent protein operably linked to a promoter capable of expressing the fluorescent protein in the ES cell, and (ii) at least a subset of the micronucleated cells comprises a recombinant chromosome, the recombinant chromosome comprising a second fluorescent protein different from the first fluorescent protein, the second fluorescent protein operably linked to a promoter capable of expressing the fluorescent protein in the ES cell; (b) selecting ES cells expressing both the first and second fluorescent proteins; (c) culturing the ES cells selected in step (c) until the homologous chromosome is lost by at least a subset of the ES cells; and (d) selecting ES cells expressing the second fluorescent protein and not the first fluorescent protein.

[0027] In some embodiments, culturing the cells in step (c) comprises culturing the cells for at least 5 days, at least 7 days, at least 10 days, or at least 14 days. In some embodiments, selecting the cells in steps (b) and (d) comprises fluorescence activated cell sorting (FACS).

[0028] The present disclosure provides mouse ES cells produced by the methods of the present disclosure.

[0029] The present disclosure provides a transgenic mouse produced from the mouse ES cells of the present disclosure.

[0030] In some embodiments, the generation of transgenic mice comprises injection of ES cells into diploid blastocysts, nuclear transfer from ES cells into enucleated mouse embryos, or tetraploid embryo complementation. In some embodiments, mouse chromosome 12 comprises a sequence of a human IGH variable region in place of a mouse Igh variable region. In some embodiments, the mouse Igh variable region comprises VH, DH and JH1-6 gene segments and intervening non-coding sequences. In some embodiments, the human IGH variable region comprises VH, DH and JH1-6 gene segments and intervening non-coding sequences. In some embodiments, mouse chromosome 6 comprises a sequence of a human IGK variable region in place of a mouse Igk variable region. In some embodiments, the mouse Igk variable region sequence comprises a sequence of a mouse V k , and J k1-5 The template sequence includes sequences encoding gene segments, as well as intervening non-coding sequences. In some embodiments, the template sequence includes a human IGK variable region sequence. In some embodiments, the human IGK variable region sequence is a human V k , and J k1-5 A gene segment includes sequences that code for the gene and intervening non-coding sequences.

[0031] The present disclosure provides a method of producing an antibody, comprising: (a) challenging a transgenic mouse of the present disclosure with an antigen, whereby the transgenic mouse produces a plurality of antibodies comprising human V, D, and J segments from the human IGH variable region; and (b) isolating an antibody specific to the antigen.

[0032] The present disclosure provides a method of producing an antibody, comprising: (a) exposing a transgenic mouse of the present disclosure to an antigen, whereby the transgenic mouse produces a plurality of antibodies comprising human V and J segments from the human IGK or IGL variable region; and (b) isolating an antibody specific to the antigen.

[0033] The present disclosure provides antibodies derived from antibodies produced by the transgenic mice of the present disclosure. In some embodiments, the antibodies comprise single chain variable fragments (scFv), bispecific antibodies, or multispecific antibodies.

[0034] The present disclosure provides a method of generating a chromosomal rearrangement, comprising: (a) providing a cell comprising a target chromosome comprising a target location and a template chromosome comprising a template sequence; (b) contacting the cell with a nucleic acid molecule comprising, from 5' to 3', a 5' homology arm comprising a nucleotide sequence upstream of the 5' end of the target location, a marker, and a 3' homology arm comprising a nucleotide sequence upstream of the 5' end of the template sequence; (c) generating a double-stranded break at the target location and the 5' end of the template sequence, thereby inserting the marker into the target chromosome 3' of the sequence of the 5' homology arm, followed by insertion of the template sequence, thereby generating a chromosomal rearrangement; and (d) selecting one or more cells that express the marker.

[0035] In some embodiments, the length of the 5' and 3' homology arms of the nucleic acid molecule is between about 20 base pairs (bp) and 2,000 base pairs (bp), between about 50 bp and 1,500 bp, between about 100 bp and 1,400 bp, between about 150 bp and 1,300 bp, between about 200 bp and 1,200 bp, between about 300 bp and 1,100 bp, between about 400 bp and 1,000 bp, or between about 500 bp and 900 bp, or between about 600 bp and 800 bp. In some embodiments, the length of the 5' and 3' homology arms of the nucleic acid molecule is between about 400 bp and 1,500 bp, between about 500 bp and 1,300 bp, or between about 600 bp and 1,000 bp. In some embodiments, the 5' and 3' homology arms of the nucleic acid molecule are between about 600 bp and 1,000 bp in length.

[0036] In some embodiments, generating the double-stranded break in (c) comprises inducing the double-stranded break using a CRISPR / Cas endonuclease and at least one sgRNA, one or more zinc finger nucleases, one or more transcription activator-like effector nucleases (TALENs), or one or more CRE recombinases. In some embodiments, the CRISPR / Cas endonuclease is selected from the group consisting of CasI, CasIB, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, CasX, CasY, Cas12a (Cpf1), Cas12b, Cas13a, CsyI, Csy2, Csy3, CseI, Cse2, CscI, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, CmrI, Cmr3, Cmr4, Cmr5, Cmr6, CsbI, Csb2, Csb3, Csx17, CsxI4, Csx10, Csx16, CsaX, Csx3, In some embodiments, the CRISPR / Cas endonuclease comprises Cas9, Cpf1, CasX, CasY, C2c1, C2c2, or C2c3, or a homolog, ortholog, or modified form thereof. In some embodiments, the CRISPR / Cas endonuclease comprises Cas9. In some embodiments, the generation of the double-stranded break comprises contacting the cell with a CRISPR / Cas endonuclease, at least a first gNA comprising a targeting sequence specific to the target location, such that the CRISPR / Cas endonuclease cleaves the target location, and a second gNA comprising a targeting sequence specific to the 5' end of the template sequence. In some embodiments, contacting the cell with the CRISPR / Cas endonuclease and the sgRNA comprises transfecting the cell with one or more nucleic acid molecules encoding the CRISPR / Cas endonuclease and the sgRNA. In some embodiments, the one or more nucleic acid molecules are plasmids.

[0037] In some embodiments, the marker comprises a fluorescent protein operably linked to a promoter capable of expressing the fluorescent protein in the cell. In some embodiments, the fluorescent protein comprises GFP, YFP, RFP, CFP, BFP, dsRed, mCherry, or tdTomato. In some embodiments, the marker further comprises a selectable marker. In some embodiments, the selectable marker is selected from the group consisting of dihydrofolate reductase (DHFR), glutamine synthetase (GS), puromycin acetyltransferase, blasticidin deaminase, histidinol dehydrogenase, hygromycin phosphotransferase (hph), bleomycin resistance gene, and aminoglycoside phosphotransferase (neomycin resistance).

[0038] In some embodiments, the cells comprise embryonic stem (ES) cells.

[0039] In some embodiments, the nucleic acid molecule is a plasmid.

[0040] The present disclosure provides a cell comprising a chromosomal rearrangement produced by the methods of the present disclosure. In some embodiments, the cell is a mouse ES cell.

[0041] The present disclosure provides transgenic mice from mouse ES cells produced by the methods of the present disclosure. [Brief description of the drawings]

[0042] A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments and the accompanying drawings.

[0043] 1 shows, from top to bottom, mouse immunoglobulin heavy chain complex (Igh), human Igh, and mouse Igh with humanized variable domains (VH, DH, and JH1-6). Chro: chromosome.

[0044] Figure 2 shows the hybridization of recombinant mouse and human embryonic stem (ES) cells by electrofusion. Mouse ES cells express the marker neomycin, and human ES cells express mCherry. Embryonic hybrid stem cells (hybridoma cells) are resistant to G418 and positive for Cherry.

[0045] FIG. 3A shows the arrangement of three pairs of PCR primers (indicated by arrows) in the VH, DH, and JH1-6 regions of the human Igh gene used for genotyping of embryonic hybrid stem (EHS) cells.

[0046] FIG. 3B is an exemplary gel showing PCR results for 12 embryonic hybrid stem (EHS) cell clones that were genotyped using the primers shown in FIG. 3A. 4A-4B are diagrams showing the pipeline for establishing recombinant humanized chromosomes in EHS cells (Figure 4A) via HDR-mediated chromosomal rearrangement (HCMR). EHS cells were co-transfected with a 5' arm homologous to the 5' of mouse Igh gene, a 3' arm homologous to the 5' of human Igh gene, and a 5'HMCR plasmid containing a pCMV-EGFP-PolyA-PGK-Puromycin-PolyA cassette; As shown in TIFF2024533683000002.tif5170, the 3'HMCR plasmid contains a 5' arm homologous to the 3' end of the human Igh variable locus, a 3' arm homologous to the 3' end of the mouse Igh variable locus, and a PGK-Hygromycin-polyA cassette, as well as four plasmids containing Cas9 and sgRNA targeting the 5' and 3' variable domains of mouse Igh and human Igh. Or, (Figure 4B) via CRE-Loxp-mediated chromosomal rearrangement (CMCR): Four plasmids were designed to mediate the CMCR process. The mouse Igh 5' (pCMV-GFP-BGH PolyA-Loxp) and 3' (BGH polyA-Loxp-511-Hygromycin-BGH polyA-PGK-BSD-BGH PolyA) plasmids were designed to insert into the 5' and 3' ends of the mouse Igh variable locus, respectively. At the same time, we designed human IGH 5' (BGH polyA-Loxp-Puro-BGH PolyA-PGK-Neomycin-BGH PolyA) and 3' (pCMV-BGP-BGH PolyA-PGK-Loxp-511) plasmids to be inserted into the 5' and 3' ends of the human IGH variable locus, respectively. For CMCR, the successfully integrated EHS cells were transfected with Cre.

[0047] FIG. 5A shows the location of the PCR primers (indicated by arrows) used to verify recombinant human chromosomes. Figure 5B shows the PCR results using the four pairs of primers shown in Figure 5A. Results of 192 single clones are shown.

[0048] Figure 6 shows the replacement of mouse chromosomes with genetically engineered human chromosomes in mouse ES cells. EHS cells carrying recombinant human chromosomes tagged with GFP are micronized by exposure to colcemid, the microcells are collected by centrifugation, and electrofused to mouse ES cells carrying the corresponding mouse chromosomes tagged with mCherry. GFP+ mCherry+ cells are isolated by fluorescence-activated cell sorting (FACS). The cells are then cultured, and GFP+ mCherry- cells that have lost the mouse chromosome are isolated by FACS.

[0049] FIG. 7A shows the location of the PCR primers (indicated by arrows) used for validation of Igh humanized mice.

[0050] FIG. 7B shows the PCR results of an exemplary Igh humanized mouse using the seven primer pairs shown in the figure.

[0051] FIG. 8A shows the results of fluorescence in situ hybridization (FISH) of Igh humanized mice.

[0052] FIG. 8B shows G-banding karyotype analysis of Igh humanized mice.

[0053] Figure 9A shows the results of whole genome sequencing (WGS) analysis of IGH-V in Igh-humanized mice. H The copy number of WGS sequences for each variable (V) gene segment located in the region is shown.

[0054] Figure 9B shows WGS analysis of IGH-D and IGH-J in Igh-humanized mice. H and J. H The copy numbers of WGS sequences of each diversity (D) gene segment and six joining (J) segments located in the 1-6 region are shown.

[0055] FIG. 10 shows the humanization of the variable domains of the mouse Igk genes.

[0056] 11A-11B show the PCR validation results of Igk-humanized mice. Fig. 11A, Position of the designed primers used in the PCR experiment. Fig. 11B, PCR results of Igk-humanized mice using the five pairs of primers shown in panel A.

[0057] Figure 12 shows the results of WGS analysis of Igk-humanized mice. K and J. k Copy number from WGS sequence of each antibody gene located on the segment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0058] Detailed Description The present disclosure provides methods of chromosome recombination, including transferring large sequence fragments between chromosomes. Using the methods disclosed herein, at least 5 megabase pairs (MB) of sequence can be transferred from a chromosome template to a target chromosome. The methods disclosed herein can also be used to generate chromosomal rearrangements, such as inversions and translocations. Also provided herein are recombinant chromosomes produced by the methods disclosed herein, as well as cells and animals containing the chromosomes, and methods of using the same.

[0059] Manipulating large fragments of genes or chromosomes holds great potential in both basic and translational research, and in the development of therapeutics. Genetic humanization is one of the most common applications in which genes in model organisms, such as mice, are replaced with their human counterparts. For example, mice with humanized Ig genes provide a powerful platform for producing human antibodies in a mouse background. However, the manipulation of large fragments remains one of the most important challenges in the field of gene editing, due to the unavailability of delivery vectors capable of carrying large fragments of chromosomes up to 1 million base pairs (MBs). The payload of conventional delivery vectors, such as adeno-associated virus vectors and other viral vectors, is limited by the size of the viral genome from which the vector is derived.

[0060] The methods disclosed herein allow for efficient in situ replacement of large sequences between chromosomes. These methods, called Massive fragment Across Species In situ Replacement Technology (MASIRT), can replace large portions of a chromosome in a single edit, in some cases replacing up to megabase pairs (MB) of sequence. These methods can be used to efficiently introduce large sequences between species or between chromosomes within a single species. In one example, MASIRT was used to obtain mice with humanized variable domains of the mouse Igh gene. Humans and mice show high similarity in the sequence and expression of antibody genes, and the genomic organization of the heavy chain is also similar between these species. Thus, MASIRT was used to generate humanized variable domains of the V H , D H , J H Approximately 3 MB of mouse genomic sequence containing the entire gene segment was replaced with approximately 1 MB of contiguous human genomic sequence containing the equivalent human gene fragment, resulting in a humanized mouse Igh gene.

[0061] Unlike other methods that only work on embryonic stem cells, the disclosed method can be advantageously used for the replacement of large sequences in zygotes. Embryonic stem cell lines are not generally available for species other than mice. In contrast, zygotes are available for many mammals, and therefore can be used to obtain animals such as rabbits and cows in which genes or gene fragments are humanized using the disclosed method. Furthermore, the methods disclosed herein can be used to replace large sequence fragments at once, e.g., up to at least 5 MB of sequence, which is about five times larger than the methods used by other methods known in the art. This increases efficiency and reduces the time and cost required to generate animals with humanized genes. For example, an Igh humanized mouse can be generated with only three replacements. An additional advantage is that when used in mice, a single replacement takes only 1-3 months, which is half or a third of the time required by other methods known in the art.

[0062] definition A chromosome is a long DNA molecule that contains all or part of an organism's genetic material. Most eukaryotic chromosomes contain packaging proteins called histones, which bind to DNA molecules with the help of chaperone proteins and condense to maintain their integrity. Eukaryotic chromosomes consist of long linear DNA molecules bound to proteins, forming a compact complex of proteins and DNA called chromatin. Each chromosome has one centromere, from which extend one or two arms. Chromosome arms are the ends of chromosomes, and telomeres are regions of repeated nucleotide sequences bound to specialized proteins that ensure the integrity of linear chromosomes by protecting the terminal regions of chromosomal DNA from progressive degradation and by preventing DNA repair systems from misidentifying the ends of DNA strands as double-strand breaks.

[0063] A "gene" includes DNA regions that code for gene products (e.g., proteins, non-coding RNAs) and all DNA regions that control the production of gene products, whether or not such control sequences are adjacent to the coding and / or transcribed sequences. Thus, genes include, but are not necessarily limited to, regulatory element sequences such as promoter sequences, terminators, translational control sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, origins of replication, matrix attachment sites, locus control regions, and the like. A coding sequence encodes a gene product upon transcription or transcription and translation. Coding sequences of the present disclosure may include fragments and need not include full-length open reading frames (ORFs). A gene may include both the transcribed strand and the complementary strand, including the anticodon. A gene may also include exons, including protein-coding sequences and untranslated regions, and introns, which are removed from the final RNA product by splicing.

[0064] The term "promoter" as used herein can refer to a DNA sequence located adjacent to a DNA sequence encoding a recombinant product. A promoter is preferably operably linked to an adjacent DNA sequence. A promoter typically increases the amount of protein or RNA product expressed from a DNA sequence compared to the amount expressed in the absence of the promoter. A promoter from one organism can be utilized to enhance protein expression from a DNA sequence from another organism. For example, a vertebrate promoter can be used to express jellyfish GFP in a vertebrate. Furthermore, a single promoter element can increase the amount of recombinant product expressed for multiple DNA sequences linked in tandem. Thus, a single promoter element can enhance the expression of one or more recombinant products. Multiple promoter elements are well known to those skilled in the art.

[0065] The term "enhancer" as used herein can refer to a DNA sequence located adjacent to a DNA sequence that codes for a protein or RNA product, or a DNA sequence located distal to the DNA sequence that codes for a protein or RNA product. Enhancer elements are usually located upstream of promoter elements, but can also be located downstream or within a coding DNA sequence, such as within an intron. In some cases, enhancers can be located kilobases away, or even tens or hundreds of kilobases away, from the gene whose expression they control. Enhancer elements can increase the amount of protein or RNA product expressed from a DNA sequence more than the increase in expression caused by the promoter element. Multiple enhancer elements are readily available to one of skill in the art.

[0066] As used herein, the term "foreign chromosome" or "foreign sequence" refers to a foreign chromosome or sequence with respect to the genome of an animal. For example, in a mouse cell in which all chromosomes except one human chromosome are mouse chromosomes, the human chromosome is the foreign chromosome. Similarly, in a mouse chromosome in which some of the mouse sequences are replaced with human sequences, the human sequences are referred to as foreign sequences. Similarly, "endogenous" refers to a chromosome or sequence that originates from the organism, such as the mouse chromosomes and sequences described above.

[0067] As used herein, the term "homologous recombination" refers to a type of genetic recombination in which nucleotide sequences are exchanged between two similar or identical DNA molecules known as homologous sequences or homology arms. Homologous recombination often involves the following basic steps: After a double-strand break (DSB) is created in both strands of DNA, the DNA portion around the 5' end of the DSB is excised in a process called resection. In a subsequent step of strand invasion, the overhanging 3' end of the cut DNA molecule "invades" an uncut similar or identical (or homologous) DNA molecule, e.g., a homology arm. After strand invasion, the additional sequence follows one of two pathways: the DSBR (double-strand break repair) pathway or the SDSA (synthesis-dependent strand annealing) pathway.

[0068] As used herein, "DNA repair pathway" refers to a cellular mechanism that allows cells to maintain genome integrity function in response to detection of DNA damage, such as single- or double-strand breaks in DNA. Depending on the type and extent of DNA damage and the phase of the cell cycle, DNA repair pathways include resection, canonical homology-directed repair (canonical HDR), homologous recombination (HR), alternative homology-directed repair (alt-HDR), double-strand break repair (DSBR), single-strand annealing (SSA), synthesis-dependent strand annealing (SDSA), break-induced replication (BIR), alternative end joining (alt-EJ), microhomology-mediated end joining (MMEJ), DNA synthesis-dependent microhomology-mediated end joining (SD-MMEJ), non-homologous end joining (NHEJ) pathways such as canonical non-homologous end joining (C-NHEJ) repair, alternative non-homologous end joining (A-NHEJ) pathway, transleucine DNA synthesis (TLS) repair, These include, but are not limited to, base excision repair (BER), nucleotide excision repair (NER), mismatch repair (MMR), DNA damage response (DDR), blunt end joining, single strand break repair (SSBR), interstrand crosslink repair (ICL), and Fanconi anemia pathway (FA).

[0069] As used herein, homology-directed repair (HDR) refers to the process of repairing DNA damage using homologous nucleic acids (e.g., sister chromatids or exogenous nucleic acids). In normal cells, HDR usually involves a series of steps, such as recognition of breaks, stabilization of breaks, resection, stabilization of single-stranded DNA, formation of DNA crossover intermediates, resolution of crossover intermediates, and ligation.

[0070] As used herein, "homolog" refers to a protein in a group of proteins that perform the same biological function, for example, proteins that belong to the same protein family, provide a common trait, or perform the same or similar biological function. Homologs are expressed by homologous genes. A homologous gene is a gene that encodes a protein that has the same or similar biological function as a protein encoded by a second gene. Homologous genes are generated by speciation (orthologs) or gene duplication (paralogs). "Orthologs" refers to a set of homologous genes in different species that have evolved from a common ancestral gene by speciation. Orthologs usually retain the same function during evolution. "Paralogues" refers to a set of homologous genes in the same species that have diverged from each other as a result of gene duplication. Thus, homologous genes can be from the same organism or from different organisms. Homologous genes include naturally occurring alleles and artificially created variants. The percent identity between homologous proteins depends on the origin of the protein and the extent to which the organism species from which the protein is derived have diverged. Homologous proteins from more closely related species (e.g., two mammals such as human and mouse) are generally more similar than proteins from more distantly related species (e.g., chicken and mouse). When optimally aligned, homologous proteins usually have at least about 40% identity, about 50% identity, about 60% identity, and sometimes at least about 70%, such as about 80%, or even at least about 90% identity over the entire length of the protein. In other cases, such as when comparing proteins from highly divergent species, homologous proteins have at least about 40% identity, about 50% identity, about 60% identity, about 70%, about 80% identity, or about 90% identity over the length of a conserved protein domain, such as a DNA-binding domain.

[0071] Homologous genes or proteins are identified by comparison of DNA or amino acid sequences, e.g., manually or by the use of computer-based tools that use known homology-based search algorithms, such as those commonly referred to as BLAST, FASTA, Smith-Waterman. Local sequence alignment programs, e.g., BLAST, can be used to search sequence databases to find similar sequences, and summary expectation values ​​(E-values) can be used to measure sequence base similarity. Since proteins with the highest E-values ​​in a particular organism are not necessarily orthologs, i.e., have the same function, or are not necessarily the only orthologs, a reciprocal query can be used to filter hit sequences with significant E-values ​​for ortholog identification. A reciprocal query searches for significant hits against a database of amino acid sequences from base organisms that are similar to the sequence of the query protein. If the best hit of the reciprocal query is the query protein itself or a protein encoded by a gene that has been duplicated after speciation, the hit can be identified as an ortholog.

[0072] As used herein, "percent identity" refers to the degree to which two optimally aligned DNA or protein segments are invariant throughout a window of constituent sequences, such as nucleotide or amino acid sequences. The "percent identity" of aligned segments of a test sequence and a reference sequence is the number of identical components shared by the sequences of the two aligned segments divided by the total number of sequence components of the reference segment over the smaller of the alignment window of the entire test sequence or the entire reference sequence. The "percent identity" ("% identity") is the percent identity multiplied by 100. It is understood that such optimal alignment is considered to be the local alignment of DNA sequences. In the alignment of protein sequences, the local alignment of protein sequences should allow for the introduction of gaps in order to achieve optimal alignment. The percent identity can be calculated over the aligned length, not including gaps introduced by the alignment itself.

[0073] As used herein, "specific for" when used in reference to a nucleotide sequence, such as a homology arm or targeting sequence of a guide RNA, refers to a sequence that is identical or substantially identical to another nucleotide sequence or the reverse complement of the other nucleotide sequence. A sequence that is "specific" for another sequence can hybridize to the other sequence or its reverse complement by Watson-Crick base pairing. Thus, one skilled in the art will understand that a sequence that is specific for another sequence is very similar to the other sequence or its reverse complement, but need not be completely identical. For example, a sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 97% or at least 99% identical to another sequence is still specific for that sequence if it can hybridize to the other sequence. As a further example, a guide nucleic acid targeting sequence can contain one, two, three or more mismatches to the target sequence, depending on the position of the mismatch in the targeting sequence, but is still specific for the target sequence if it can target a ribonucleoprotein complex that includes a gNA and an endonuclease to the target sequence.

[0074] As used herein, "selection" refers to the separation of two populations of distinct products using any one of the methods known in the art. Selection applied to one or more cells, chromosomes, or sequences can be based on a marker, such as a selectable marker. To select cells expressing a selectable marker, a mixed cell population containing cells that express and cells that do not express the marker is cultured in a selective medium, such that cells that do not express the marker are killed or inhibited from proliferating. Sequences or chromosomes containing the marker can also be selected by placing them in cells and applying a selective regimen. Similarly, selection can be based on a detectable marker, such as a fluorescent protein. Cells expressing a detectable marker can be physically removed from a mixed cell population based on the detectable marker using methods known in the art, such as fluorescence activated cell sorting (FACS). Alternatively, or in addition, the mixed cell population can be diluted so that single cells can be isolated and cultured, and clones derived from the isolated cells can be assayed for the presence of one or more traits, such as a marker.

[0075] As used herein, "derived" refers to the source or origin of a molecular entity, e.g., a nucleic acid or a protein. The source of a molecular entity may be naturally occurring, recombinant, unpurified, or purified. For example, a polypeptide derived from a second polypeptide contains an amino acid sequence that is identical or substantially similar, e.g., 50% or more homologous, to the amino acid sequence of the second protein. A derived molecular entity, e.g., a nucleic acid or protein, may contain one or more modifications, e.g., one or more amino acid or nucleotide changes.

[0076] "Isolate" refers to a molecular entity that has been purified, removed or isolated from its source or origin.

[0077] A "naturally occurring" sequence refers to a sequence that is found in at least one species found in nature.

[0078] "Artificial sequence" refers to a sequence that does not occur in nature. An artificial sequence may be similar to a naturally occurring sequence, but contains one or more modifications compared to a naturally occurring sequence. Alternatively, an artificial sequence may have little or no similarity to a naturally occurring sequence. A chimeric or recombinant sequence is an operably linked sequence of two sequences that are derived from different sources or that are not found adjacent to each other, and is one type of artificial sequence.

[0079] "Operably linked" refers to the juxtaposition of genetic elements where the elements are in a relationship that allows them to operate in an expected manner. For example, a promoter is operably linked to a coding region if it helps initiate transcription of the coding sequence. There may be intervening residues between the promoter and the coding region so long as this functional relationship is maintained.

[0080] Stem cells are classified herein as follows: The most pluripotent and the earliest in development are "embryonic stem (ES) cells" or "ES cells". ES cells may be fresh primary cells or derived from ES cell lines. All other stem cells derived from somatic tissues (all tissues except germ tissues) are generally defined as "somatic stem cells", but may generally be known as one or more of the following: "adult stem cells", "mature stem cells", "progenitor cells", "precursor stem cells", "progenitor cells". Another class of non-ES cells is defined as "germline stem cells". Finally, non-stem cells are referred to herein as "mature cells", but are also known as "differentiated cells", "mature differentiated cells", "terminally differentiated cells", "somatic cells". Mature cells may also be primary isolated cells from tissues, or immortal or tumor-derived cell lines. The disclosure further encompasses "precursor forms of mature cells," which includes all cells that do not meet the commonly used scientific definitions of stem cells or mature cells. ES cells can be cultured in vitro for extended periods of time and induced to resume the normal program of embryonic development and differentiate into all cell types of the adult animal, including germ cells, before being inserted / injected into the cavity of a normal blastocyst.

[0081] As used herein, a "hybrid cell" refers to a cell that contains elements of two genomes. One skilled in the art will understand that a hybrid cell can contain two complete or nearly complete genomes from separate sources. Alternatively, a hybrid cell can contain a complete genome from one source and only a few chromosomes, one chromosome, or a portion of one chromosome from a second source. A cell that contains a mixture of elements of two genomes between the two extremes described above is still considered a hybrid cell. The two genomes contained in the hybrid can be from different individuals, different strains of the same species, or different species. Hybrid cells can be generated by any method known in the art, including, but not limited to, cell fusion, micronuclear cell fusion (MMCT), which transfers a small number of chromosomes from one cell to another, and the like.

[0082] As used herein, "hybrid embryonic stem (EHS)" cells refer to hybrid cells that have the properties of embryonic stem cells. EHS cells can be generated by fusion of ES cells from two different species or by chromosomal transfer via MMCT from a cell of one species to a stem cell of another species.

[0083] As used herein, "cancer" refers to a disease, condition, trait, genotype, or phenotypes characterized by unregulated cell growth or replication, as known in the art. Cancer includes both solid and liquid tumors. Exemplary cancers include, but are not limited to, leukemia, breast cancer, bone cancer, brain cancer, head and neck cancer, retinal cancer, esophageal cancer, gastric cancer, multiple myeloma, ovarian cancer, uterine cancer, thyroid cancer, testicular cancer, endometrial cancer, melanoma, colon cancer, lung cancer, bladder cancer, prostate cancer, lung cancer (including small cell lung cancer and non-small cell lung cancer), pancreatic cancer, sarcoma, cervical cancer, head and neck cancer, and skin cancer.

[0084] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.

[0085] Chromosome recombination method The present disclosure provides a method for recombining chromosomes using a template chromosome, a target chromosome, one or more nucleic acid molecules such as vectors or plasmids, and homology-directed repair. A nuclease is used to generate double-stranded breaks flanking the template sequence in the template chromosome and double-stranded breaks flanking the target sequence in the target chromosome, or double-stranded breaks at a target location in the target chromosome. One or more nucleic acid molecules including a marker and homology arms containing the sequences of the target chromosome and the template chromosome are used to induce chromosomal rearrangement by replacing the target sequence with the template sequence, inserting the template sequence into the target location, or binding the target sequence with the template sequence at the double-stranded break site.

[0086] In some embodiments, the methods involve replacing the target sequence with the template sequence, i.e., the target sequence is deleted by insertion of the template sequence.

[0087] In some embodiments, the method comprises replacing the target sequence with the template sequence. Any suitable template sequence and any suitable target sequence can be used in the methods described herein. For example, the method can be used to replace a part of a chromosome of a model organism with a homologous human sequence, thereby humanizing that part of the genome of the model organism. Alternatively, a large sequence can be inserted into the target location with little or no deletion of the target sequence.

[0088] In some embodiments, the disclosure provides a method of generating a recombinant chromosome, comprising: (a) providing a cell comprising a target chromosome comprising a target sequence and a template chromosome comprising a template sequence; (b) contacting the cell with (i) a first nucleic acid molecule comprising, from 5' to 3', a 5' homology arm comprising a nucleotide sequence upstream of the 5' end of the target sequence, at least a first marker, and a 3' homology arm comprising a nucleotide sequence upstream of the 5' end of the template sequence; and (ii) a second nucleic acid molecule comprising, from 5' to 3', a 5' homology arm comprising a nucleotide sequence downstream of the 3' end of the template sequence, at least a second marker, and a 3' homology arm comprising a nucleotide sequence downstream of the 3' end of the target sequence; (c) generating double-stranded breaks on both sides of the target sequence and at the 5' and 3' ends of the template sequence, thereby inserting the template sequence and the first and second markers into the target chromosome; and (d) selecting a cell or a plurality of cells expressing the first and second markers. In some embodiments, the first and / or second nucleic acid molecule is a plasmid. In some embodiments of the methods described herein, the arrangement of the template sequence, the target sequence, and the homology arms of the first and second nucleic acid molecules is shown in Figure 4A-4B. In some embodiments, the first marker is located at the 5' end of the template sequence, and the second marker is located at the 3' end of the template sequence after the insertion of the template sequence. For example, the recombinant chromosome produced by the methods described herein comprises, from 5' to 3', the target chromosome sequence upstream of the target sequence, the first marker, the template sequence, the second marker, and the target chromosome sequence downstream of the target sequence after the insertion of the template sequence and the deletion of the target sequence.

[0089] Those skilled in the art will understand that many lengths of template sequences are suitable for the methods described herein. Suitable template sequences can range from as little as a few hundred base pairs to as long as hundreds of megabase pairs, including most of a chromosome. In some embodiments of the methods described herein, the length of the template sequence is at least 25KB, at least 50KB, at least 100KB, at least 200KB, at least 400KB, at least 500KB, at least 600KB, at least 700KB, at least 800KB, at least 900KB, at least 1MB, at least 2MB, at least 3MB, at least 4MB, at least 5MB, at least 10MB, at least 15MB, at least 20MB, at least 50MB, at least 100MB, at least 150MB, at least 200MB, or at least 250MB. In some embodiments, the length of the template sequence is between 50KB and 250MB, between 100KB and 200MB, between 200KB and 50MB, between 500KB and 50MB, between 1MB and 100MB, between 1MB and 10MB, between 1MB and 5MB, between 1MB and 3MB, between 5MB and 50MB, between 5MB and 10MB, between 3MB and 10MB, or between 5MB and 50MB.

[0090] In some embodiments of the methods described herein, the template chromosome comprises, from 5' to 3', the sequence of the 3' homology arm of the first nucleic acid molecule, the template sequence, and the sequence of the 5' homology arm of the second nucleic acid molecule. In some embodiments, the template chromosome comprises, from 5' to 3', the sequence of the 3' homology arm of the first nucleic acid molecule, a third endonuclease site, the template sequence, a fourth endonuclease site, and the sequence of the 5' homology arm of the second nucleic acid molecule.

[0091] Those skilled in the art will understand that many lengths of target sequence are suitable for the methods described herein. Suitable target sequences can be as small as the endonuclease site used to generate double-stranded breaks (target locations) or can include a large portion of a chromosome, and thus can be up to hundreds of megabase pairs in length. In some embodiments of the methods described herein, the length of the target sequence is at least 25KB, at least 50KB, at least 100KB, at least 200KB, at least 400KB, at least 500KB, at least 600KB, at least 700KB, at least 800KB, at least 900KB, at least 1MB, at least 2MB, at least 3MB, at least 4MB, at least 5MB, at least 10MB, at least 15MB, at least 20MB, at least 50MB, at least 100MB, at least 150MB, at least 200MB, or at least 250MB. In some embodiments, the length of the target sequence is between 50KB and 250MB, between 100KB and 200MB, between 200KB and 50MB, between 500KB and 50MB, between 1MB and 100MB, between 1MB and 10MB, between 1MB and 5MB, between 1MB and 3MB, between 5MB and 50MB, between 5MB and 10MB, between 3MB and 10MB, or between 5MB and 50MB.

[0092] In some embodiments of the methods described herein, the target chromosome comprises, from 5' to 3', the sequence of the 5' homology arm of the first nucleic acid molecule, the target sequence, and the sequence of the 3' homology arm of the second nucleic acid molecule. In some embodiments, the target chromosome comprises, from 5' to 3', the sequence of the 5' homology arm of the first nucleic acid molecule, the first endonuclease site, the target sequence, the second endonuclease site, and the sequence of the 3' homology arm of the second nucleic acid molecule.

[0093] In some embodiments, the nucleic acid molecule used in the methods described herein is a DNA molecule. In some embodiments, the nucleic acid molecule used in the methods described herein is circular, such as a plasmid. Alternatively, additional endonuclease sites can be used to linearize the nucleic acid molecule of the present disclosure. Exemplary endonuclease sites include, but are not limited to, restriction endonucleases, as well as CRISPR / Cas endonucleases, ZFNs and TALENs described herein. Those skilled in the art can incorporate suitable endonuclease sites into the nucleic acid molecule, for example, adjacent to or close to either or both homology arms of the nucleic acid molecule. Those skilled in the art will be able to incorporate suitable CRE recombinase sites into the nucleic acid molecule.

[0094] In some embodiments, the target sequence is deleted by the insertion of the template sequence, and the template chromosome and the target chromosome are cut by the CRISPR / Cas ribonucleoprotein on both sides of the template sequence and the target sequence. In some embodiments, (a) the target chromosome comprises, from 5' to 3', the sequence of the 5' homology arm of the first nucleic acid molecule, the first sgRNA target sequence, the target sequence, the second sgRNA target sequence, and the sequence of the 3' homology arm of the second nucleic acid molecule; and (b) the template chromosome comprises, from 5' to 3', the third sgRNA target sequence, the sequence of the 3' homology arm of the first nucleic acid molecule, the template sequence, the sequence of the 5' homology arm of the second nucleic acid molecule, and the fourth sgRNA target sequence. In some embodiments, the first, second, third, and fourth sgRNA comprise different targeting sequences. For example, the first sgRNA comprises a targeting sequence specific to the first sgRNA target sequence on the target chromosome, the second sgRNA comprises a targeting sequence specific to the second sgRNA target sequence on the target chromosome, the third sgRNA comprises a targeting sequence specific to the third sgRNA target sequence on the template chromosome, and the fourth sgRNA comprises a targeting sequence specific to the fourth sgRNA target sequence on the target chromosome. Alternatively, one or more sgRNA target sequences and the corresponding sgRNA target sequences may be the same sequence.

[0095] In some embodiments, the insertion of template sequence comprises little or no deletion of the sequence of target sequence.Those skilled in the art will understand that many mechanisms of double-strand break repair involve the excision of the cut ends, so that deletions are generated around the endonuclease site described herein.For example, deletions of about 5bp, 10bp, 15bp, 20bp, 25bp, 30bp, 35bp, 40bp, 45bp, or 50bp can be generated by the method described herein around the target position or around the endonuclease site flanking the target sequence.

[0096] In some embodiments, for example in embodiments in which little or no target sequence is deleted by the methods described herein, (a) the target chromosome comprises, from 5' to 3', the sequence of the 5' homology arm of a first nucleic acid molecule, the first sgRNA target sequence, and the sequence of the 3' homology arm of a second nucleic acid molecule; and (b) the template chromosome comprises, from 5' to 3', the second sgRNA target sequence, the sequence of the 3' homology arm of the first nucleic acid molecule, the template sequence, the sequence of the 5' homology arm of the second nucleic acid molecule, and the third sgRNA target sequence. In some embodiments, the first, second, and third sgRNA comprise different target sequences. For example, a first sgRNA includes a targeting sequence specific for a first sgRNA target sequence on a target chromosome, a second sgRNA includes a targeting sequence specific for a second sgRNA target sequence on a target chromosome, and a third sgRNA includes a targeting sequence specific for a third sgRNA target sequence on a template chromosome.

[0097] In some embodiments, the insertion of the template sequence disrupts one or more functions of the target sequence.For example, when the template sequence is inserted into the coding sequence of a gene, it can prevent the expression of the appropriate gene product by generating premature stop codons, mutations in the protein coding sequence, abnormal splice products, etc.Similarly, the insertion of the template sequence into the regulatory sequence of a gene, such as enhancer or promoter, can prevent the expression of the gene.

[0098] In some embodiments, the method of the present disclosure comprises deleting the first and / or second marker after the insertion of the target sequence. The marker can be deleted by any suitable method known in the art. For example, the cell containing the recombinant chromosome can be contacted with a CRISPR / Cas ribonucleoprotein that contains a gNA targeting sequence specific to the sequence that codes for the marker, thereby inducing the deletion of all or part of the marker sequence.

[0099] The disclosed methods can also be used to generate chromosomal rearrangements, such as inversions and translocations. Many of the chromosomal rearrangements are involved in human diseases and disorders, such as cancer. Reproducing such chromosomal rearrangements in model organisms, such as mice, facilitates the study of these diseases and disorders. The chromosomal abnormalities involved are known to those of skill in the art and are described in the Mitelman database, available at mitelmandatabase.isb-cgc.org / . Further information regarding chromosomal abnormalities involved in humanized diseases is also available at rarediseases.info.nih.gov / diseases / diseases-by-category / 36 / chromosome-disorders.

[0100] The present disclosure provides a method of generating a chromosomal rearrangement, comprising: (a) providing a cell comprising a target chromosome comprising a target location and a template chromosome comprising a template sequence; (b) contacting the cell with a nucleic acid molecule comprising, from 5' to 3', a 5' homology arm comprising a nucleotide sequence upstream of the 5' end of the target location, a marker, and a 3' homology arm comprising a nucleotide sequence upstream of the 5' end of the template sequence; (c) generating a double-stranded break at the target location and the 5' end of the template sequence, thereby inserting the marker into the target chromosome 3' of the sequence of the 5' homology arm, followed by insertion of the template sequence, thereby generating a chromosomal rearrangement; and (c) selecting one or more cells that express the marker. Alternatively, the method includes: (a) providing a cell that includes a target chromosome that includes a target position and a template chromosome that includes a template sequence; (b) contacting the cell with a nucleic acid molecule that includes, from 5' to 3', a 5' homology arm that includes a nucleotide sequence downstream of the 3' end of the template sequence, a marker, and a 3' homology arm that includes a nucleotide sequence downstream of the 3' end of the target sequence; (c) generating a double-stranded break at the target position and the 3' end of the template sequence, thereby inserting the marker into the target chromosome 3' of the sequence of the 5' homology arm, followed by inserting the template sequence, thereby generating a chromosomal rearrangement; and (c) selecting one or more cells that express the marker. In some embodiments, generating the double-stranded break includes contacting the cell with a CRISPR / Cas endonuclease, at least a first gNA that includes a targeting sequence specific to the target position, such that the CRISPR / Cas endonuclease cuts the target position, and a second gNA that includes a targeting sequence specific to the 5' end of the template sequence. In some embodiments, generating double-strand break comprises contacting the cell with CRISPR / Cas endonuclease, at least a first gNA comprising a targeting sequence specific to the target position, such that CRISPR / Cas endonuclease cuts the target position, and a second gNA comprising a targeting sequence specific to the 3' end of the template sequence.In some embodiments, the nucleic acid molecule comprises DNA.In some embodiments, the nucleic acid molecule comprises a plasmid.

[0101] To generate double-stranded breaks in target chromosome and template chromosome, suitable methods known in the art can be used. This can be achieved by selecting homology arm sequences that overlap or include endonuclease sites on target chromosome and template chromosome, particularly for nucleic acid molecules (e.g., plasmids) used as guides for chromosome rearrangement via HDR. In some embodiments, generating double-stranded breaks in (c) comprises using CRISPR / Cas endonuclease and one or more guide nucleic acids (gNA), one or more zinc finger nucleases, one or more transcription activator-like effector nucleases (TALEN), or one or more CRE recombinases to induce double-stranded breaks. For example, Cre recombinase induces inversion of the chromosomal region between two LoxP sites, thereby inserting the template sequence and the first and second markers into the target chromosome. In some embodiments, the CRISPR / Cas endonuclease is selected from the group consisting of CasI, CasIB, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, CasX, CasY, Cas12a (Cpf1), Cas13a, CsyI, Csy2, Csy3, CseI, Cse2, CscI, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, CmrI, Cmr3, Cmr4, Cmr5, Cmr6, CsbI, Csb2, Csb3, Csx17, CsxI4, Csx10, Csx16, CsaX, Csx3, In some embodiments, the CRISPR / Cas endonuclease comprises Cas9, Cas12a (Cpf1), Cas13a, CasX, CasY, C2c1, or C2c3. In some embodiments, the CRISPR / Cas endonuclease comprises Cas9. In some embodiments, the gNA comprises a single guide RNA (sgRNA).

[0102] Any suitable method known in the art can be used to contact cells with the endonuclease described herein.For example, in the case of CRISPR / Cas endonuclease, the nucleic acid molecule (e.g., plasmid, etc.) that contains the endonuclease and the sequence that codes for gNA can be used to transfect cells.Alternatively, the endonuclease or the nucleic acid molecule that codes for the endonuclease can be introduced into cells by electroporation, lipofection, transduction, etc.

[0103] The cells used to carry out the methods described herein can be any suitable cells known in the art. In some embodiments, the cells include embryonic stem (ES) cells. In some embodiments, the cells include embryonic hybrid stem (EHS) stem cells. EHS cells can be generated by fusing ES cells of two different species, for example, human and mouse, human and rat, or mouse and monkey. All fusion methods known in the art are contemplated within the scope of this disclosure, including, but not limited to, electrical fusion, virus-induced fusion, chemically induced fusion, and the like. In some embodiments, the method includes fusing a human EH cell with an EH cell selected from the group consisting of mouse, rat, rabbit, guinea pig, hamster, sheep, goat, donkey, cow, horse, camel, chicken, and monkey. In some embodiments, the method includes fusing EH cells from any two different species selected from the group consisting of mouse, rat, rabbit, guinea pig, hamster, sheep, goat, donkey, cow, horse, camel, chicken, and monkey.

[0104] In some embodiments, the cell comprises a zygote. As used herein, the term "zygote" refers to a eukaryotic cell formed by the fertilization event between two gametes, such as a mammalian egg and a mammalian sperm. Zygotes at the single cell, two cell, four cell, eight cell or higher stage are suitable for the methods described herein.

[0105] After the recombinant chromosomes described herein are generated, any suitable method can be used to recover the recombinant chromosomes. In some embodiments, the recovery of the recombinant chromosomes described herein comprises micronuclear cell fusion (MMCT). The recovered chromosomes are introduced into a cell type suitable for downstream applications by fusing micronucleated cells containing the recombinant chromosomes with target cells, such as ES cells. These methods are described in more detail below.

[0106] Template chromosome The present disclosure provides a template chromosome comprising a template sequence for use in the methods described herein.

[0107] As used herein, a "template chromosome" refers to a chromosome that contains a "template sequence." A template sequence refers to a sequence that is introduced into a target chromosome, or target location, using the methods of the present disclosure.

[0108] The template chromosome can be isolated or derived from any suitable source. In some embodiments, the template chromosome is from a eukaryotic organism. In some embodiments, the eukaryotic organism is a vertebrate, such as a bird, a reptile, or a mammal. In some embodiments, the template chromosome is from a mouse, a rat, a rabbit, a guinea pig, a hamster, a sheep, a goat, a donkey, a cow, a horse, a camel, a monkey, or a chicken. In some embodiments, the template chromosome is from a human.

[0109] In some embodiments, the template chromosome is an exogenous chromosome and the template sequence is an exogenous sequence, for example, the target chromosome is a mouse chromosome and the template chromosome and the corresponding template sequence are from a non-mouse species, such as human.

[0110] In some embodiments, the template chromosome is an endogenous chromosome and the template sequence is an endogenous sequence. For example, the template chromosome is a mouse chromosome and the target chromosome is a different second mouse chromosome.

[0111] In some embodiments, the template chromosome is an artificial chromosome.

[0112] In some embodiments, the template chromosome is a naturally occurring chromosome.

[0113] In some embodiments, the template chromosome comprises one or more modifications of naturally occurring chromosomes.Modifications include, inter alia, insertion, deletion, and rearrangement of sequences.Examples of sequences that are inserted into the template chromosome include, inter alia, markers, promoters, cDNA sequences, non-coding sequences, etc.

[0114] In some embodiments, the template chromosome comprises an endonuclease site located 5' to the template sequence. In some embodiments, the template chromosome comprises an endonuclease site located 3' to the template sequence. In some embodiments, the endonuclease site is located immediately adjacent to the template sequence. In some embodiments, the endonuclease site is located near the template sequence.

[0115] In some embodiments, the template chromosome comprises an endonuclease site on each side of the template sequence. For example, the template chromosome comprises a first endonuclease site located 5' of the template sequence and a second endonuclease site located 3' of the template sequence. In some embodiments, both the first endonuclease site and the second endonuclease site are recognized and cleaved by the same endonuclease. For example, both the first endonuclease site and the second endonuclease site comprise the same DNA sequence recognized by the same endonuclease. In some embodiments, the first endonuclease site is cleaved by a first endonuclease and the second endonuclease site is cleaved by a second endonuclease. For example, the first and second endonuclease sites comprise different DNA sequences recognized by two different zinc finger nucleases (ZFNs), or comprise two different CRISPR / Cas target sequences recognized by CRISPR / Cas ribonucleoprotein complexes that comprise guide nucleic acids (gNAs) that comprise different targeting sequences. In some embodiments, the first and / or second endonuclease sites are located immediately adjacent to the template sequence. In some embodiments, the first and / or second endonuclease sites are located near the template sequence.

[0116] A sequence within 5 base pairs (bp), 10 bp, 15 bp, 20 bp, 30 bp, 40 bp, 50 bp, 70 bp, 80 bp, 90 bp, 100 bp, 120 bp, 140 bp, 160 bp, 180 bp, 200 bp, 250 bp, 300 bp, 400 bp, or 500 bp of the template sequence can be considered to be near the template sequence.

[0117] In some embodiments, the template chromosome comprises one or more sequences of the homology arm of the nucleic acid molecule used to promote homology-directed repair. In some embodiments, the template chromosome comprises a sequence of the homology arm located at or near the 5' end of the template sequence. In some embodiments, the homology arm is located upstream, i.e., 5', of the template sequence. In some embodiments, the template chromosome comprises, from 5' to 3', an endonuclease site, a homology arm sequence, and a template sequence. In some embodiments, the template chromosome comprises a sequence of the homology arm located at or near the 3' end of the template sequence. In some embodiments, the homology arm is located downstream, i.e., 3', of the template sequence. In some embodiments, the template chromosome comprises, from 5' to 3', a template sequence, a homology arm sequence, and an endonuclease site. In some embodiments, the homology arm sequence is located between the endonuclease site and the template sequence.

[0118] In some embodiments, the template chromosome comprises a first homology arm sequence located at or near the 5' end of the template sequence, and a second homology arm sequence located at or near the 3' end of the template sequence. That is, the template chromosome comprises homology arms upstream and downstream of the template sequence. In some embodiments, the first homology arm is a 3' homology arm of a first nucleic acid molecule comprising, from 5' to 3', a 5' homology arm comprising a nucleotide sequence upstream of the 5' end of the target sequence, at least a sequence of a first marker, and a first homology arm sequence. In some embodiments, the second homology arm is a 5' homology arm of a second nucleic acid molecule comprising, from 5' to 3', a 3' homology arm comprising a second homology arm sequence, at least a sequence of a second marker, and a nucleotide sequence downstream of the 3' end of the template sequence. In some embodiments, the template chromosome comprises, from 5' to 3', a first endonuclease site, a first homology arm sequence, a template sequence, a second homology arm sequence, and a second endonuclease site.

[0119] In some embodiments, the first and / or second homology arm sequence is located immediately adjacent to the first and / or second endonuclease site. In some embodiments, the first homology arm sequence is located immediately adjacent to the first endonuclease site, and the second homology arm sequence is located immediately adjacent to the second endonuclease site, where the first homology arm is between the first endonuclease site and the template sequence, and the second homology arm is between the template sequence and the second template sequence. In some embodiments, the first homology arm is between the first endonuclease site and the template sequence, and the second homology arm is between the template sequence and the second template sequence.

[0120] In some embodiments, the first and / or second homology arm sequences are located near the template sequence. A homology arm that is within 0 bp, 5 base pairs (bp), 10 bp, 15 bp, 20 bp, 30 bp, 40 bp, 50 bp, 70 bp, 80 bp, 90 bp, 100 bp, 120 bp, 140 bp, 160 bp, 180 bp, 200 bp, or 250 bp of the template sequence may be considered to be near the template sequence.

[0121] In some embodiments, the template chromosome comprises, from 5' to 3', a first endonuclease site, a first homology arm, a template sequence, a second homology arm, and a second endonuclease site.

[0122] In some embodiments, the length of the first and / or second homologous sequence of the template chromosome is between about 20-2,000 bp, between about 50-1,500 bp, between about 100-1,400 bp, between about 150-1,300 bp, between about 200-1,200 bp, between about 300-1,100 bp, between about 400-1,000 bp, or between about 500-900 bp, or between about 600-800 bp. In some embodiments, the length of the homologous sequence of the template chromosome is between about 400 bp and 1,500 bp. In some embodiments, the length of the homologous sequence of the template chromosome is between about 500-1,300 bp. In some embodiments, the length of the homologous sequence of the template chromosome is between about 600-1,000 bp.

[0123] Template sequence The template chromosome comprises a template sequence and serves as a source of the template sequence in recombinant chromosome and the method described herein.The template sequence can be located at any suitable position on the template chromosome.For example, without wishing to be bound by theory, the template sequence can be present in the euchromatin region on the chromosome.

[0124] The template sequence can be isolated or derived from any suitable source. In some embodiments, the template sequence comprises an endogenous sequence, e.g., a sequence endogenous to the template chromosome or a sequence endogenous to the species from which the target chromosome originated. In some embodiments, the template sequence is an exogenous sequence. For example, the template sequence is of a sequence exogenous to the species from which the target chromosome originated. In some embodiments, the template sequence comprises a naturally occurring sequence. In some embodiments, the template sequence comprises one or more modifications of a naturally occurring sequence. Modifications include, inter alia, insertions, deletions, and rearrangements of sequences such as artificial sequences or markers. In some embodiments, the template sequence comprises an artificial sequence. In some embodiments, the template sequence comprises both naturally occurring and artificial sequences. Exemplary artificial sequences include, inter alia, markers, cDNA sequences, promoters, and recombination sequences. Exemplary markers include, but are not limited to, the selectable markers disclosed in Table 3 below, as well as detectable markers such as green fluorescent protein (GFP), mCherry, and the like.

[0125] In some embodiments, the template sequence is from a eukaryote. In some embodiments, the eukaryote is a vertebrate, such as a bird, a reptile, or a mammal. In some embodiments, the template sequence comprises a mouse, rat, rabbit, guinea pig, hamster, sheep, goat, donkey, cow, horse, camel, monkey, or chicken sequence. In some embodiments, the template sequence comprises a human sequence.

[0126] In some embodiments, the length of the template sequence is at least 25 KB, at least 50 KB, at least 100 KB, at least 200 KB, at least 400 KB, at least 500 KB, at least 600 KB, at least 700 KB, at least 800 KB, at least 900 KB, at least 1 MB, at least 2 MB, at least 3 MB, at least 4 MB, at least 5 MB, at least 6 MB, at least 7 MB, at least 8 MB, at least 9 MB, at least 10 MB, at least 15 MB, at least 20 MB, at least 25 MB, at least 30 MB, at least 40 MB, at least 50 MB, at least 60 MB, at least 70 MB, at least 80 MB, at least 90 MB, at least 100 MB, at least 120 MB, at least 140 MB, at least 160 MB, at least 180 MB, at least 200 MB, at least 220 MB, or at least 250 MB. In some embodiments, the length of the template sequence is at least 50 KB, at least 100 KB, at least 200 KB, at least 500 KB, at least 700 KB, at least 1 MB, at least 2 MB, at least 3 MB, at least 4 MB, at least 5 MB, at least 6 MB, at least 7 MB, at least 8 MB, at least 9 MB, at least 10 MB, at least 20 MB, at least 30 MB, at least 40 MB, or at least 50 MB. In some embodiments, the length of the template sequence is at least 1 MB. In some embodiments, the length of the template sequence is at least 2 MB. In some embodiments, the length of the template sequence is at least 3 MB. In some embodiments, the length of the template sequence is at least 4 MB. In some embodiments, the length of the template sequence is at least 5 MB. In some embodiments, the length of the template sequence is at least 10 MB. In some embodiments, the length of the template sequence is at least 20 MB.

[0127] In some embodiments, the length of the template sequence is between 50KB and 250MB, 50KB and 100MB, 50KB and 50MB, 50KB and 20MB, 50KB and 10MB, 50KB and 5MB, 50KB and 3MB, 50KB and 2MB, 50KB and 1MB, 100KB and 200MB, 100KB and 100MB, 100KB and 50MB, 100KB and 20MB, 100KB and 10MB, 100KB and 5MB, 100KB and 3MB, 100KB and 2MB, 100KB and 1MB, 100KB and 500KB, 200KB and 100MB, 200KB and 50MB, 200KB and 20MB, 200KB and 10MB, 200KB and 5MB, 200KB and 3MB, The file size is between 200KB and 2MB, 200KB and 1MB, 200KB and 500KB, 500KB and 100MB, 500KB and 50MB, 500KB and 20MB, 500KB and 10MB, 500KB and 5MB, 500KB and 3MB, 500KB and 2MB, 500KB and 1MB, 1MB and 100MB, 1MB and 50MB, 1MB and 20MB, 1MB and 10MB, 1MB and 5MB, 1MB and 3MB, 1MB and 2MB, 3MB and 100MB, 3MB and 50MB, 3MB and 20MB, 3MB and 10MB, 3 MB and 5MB, 5MB and 100MB, 5MB and 50MB, 5MB and 20MB, 5MB and 10MB, 10MB and 100MB, 10MB and 50MB, or 10MB and 20MB. In some embodiments, the template sequence is between 50 KB and 250 MB in length. In some embodiments, the template sequence is between 500 KB and 200 MB in length. In some embodiments, the template sequence is between 200 KB and 50 MB, between 1 MB and 20 MB, between 1 MB and 10 MB, between 1 MB and 5 MB, between 1 MB and 3 MB, between 3 MB and 20 MB, between 3 MB and 10 MB, between 3 MB and 7 MB, or between 3 MB and 5 MB in length. In some embodiments, the template sequence is between 1 MB and 10 MB in length. In some embodiments, the template sequence is between 1 MB and 5 MB in length. In some embodiments, the template sequence is between 3 MB and 5 MB in length.

[0128] In some embodiments, the template sequence comprises the sequence of one or more genes. In some embodiments, the template sequence comprises the sequence of multiple genes. In some embodiments, the template sequence comprises the sequence of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1000, 1500, or 2000 genes.

[0129] In some embodiments, the template sequence comprises a human sequence, such as the sequence of one or more human genes. In some embodiments, the template sequence comprises a subsequence of a human gene. In some embodiments, the template sequence comprises a subsequence of a human gene and an artificial sequence, such as a marker or a fusion protein. In some embodiments, the template sequence comprises the sequence of one or more human genes and an artificial sequence.

[0130] In some embodiments, the template sequence comprises the sequence of a human gene.All human genes are contemplated within the scope of this disclosure.Without wishing to be bound by theory, human genes involved in disease development or potential therapeutic targets can be introduced into model organisms such as mice to facilitate disease research and development of suitable therapeutics.

[0131] Exemplary genes for inclusion in the template sequence include, but are not limited to, immunoglobulin genes, T cell receptor (TCR) genes, immune checkpoint genes, cytokines, chemokines, receptors, transcription factors, cytoskeleton genes, cell cycle checkpoint genes, oncogenes, and genes involved in development, immunology, or neurobiology. Exemplary immune checkpoint genes include BTLA, CTLA-4, TIM-3, PD-1, and PD-L1. Exemplary cytokines include interleukins (CTNF, IL-16, IL-1B, IL-6, IL-12, IL-17F, IL-2, IL-3, IL-9, IL-12B, IL-21, IL33, leptin, IL-13, IL1A, IL-23, IL-4), interferons (IFN10, IFNα7, IFNa4Fc, IFNβ, IFNα4, IFNγ, IFNα5, IFNω), tumor necrosis factors (TNFs, e.g., BAFF, TNFβ, CD30 ligand, TNFα, CD40 ligand, TNFSF10, CD27 ligand). Exemplary chemokines include CXC, CC CX3C and C family chemokines. Exemplary receptors include G protein-coupled receptors, ligand-gated ion channels (ionotropic receptors), kinase-coupled and related receptors, and nuclear receptors. Exemplary transcription factors include, but are not limited to, helix-turn-helix transcription factors (e.g., Oct-1), helix-loop-helix transcription factors (e.g., E2A), zinc finger transcription factors (e.g., glucocorticoid receptor, GATA proteins), basic protein leucine zipper transcription factors (e.g., cyclic AMP response element binding factor (CREB), activator protein-1 (AP-1)), and beta-sheet motif transcription factors (e.g., nuclear factor-κB (NF-κB)). Exemplary cell cycle control genes include, but are not limited to, cyclins, cyclin-dependent kinases, cell cycle checkpoint genes, and the like.

[0132] In some embodiments, the template sequence comprises an oncogene or a tumor suppressor gene. Exemplary oncogenes and tumor suppressor genes suitable for inclusion in the template sequence are shown in Table 1 below.

[0133] TIFF2024533683000003.tif228170TIFF2024533683000004.tif230170TIFF2024533683000005.tif229170TIFF2024533683000006.tif222170

[0134] In some embodiments, the template sequence comprises a sequence of a human gene associated with a genetic disease or disorder. In some embodiments, the template sequence comprises a sequence of a human chromosomal region associated with a genetic disease or disorder. Non-limiting examples of genes and chromosomal regions associated with diseases and disorders are shown in Table 2 below.

[0135] TIFF2024533683000007.tif222170TIFF2024533683000008.tif226170TIFF20245336830 00009.tif228170TIFF2024533683000010.tif226170TIFF2024533683000011.tif175170

[0136] In some embodiments, the template sequence comprises an immunoglobulin sequence. Both surface and secretory immunoglobulins are contemplated within the scope of this disclosure. Immunoglobulins recognize foreign antigens and initiate an immune response. In humans, each immunoglobulin molecule is composed of two identical heavy chains encoded by the IGH locus on chromosome 14 and two identical light chains encoded by the immunoglobulin kappa locus (IGK) on chromosome 2 and the immunoglobulin lambda locus (IGL) on chromosome 22. The IGH locus includes the V (variable region), D (diverse region), J (joining region), and C (constant region). The V, D, and J regions each contain multiple different gene segments and are collectively referred to herein as the IGH variable region. During B cell development, a recombination event occurs at the DNA level, resulting in the joining of one D segment with a J segment. The fused DJ exons of this partially rearranged DJ region are then joined to the V segment. The rearranged VDJ region containing the fused VDJ exons is then transcribed and fused to the constant region by RNA splicing. This transcript encodes the mu heavy chain. Late in development, B cells produce VDJ-Cmu-Cdelta pre-messenger RNA, which is alternatively spliced ​​to encode either mu or delta heavy chains. Mature B cells in lymph nodes undergo switch recombination, and the fused VDJ gene segment is brought into close proximity with either the IGHG, IGHA, or IGHE gene segments, allowing each cell to express either gamma, alpha, or epsilon heavy chains. Potential recombination of many different V segments with several J segments increases the breadth of antigen recognition. Further diversity is achieved by junction diversity generated by random nucleotide addition by terminal deoxynucleotidyl transferase and by somatic hypermutation. Each light chain consists of two tandem immunoglobulin domains and a constant domain (C L ) and the variable domain (V L). In the case of the light chain, the V domain is encoded by two separate DNA segments. The first segment encodes most of the V domain and is referred to as the V gene segment. The second segment encodes the remainder of the V domain and is referred to as the joining or J gene segment. Like the heavy chain, the light chain also undergoes rearrangement to join the V segment to the J gene segment, bringing the V gene into proximity with the constant region sequence bounded only by introns. Any of the IGH sequences IGHV, IGHD, IGHJ, IGHG, IGHA, or combinations thereof are contemplated within the template sequences of the present disclosure. Any of the light chain sequences IGK or IGL, or combinations thereof are contemplated within the template sequences of the present disclosure.

[0137] In some embodiments, the recombinant chromosome comprises a mouse chromosome with one or more non-coding sequences introduced into said chromosome. For example, one or more non-coding sequences capable of controlling antibody generation, maturation and / or diversification may be introduced into said chromosome. For example, one or more non-coding sequences capable of controlling antibody diversification may be introduced into said chromosome. For example, one or more non-coding sequences capable of controlling antibody class switching may be introduced into said chromosome. For example, one or more non-coding sequences in a switch region may be introduced into said chromosome. For example, class switch recombination, somatic hypermutation and / or activation-induced cytidine deaminase may be controlled when one or more non-coding sequences are introduced into said chromosome. For example, diversity of a repertoire of Ig sequences may be controlled when one or more non-coding sequences are introduced into said chromosome. For example, about 2 kb of variable region including rearranged genes on the heavy chain, kappa light chain and lambda light chain loci, and / or about 4 kb of switch region including extensive stretches of G:C rich DNA on the heavy chain locus may be introduced into said chromosome.

[0138] In some embodiments, the template sequence comprises a human IGH sequence. Human IGH spans nucleotide positions 105,586,437-106,879,844 of chromosome 14 of the GRCh38.p13 assembly of the human genome. Those skilled in the art will understand that human IGH sequences having 5' and 3' boundaries that deviate from those described above, for example, at least 100bp, 500bp, 1,000bp, 2,000bp, 5,000bp, 10,000bp or more, are suitable template sequences.

[0139] In some embodiments, the template sequence comprises a human IGH variable region sequence. In some embodiments, the human IGH variable region sequence is a human V H , D Hand JH1-6 gene segments and sequences encoding the intervening non-coding sequences. In some embodiments, the human IGH variable region sequence comprises nucleotide positions 105,862,994-106,811,028 of chromosome 14 of the GRCh38.p13 assembly of the human genome. In some embodiments, the human IGH variable region sequence comprises nucleotide positions 105,862,994-106,811,028 of chromosome 14 of the GRCh38.p13 assembly of the human genome, minus at least about 50 bp, 100 bp, 500 bp, 1,000 bp, 2,000 bp, 5,000 bp, 7,000 bp, 10,000 bp, 15,000 bp, 20,000 bp, or 50,000 bp from the 5' end or the 3' end, or both. In some embodiments, the human IGH variable region sequence comprises nucleotide positions 105,862,994-106,811,028 of chromosome 14 of the GRCh38.p13 assembly of the human genome, and at least about 50 bp, 100 bp, 500 bp, 1,000 bp, 2,000 bp, 5,000 bp, 7,000 bp, 10,000 bp, 15,000 bp, 20,000 bp, or 50,000 bp of additional flanking sequence at the 5' or 3' end, or both. In some embodiments, the human IGH variable region sequence comprises nucleotide positions 105,862,994-106,811,028 of chromosome 14 of the GRCh38.p13 assembly of the human genome, and one or more modifications thereto. Exemplary modifications include, but are not limited to, a deletion, such as a deletion of one or more V, D or J segments, an insertion, such as an insertion of a marker, a rearrangement, or a combination thereof.

[0140] In some embodiments, the template sequence comprises a sequence of a T cell receptor subunit (TCR). The T cell receptor (TCR) is a protein complex found on the surface of T cells or T lymphocytes [1] that is responsible for recognizing antigen fragments as peptides bound to major histocompatibility complex (MHC) molecules. The TCR comprises a disulfide-linked membrane-bound heterodimeric protein, composed of highly variable α and β chains, most often expressed as part of a complex with the invariant CD3 chain molecule (CD3δ, CD3ε, CD3γ, CD3ζ). T cells expressing these two chains are called α:β (or αβ) T cells. A small number of T cells express an alternative receptor formed by variable γ and σ chains, called γσ T cells. The development of TCRs occurs through a process of lymphocyte-specific genetic recombination. This genetic recombination occurs through recombination of TCR gene segments in T cells of the thymus, resulting in the formation of the final sequence from a large number of potential segments. The TCR alpha locus contains variable (V) and joining (J) gene segments (Vβ and Jβ), while the TCR beta locus contains a D gene segment in addition to the Vα and Jα segments. Thus, the alpha chain is generated from VJ recombination, and the beta chain is involved in VDJ recombination. This is also true for the development of gamma delta TCRs, where the TCR gamma chain is involved in VJ recombination, and the TCR delta genes are generated from VDJ recombination. The TCR alpha chain locus consists of 46 variable segments, 8 joining segments, and a constant region. The TCR beta chain locus consists of 48 variable segments, followed by 2 diversity segments, 12 joining segments, and 2 constant regions. Template sequences comprising any of the sequences, subsequences, or combinations thereof of the TCR subunits described herein are contemplated within the scope of this disclosure. In some embodiments, the template sequence comprises a TCR alpha chain variable region sequence (encoded by the T cell receptor alpha locus, or TRA), a TCR beta chain variable region sequence (encoded by the T cell receptor beta locus, or TRB), a TCR gamma variable region sequence (encoded by the T cell receptor gamma locus, or TRG), or a TCR delta variable region sequence (encoded by the T cell receptor delta locus, or TRD).

[0141] In some embodiments, the template sequence comprises a sequence that encodes an antibody or antigen-binding fragment.

[0142] As used herein, the term "antibody" refers to an immunoglobulin molecule that specifically binds to or immunologically reacts with a particular antigen, including polyclonal, monoclonal, genetically engineered, or otherwise modified forms of antibodies (including, but not limited to, chimeric, humanized, heteroconjugate, e.g., bi-, tri-, and tetra-specific antibodies, bispecific, trispecific, tetraspecific antibodies), and antigen-binding fragments (e.g., Fab', F(ab')2, Fab, Fv, rlgG, scFv fragments, etc.). Unless otherwise indicated, the term "monoclonal antibody" (mAb) is meant to include both intact molecules and antibody fragments (e.g., including Fab and F(ab')2 fragments) that are capable of specifically binding to a target protein. As used herein, Fab and F(ab')2 fragments refer to antibody fragments that lack the Fc fragment of an intact antibody. Examples of these antibody fragments are described herein.

[0143] As used herein, the term "antigen-binding fragment" refers to one or more fragments of an antibody that retain the ability to specifically bind to a target antigen. The antigen-binding function of an antibody can be performed by a fragment of a full-length antibody. The antibody fragment can be, for example, a Fab, F(ab')2, scFv, a bibody, a tribody, an affibody, a nanobody, an aptamer, or a domain antibody. Examples of binding fragments encompassed by the term "antigen-binding fragment" of an antibody include, but are not limited to, (i) a Fab fragment, which is a monovalent fragment consisting of the VL, VH, CL, and CH1 domains; (ii) an F(ab')2 fragment, which is a bivalent fragment comprising two Fab fragments disulfide-linked at the hinge region; (iii) an Fd fragment consisting of the VH and CH1 domains; (iv) an Fv fragment consisting of the VL and VH domains of a single arm of an antibody; (v) a dAb comprising the VH and VL domains; (vi) a dAb fragment consisting of the VH domain (see, e.g., Ward et al., Nature 341:544-546, 1989); (vii) a dAb consisting of a VH or VL domain; (viii) an isolated complementarity determining region (CDR); and (ix) a combination of two or more (e.g., 2, 3, 4, 5 or 6) isolated CDRs, optionally linked by a synthetic linker. Furthermore, although the two domains of an Fv fragment, VL and VH, are encoded by separate genes, they can be joined using recombinant techniques by a linker that allows them to be made into a single protein chain in which the VL and VH regions pair to form a monovalent molecule (known as a single chain Fv (scFv)); see, for example, Bird et al., Science 242:423-426, 1988 and Huston et al., Proc. Natl. Acad. Sci. USA 85:5879-5883, 1988). These antibody fragments can be obtained using conventional techniques known to those of skill in the art, and the fragments can be screened for utility in the same manner as intact antibodies.Antigen-binding fragments can be produced by recombinant DNA techniques, enzymatic or chemical cleavage of intact immunoglobulins, or, in some cases, chemical peptide synthesis procedures known in the art.

[0144] As used herein, the term "complementarity determining region" (CDR) refers to the hypervariable regions found in both the light and heavy chain variable domains of an antibody. The more highly conserved portions of the variable domains are called framework regions (FR). The amino acid positions that define the hypervariable regions of an antibody may vary depending on the context and the various definitions known in the art. Some positions within the variable domains may be considered hybrid hypervariable positions in that they are considered within the hypervariable region under some criteria, while being considered outside the hypervariable region under other criteria. One or more of these positions may also be present in an extended hypervariable region. The antibodies described herein may contain modifications at these hybrid hypervariable positions. Naturally occurring heavy and light chain variable domains each contain four FR regions, which are largely in a β-sheet conformation and are connected by three CDRs to form loops that connect and in some cases form part of the β-sheet structure. The CDRs of each chain are held in close proximity by framework regions in the order FR1-CDR1-FR2-CDR2-FR3-CDR3-FR4 and contribute, together with the CDRs of the other antibody chain, to the formation of the antibody target binding site (see Kabat et al., Sequences of Proteins of Immunological Interest, National Institute of Health, Bethesda, Md., 1987). Numbering of immunoglobulin amino acid residues herein is done according to the immunoglobulin amino acid residue numbering system of Kabat et al., unless otherwise indicated.

[0145] In some embodiments, the antibody or antigen-binding fragment comprises a human antibody or antigen-binding fragment, hi some embodiments, the antibody or antigen-binding fragment is humanized.

[0146] Those skilled in the art will understand that the template sequence may also include sequences necessary for expression of a gene, such as an antibody, in a particular tissue, cell or cells, or organism. Such sequences include, but are not limited to, promoters, enhancers, untranslated sequences such as 5' and 3' untranslated regions of messenger RNA (mRNA), polyadenylation (polyA) sequences, introns, internal ribosome entry sites (IRES), etc. The selection of appropriate sequences will be apparent to those skilled in the art.

[0147] In some embodiments, the template sequence comprises a promoter. In some embodiments, the promoter comprises an endogenous promoter, i.e., the promoter is the promoter normally associated with the gene contained within the template sequence. In some embodiments, the promoter is not an endogenous promoter, e.g., a promoter isolated or derived from another gene or organism than the gene in the template sequence to which the promoter is operably linked. For example, the template sequence comprises a sequence encoding an antibody or antigen-binding fragment operably linked to a promoter that is not an immunoglobulin promoter. In some embodiments, the promoter is a constitutive promoter, an inducible promoter, or a tissue-specific promoter. In some embodiments, the promoter is isolated or derived from a mammalian gene, e.g., a gene expressed in lymphocytes.

[0148] Exemplary promoters that can be used to express the genes of the template sequence include, but are not limited to, the SV40 early promoter region, the promoter contained in the 3' long terminal repeat of Rous sarcoma virus, the metallothionein gene regulatory sequence, promoter elements from yeast and other fungi, such as the tetracycline (Tet) promoter, the Gal4 promoter, the ADC (alcohol dehydrogenase) promoter, the PGK (phosphoglycerol kinase) promoter, the alkaline phosphatase promoter, and the following animal transcriptional control regions that exhibit tissue specificity and have been utilized in transgenic animals: the elastase I gene regulatory region active in pancreatic acinar cells; the insulin gene regulatory region active in pancreatic beta cells, the immunoglobulin gene regulatory region active in lymphoid cells, the mouse mammary tumor virus regulatory region active in testicular, breast, lymphoid and mast cells, the albumin gene regulatory region active in the liver, the alpha-fetoprotein gene regulatory region active in the liver, the alpha1-antitrypsin gene regulatory region active in the liver, the beta-globin gene regulatory region active in myeloid cells, Myelin basic protein gene regulatory region active in oligodendrocyte cells in the brain, myosin light chain-2 gene regulatory region active in skeletal muscle, neuron-specific enolase (NSE) active in neurons, brain-derived neurotrophic factor (BDNF) gene regulatory region active in neurons, glial fibrillary acidic protein (GFAP) promoter active in astrocytes, and gonadotropin-releasing hormone gene regulatory region active in the hypothalamus.

[0149] Target chromosome The present disclosure provides target chromosomes comprising target sequences for use in the methods described herein.

[0150] As used herein, "target chromosome" refers to a chromosome that contains a "target sequence" or, if there is no significant deletion of the target sequence by the insertion of a template sequence, a "target location." Target sequence refers to the sequence of a target chromosome that is deleted by inserting a template sequence using the methods described herein. Target location refers to the location in a target chromosome where a template sequence is inserted (in the case of an insertion) or joined (in the case of a chromosomal translocation or chromosomal rearrangement).

[0151] The target chromosome can be isolated or derived from any suitable source. In some embodiments, the target chromosome is from a eukaryotic organism. In some embodiments, the eukaryotic organism is a vertebrate, such as a bird, a reptile, or a mammal. In some embodiments, the target chromosome is from a mouse, a rat, a rabbit, a guinea pig, a hamster, a sheep, a goat, a donkey, a cow, a horse, a camel, a monkey, or a chicken. In some embodiments, the target chromosome is from a mouse. In some embodiments, the target chromosome is from a rat. In some embodiments, the target chromosome is from a monkey.

[0152] In some embodiments, the template chromosome and the target chromosome are from different species. For example, the template chromosome is from a human and the target chromosome is from a mouse. In some embodiments, the template chromosome and the target chromosome are from the same species.

[0153] In some embodiments, the target chromosome is an artificial chromosome.

[0154] In some embodiments, the target chromosome is a naturally occurring chromosome.

[0155] In some embodiments, the target chromosome comprises one or more modifications of a naturally occurring chromosome. Modifications include, inter alia, insertion, deletion, and rearrangement of sequences. Examples of sequences to be inserted into the target chromosome include, inter alia, markers, promoters, cDNA sequences, non-coding sequences, and the like. Suitable markers include selectable markers such as those disclosed in Table 3, and detectable markers such as GFP, mCherry, and the like.

[0156] In some embodiments, the target chromosome comprises an endonuclease site located 5' of the template sequence. In some embodiments, the target chromosome comprises an endonuclease site located 3' of the template sequence. In some embodiments, the endonuclease site is located immediately adjacent to the target sequence. In some embodiments, the endonuclease site is located near the target sequence.

[0157] In some embodiments, the target chromosome comprises an endonuclease site on each side of the target sequence. For example, the target chromosome comprises a first endonuclease site located 5' of the target sequence and a second endonuclease site located 3' of the target sequence. In some embodiments, both the first endonuclease site and the second endonuclease site are recognized and cleaved by the same endonuclease. For example, both the first endonuclease site and the second endonuclease site comprise the same DNA sequence recognized by the same endonuclease. In some embodiments, the first endonuclease site is cleaved by the first endonuclease and the second endonuclease site is cleaved by the second endonuclease. For example, the first and second endonuclease sites comprise different DNA sequences recognized by two different zinc finger nucleases (ZFNs), or comprise two different CRISPR / Cas target sequences recognized by CRISPR / Cas ribonucleoprotein complexes that comprise guide nucleic acids (gNAs) that comprise different targeting sequences. In some embodiments, the first and / or second endonuclease sites are located immediately adjacent to the template sequence. In some embodiments, the first and / or second endonuclease sites are located near the target sequence.

[0158] An endonuclease site within 5 base pairs (bp), within 10 bp, within 15 bp, within 20 bp, within 30 bp, within 40 bp, within 50 bp, within 70 bp, within 80 bp, within 90 bp, within 100 bp, within 120 bp, within 140 bp, within 160 bp, within 180 bp, within 200 bp, within 250 bp, within 300 bp, within 400 bp, or within 500 bp of the template sequence can be considered to be near the target sequence.

[0159] In some embodiments, the target chromosome comprises one or more sequences of the homology arm of the nucleic acid molecule used to promote homology-directed repair. In some embodiments, the target chromosome comprises a sequence of the homology arm located 5' of the target sequence. In some embodiments, the target chromosome comprises, from 5' to 3', a homology arm sequence, an endonuclease site, and a target sequence. In some embodiments, the target chromosome comprises a sequence of the homology arm located 3' of the target sequence. In some embodiments, the target chromosome comprises, from 5' to 3', a target sequence, an endonuclease site, and a homology arm sequence. In some embodiments, the endonuclease site is located between the homology arm sequence and the target sequence.

[0160] In some embodiments, the target chromosome comprises a first homology arm sequence located 5' of the target sequence and a second homology arm sequence located 3' of the target sequence. That is, the target chromosome comprises homology arms both upstream and downstream of the target sequence. In some embodiments, the first homology arm is a 5' homology arm of a first nucleic acid molecule comprising, from 5' to 3', a first homology arm, at least a sequence of a first marker, and a 3' homology arm comprising a nucleotide sequence upstream of the 5' end of the template sequence. In some embodiments, the second homology arm is a 3' homology arm of a second nucleic acid molecule comprising, from 5' to 3', a 5' homology arm comprising a nucleotide sequence downstream of the 3' end of the template sequence, at least a sequence of a second marker, and a second homology arm. In some embodiments, the target chromosome comprises, from 5' to 3', a first homology arm sequence, a first endonuclease site, a target sequence, a second endonuclease site, and a second homology arm sequence.

[0161] In some embodiments, the first and / or second homology arm sequence of the target chromosome is located immediately adjacent to the first and / or second endonuclease site. In some embodiments, the first homology arm sequence is located immediately adjacent to the first endonuclease site, and the second homology arm sequence is located immediately adjacent to the second endonuclease site, where the first endonuclease site is between the first homology arm and the target sequence, and the second endonuclease site is between the target sequence and the second homology arm.

[0162] In some embodiments, the first and / or second homology arm sequences are located near the target sequence. An endonuclease site that is within 5 bp, 10 bp, 15 bp, 20 bp, 30 bp, 40 bp, 50 bp, 70 bp, 80 bp, 90 bp, 100 bp, 120 bp, 140 bp, 160 bp, 180 bp, 200 bp, or 250 bp of the target sequence may be considered to be near the target sequence.

[0163] In some embodiments, the target chromosome comprises, from 5' to 3', a first homology arm, a first endonuclease site, a target sequence, a second endonuclease site, and a second homology arm.

[0164] In some embodiments, little or no sequence of the target chromosome is deleted when the template sequence is inserted, and the target sequence is referred to herein interchangeably as "target site" or "target location." Those skilled in the art will understand that in these cases, the arrangement of the homology arms and endonuclease sites is similar to that described above, except that the homology arms flank the endonuclease sites of the target location, rather than the target sequence itself being flanked by the endonuclease sites. In some embodiments, the target chromosome comprises, from 5' to 3', the sequence of the first homology arm, the endonuclease site, and the sequence of the second homology arm. In some embodiments, the first homology arm is the 5' homology arm of the first nucleic acid molecule, which comprises, from 5' to 3', the first homology arm, the sequence of at least the first marker, and a 3' homology arm comprising a nucleotide sequence upstream of the 5' end of the template sequence. In some embodiments, the second homology arm is a 3' homology arm of a second nucleic acid molecule that comprises, from 5' to 3', a 5' homology arm that includes a nucleotide sequence downstream of the 3' end of the template sequence, at least the sequence of a second marker, and a second homology arm.

[0165] In some embodiments, the template sequence is bound to the target sequence to generate a chromosomal rearrangement or translocation. In some embodiments, the target chromosome comprises, from 5' to 3', a target chromosome homology arm sequence and an endonuclease site. In some embodiments, the target chromosome homology arm comprises, from 5' to 3', a 5' homology arm of a nucleic acid molecule comprising a target sequence homology arm, at least one marker, and a 3' homology arm comprising a nucleotide sequence upstream of the 5' end of the template sequence. In some embodiments, the target chromosome comprises, from 5' to 3', an endonuclease site and a target chromosome homology arm sequence. In some embodiments, the target chromosome homology arm comprises, from 5' to 3', a 5' homology arm comprising a nucleotide sequence downstream of the 3' end of the template sequence, at least a first marker, and a 3' homology arm of a nucleic acid molecule comprising a target sequence homology arm.

[0166] In some embodiments, the length of the first and / or second homology arm sequence of the target chromosome is between about 20-2,000 bp, between about 50-1,500 bp, between about 100-1,400 bp, between about 150-1,300 bp, between about 200-1,200 bp, between about 300-1,100 bp, between about 400-1,000 bp, or between about 500-900 bp, or between about 600 bp-800 bp. In some embodiments, the length of the homology sequence of the target chromosome is between about 400-1,500 bp. In some embodiments, the length of the homology sequence of the target chromosome is between about 500-1,300 bp. In some embodiments, the length of the homology sequence of the target chromosome is between about 600-1,000 bp.

[0167] Target sequence or target position

[0168] The target chromosome comprises a target sequence or target location into which a template sequence is inserted or bound by the methods described herein. The target sequence can be located at any suitable location on the target chromosome.

[0169] The target sequence can be isolated or derived from any suitable source. In some embodiments, the target sequence and the template sequence are from different species. For example, the template sequence is from human and the target sequence is from mouse. In some embodiments, the target sequence and the template sequence are from the same species.

[0170] In some embodiments, the target sequence comprises a naturally occurring sequence. In some embodiments, the target sequence comprises one or more modifications of the naturally occurring sequence. Modifications include, inter alia, insertions, deletions, and rearrangements of sequences such as artificial sequences or markers. In some embodiments, the target sequence comprises an artificial sequence. In some embodiments, the target sequence comprises both naturally occurring and artificial sequences. Exemplary artificial sequences include, inter alia, markers, cDNA sequences, promoters, and recombination sequences. Exemplary markers include, but are not limited to, the selectable markers disclosed in Table 3 below, as well as detectable markers such as green fluorescent protein (GFP), mCherry, and the like.

[0171] In some embodiments, the target sequence is from a eukaryote. In some embodiments, the eukaryote is a vertebrate, such as a bird, a reptile, or a mammal. In some embodiments, the template sequence comprises a mouse, rat, rabbit, guinea pig, hamster, sheep, goat, donkey, cow, horse, camel, monkey, or chicken sequence. In some embodiments, the target sequence comprises a mouse sequence. In some embodiments, the target sequence comprises a rat sequence. In some embodiments, the target sequence comprises a monkey sequence.

[0172] In some embodiments, the length of the target sequence is at least 25KB, at least 50KB, at least 100KB, at least 200KB, at least 400KB, at least 500KB, at least 600KB, at least 700KB, at least 800KB, at least 900KB, at least 1MB, at least 2MB, at least 3MB, at least 4MB, at least 5MB, at least 6MB, at least 7MB, at least 8MB, at least 9MB, at least 10MB, at least 15MB, at least 20MB, at least 25MB, at least 30MB, at least 40MB, at least 50MB, at least 60MB, at least 70MB, at least 80MB, at least 90MB, at least 100MB, at least 120MB, at least 140MB, at least 160MB, at least 180MB, at least 200MB, at least 220MB, or at least 250MB. In some embodiments, the length of the target sequence is at least 50 KB, at least 100 KB, at least 200 KB, at least 500 KB, at least 700 KB, at least 1 MB, at least 2 MB, at least 3 MB, at least 4 MB, at least 5 MB, at least 6 MB, at least 7 MB, at least 8 MB, at least 9 MB, at least 10 MB, at least 20 MB, at least 30 MB, at least 40 MB, or at least 50 MB. In some embodiments, the length of the target sequence is at least 1 MB. In some embodiments, the length of the target sequence is at least 2 MB. In some embodiments, the length of the target sequence is at least 3 MB. In some embodiments, the length of the target sequence is at least 4 MB. In some embodiments, the length of the target sequence is at least 5 MB. In some embodiments, the length of the target sequence is at least 10 MB. In some embodiments, the length of the target sequence is at least 20 MB.

[0173] In some embodiments, the length of the target sequence is between 50KB and 250MB, 50KB and 100MB, 50KB and 50MB, 50KB and 20MB, 50KB and 10MB, 50KB and 5MB, 50KB and 3MB, 50KB and 2MB, 50KB and 1MB, 100KB and 200MB, 100KB and 100MB, 100KB and 50MB, 100KB and 20MB, 100KB and 10MB, 100KB and 5MB, 100KB and 3MB, 100KB and 2MB, 100KB and 1MB, 100KB and 500KB, 200KB and 100MB, 200KB and 50MB, 200KB and 20MB, 200KB and 10MB, 200KB and 5MB, 200KB and 3MB, The file size is between 200KB and 2MB, 200KB and 1MB, 200KB and 500KB, 500KB and 100MB, 500KB and 50MB, 500KB and 20MB, 500KB and 10MB, 500KB and 5MB, 500KB and 3MB, 500KB and 2MB, 500KB and 1MB, 1MB and 100MB, 1MB and 50MB, 1MB and 20MB, 1MB and 10MB, 1MB and 5MB, 1MB and 3MB, 1MB and 2MB, 3MB and 100MB, 3MB and 50MB, 3MB and 20MB, 3MB and 10MB, 3 MB and 5MB, 5MB and 100MB, 5MB and 50MB, 5MB and 20MB, 5MB and 10MB, 10MB and 100MB, 10MB and 50MB, or 10MB and 20MB. In some embodiments, the target sequence is between 200KB and 50MB, between 1MB and 20MB, between 1MB and 10MB, between 1MB and 5MB, between 1MB and 3MB, between 3MB and 20MB, between 3MB and 10MB, between 3MB and 7MB, or between 3MB and 5MB in length. In some embodiments, the target sequence is between 1MB and 10MB in length. In some embodiments, the target sequence is between 1MB and 5MB in length. In some embodiments, the target sequence is between 3MB and 5MB in length.

[0174] In some embodiments, the target sequence comprises the sequence of one or more genes. In some embodiments, the target sequence comprises the sequence of multiple genes. In some embodiments, the target sequence comprises the sequence of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1000, 1500, or 2000 genes.

[0175] In some embodiments, the target sequence comprises a sequence that is homologous to the template sequence.For example, the template chromosome is a human chromosome that comprises a human template sequence that comprises one or more genes of the genes listed in Table 1 and Table 2 above, and the target chromosome is a mouse chromosome that comprises a mouse target sequence, and the mouse target sequence comprises a mouse sequence that is homologous to the human template sequence.As a further example, the template chromosome is a human chromosome that comprises a human IGH sequence, the target chromosome is a mouse chromosome, and the target sequence comprises a homologous mouse Igh sequence.As a further example, the template chromosome is a human chromosome that comprises a human TCR sequence, the target chromosome is a mouse chromosome, and the target sequence comprises a homologous mouse TCR sequence.

[0176] In some embodiments, the target chromosome is from a mouse, rat, rabbit, guinea pig, hamster, sheep, goat, donkey, cow, horse, camel, monkey, or chicken, and the target sequence comprises a mouse, rat, rabbit, guinea pig, hamster, sheep, goat, donkey, cow, horse, camel, monkey, or chicken homolog of the template sequence.

[0177] In some embodiments, the target sequence comprises a sequence of a mouse, rat, rabbit, guinea pig, hamster, sheep, goat, donkey, cow, horse, camel, monkey, or chicken gene. Mouse, rat, rabbit, guinea pig, hamster, sheep, goat, donkey, cow, horse, camel, monkey, or chicken gene are all contemplated within the scope of the present disclosure. Without wishing to be bound by theory, the introduction of human genes involved in disease development or potential therapeutic targets into model organisms such as mice, rats, rabbits, guinea pigs, hamsters, sheep, goats, donkeys, cows, horses, camels, monkeys, and chickens can facilitate disease research and development of appropriate treatments. In some embodiments, the target sequence comprises a mouse sequence that is homologous to a human template sequence. In some embodiments, the target sequence comprises a rat sequence that is homologous to a human template sequence. In some embodiments, the target sequence comprises a monkey sequence that is homologous to a human template sequence.

[0178] In some embodiments, the target sequence comprises an immunoglobin sequence, such as a mouse immunoglobulin sequence. In some embodiments, the target sequence comprises a mouse Igh sequence. Mouse Igh spans nucleotide positions 1112,947,269 to 116,248,693 on chromosome 12 of the GRCm39 assembly of the mouse genome. Those skilled in the art will appreciate that mouse Igh sequences having 5' and 3' boundaries that deviate from those described above, for example, at least 100 bp, 500 bp, 1,000 bp, 2,000 bp, 5,000 bp, 10,000 bp or more, are suitable template sequences.

[0179] In some embodiments, the target sequence comprises a mouse Igh variable region sequence. H , D H and J. HThe mouse Igh variable region sequence includes sequences encoding the mouse homolog of the 1-6 gene segment and intervening non-coding sequences. In some embodiments, the mouse Igh variable region sequence includes nucleotide positions 113,391,842 to 115,973,952 of chromosome 12 of the GRCm39 ensemble of the mouse genome. In some embodiments, the mouse Igh variable region sequence includes nucleotide positions 113,391,842 to 115,973,952 of chromosome 12 of the GRCm39 ensemble of the mouse genome, minus at least about 50 bp, 100 bp, 500 bp, 1,000 bp, 2,000 bp, 5,000 bp, 7,000 bp, 10,000 bp, 15,000 bp, 20,000 bp, or 50,000 bp from the 5' end or the 3' end, or both. In some embodiments, the human IGH variable region sequence comprises nucleotide positions 113,391,842-115,973,952 of chromosome 12 of the GRCm39 assembly of the mouse genome, and at least about 50 bp, 100 bp, 500 bp, 1,000 bp, 2,000 bp, 5,000 bp, 7,000 bp, 10,000 bp, 15,000 bp, 20,000 bp, or 50,000 bp of additional flanking sequence at the 5' or 3' end, or both. In some embodiments, the mouse Igh variable region sequence comprises nucleotide positions 113,391,842-115,973,952 of chromosome 12 of the GRCm39 assembly of the mouse genome, and one or more modifications thereto. Exemplary modifications include, but are not limited to, deletions, such as deletions of one or more V, D or J segments, insertions, such as insertions of markers, rearrangements, or combinations thereof. In some embodiments, the target sequence comprises a mouse Igl variable region sequence. In some embodiments, the target sequence comprises a mouse Igk variable region sequence. In some embodiments, the template sequence comprises a human IGL variable region sequence. In some embodiments, the template sequence comprises a human IGK variable region sequence.

[0180] In some embodiments, for example in embodiments in which little or no target chromosome sequence is deleted by the methods described herein, the target chromosome comprises a target site. The location of the target sequence is the location where the template sequence is inserted or where the template sequence is bound. Any location on the target chromosome can be a suitable location. In some embodiments, the target site comprises an endonuclease site for generating a double-strand break at the target site.

[0181] Recombinant chromosomes The present disclosure provides recombinant chromosomes produced by the methods described herein.

[0182] In some embodiments, the recombinant chromosome comprises a mouse chromosome comprising one or more humanized sequences. In some embodiments, the humanized sequences comprise one or more genes associated with a human disease or disorder, such as a gene associated with a genetic disease or disorder, or a cancer gene. In some embodiments, the recombinant chromosome comprises a rat chromosome comprising one or more humanized sequences. In some embodiments, the recombinant chromosome comprises a monkey chromosome comprising one or more humanized sequences.

[0183] In some embodiments, the recombinant chromosome comprises a mouse chromosome in which one or more immunoglobulin sequences have been humanized. In some embodiments, the immunoglobulin sequence comprises an IGH sequence, such as an IGH variable region. In some embodiments, the recombinant chromosome comprises mouse chromosome 12, in which the mouse Igh variable region has been replaced with a human IGH variable region from chromosome 14. In some embodiments, the mouse Igh variable region comprises a human IGH variable region from chromosome 14. H , D H and J. H 1-6 gene segments and intervening non-coding sequences. In some embodiments, the human IGH variable region comprises the V H , D H and J. H1-6 gene segments and intervening non-coding sequences. In some embodiments, the recombinant chromosome comprises mouse chromosome 12b, where a mouse IGh variable region comprising a nucleotide sequence from about 113,391,842 to 115,973,952 of chromosome 12 of the GRCm39 assemblies of the mouse genome is replaced with a human IGH variable region comprising a nucleotide sequence from about 105,862,994 to 106,811,028 of chromosome 14 of the GRCh38.p13 assemblies of the human genome. In some embodiments, the recombinant chromosome is mouse chromosome 6, where a mouse Igk variable region sequence comprises a human IGK variable region sequence in place of the mouse Igk variable region. In some embodiments, the mouse Igk variable region sequence comprises a mouse V k , and J k1-5 In some embodiments, the template sequence comprises a human IGK variable region sequence. In some embodiments, the human IGK variable region sequence comprises a human V k , and J k1-5 It includes sequences that code for gene segments, as well as intervening non-coding sequences.

[0184] Nucleic Acid Molecules, Plasmids and Vectors The present disclosure provides nucleic acid molecules for use in the methods described herein. A nucleic acid molecule, sometimes referred to as polynucleotide, refers to a chain of linked nucleotides that constitute a single molecule. The nucleic acid molecule of the present disclosure can be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The exemplary nucleic acid molecule of the present disclosure comprises homology arms that are specific to or adjacent to both target and template sequences, to facilitate the insertion of the template sequence into the target sequence, or the binding of the template sequence to the target sequence by double-strand break repair.

[0185] The present disclosure provides a nucleic acid molecule comprising a target chromosome and a template chromosome specific homology arm that promotes HDR-mediated chromosome rearrangement as described herein.In some embodiments, the nucleic acid molecule comprises a 5' homology arm comprising, from 5' to 3', a nucleotide sequence upstream of the 5' end of the target sequence, at least a first marker, and a 3' homology arm comprising, from 5' to 3', a nucleotide sequence downstream of the 3' end of the template sequence, at least a second marker, and a 3' homology arm comprising, from 5' to 3', a nucleotide sequence downstream of the 3' end of the target sequence.

[0186] The present disclosure provides a vector comprising the nucleic acid molecule described herein. According to the present disclosure, a vector is a nucleic acid molecule that can transport other nucleic acids that are linked to it. A plasmid is, for example, a type of vector. Vector sequences include sequences that are necessary for the production of vectors from host cells, such as bacteria, such as origin or replication, and selectable markers.

[0187] In some embodiments, the vector is a plasmid. In some embodiments, the plasmid comprises, from 5' to 3', a 5' homology arm comprising a nucleotide sequence upstream of the 5' end of the target sequence, at least a first marker, and a 3' homology arm comprising a nucleotide sequence upstream of the 5' end of the template sequence. In some embodiments, the plasmid comprises, from 5' to 3', a 5' homology arm comprising a nucleotide sequence downstream of the 3' end of the template sequence, at least a second marker, and a 3' homology arm comprising a nucleotide sequence downstream of the 3' end of the target sequence.

[0188] In some embodiments, the vector comprises a sequence of a homology arm located at or near the 5' end of the template sequence. In some embodiments, the homology arm is located upstream, i.e., 5', of the template sequence. In some embodiments, the vector comprises a sequence of a homology arm located at or near the 3' end of the template sequence. In some embodiments, the homology arm is located downstream, i.e., 3', of the template sequence. In some embodiments, the sequence of the template homology arm in the vector is identical or substantially identical to the sequence of the homology arm in the template sequence.

[0189] In some embodiments, the vector comprises a sequence of a homology arm located 5' of the target sequence or location, i.e., upstream of the target sequence or location, In some embodiments, the vector comprises a sequence of a homology arm located 3' of the target sequence or location, i.e., downstream of the target sequence or location.

[0190] Those skilled in the art will understand that even if there is some mismatch between the homology arm sequence of the vector and the equivalent (corresponding) sequence of the template chromosome or target chromosome, the vector will still facilitate the repair of the double-strand break from the vector to the template chromosome or target chromosome. For example, the homology arm sequence of the vector that is at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, or at least 99% identical to the equivalent sequence of the template chromosome is suitable for the method of the present disclosure.

[0191] In some embodiments, a nucleic acid molecule, plasmid, or vector described herein comprises one or more endonuclease sites.

[0192] In some embodiments, the present disclosure provides a first nucleic acid molecule comprising (i) a 5' homology arm comprising, from 5' to 3', a nucleotide sequence upstream of the 5' end of the target sequence, at least a first marker, and a 3' homology arm comprising a nucleotide sequence upstream of the 5' end of the template sequence; and (ii) a second nucleic acid molecule comprising, from 5' to 3', a 5' homology arm comprising, from 5' to 3', a nucleotide sequence downstream of the 3' end of the template sequence, at least a second marker, and a 3' homology arm comprising a nucleotide sequence downstream of the 3' end of the target sequence. In some embodiments, the first and second nucleic acid molecules are plasmids. In some embodiments, the first nucleic acid molecule comprises, from 5' to 3', a 5' homology arm comprising a nucleotide sequence upstream of the 5' end of the target sequence, a first endonuclease site, at least a first marker, a second endonuclease site, and a 3' homology arm comprising a nucleotide sequence upstream of the 5' end of the template sequence, wherein the first and second endonuclease sites overlap the homology arms such that the first and second endonuclease sites on the nucleic acid molecule and the corresponding endonuclease sites on the template and target chromosomes are cleaved by the same endonuclease. In some embodiments, the second nucleic acid molecule comprises, from 5' to 3', a 5' homology arm comprising a nucleotide sequence downstream of the 3' end of the template sequence, a third endonuclease site, at least a second marker, a fourth endonuclease site, and a 3' homology arm comprising a nucleotide sequence downstream of the 3' end of the target sequence, wherein the second and third endonuclease sites overlap the homology arm such that the third and fourth endonuclease sites on the nucleic acid molecule and the corresponding endonuclease sites on the template and target chromosomes are cleaved by the same endonuclease. In some embodiments, the first marker and the second marker are not the same marker. In some embodiments, the first marker on the first nucleic acid molecule comprises a combination of a selectable marker and a detectable marker. In some embodiments, the first marker comprises eGFP and puromycin resistance. In some embodiments, the second marker comprises a selectable marker.In some embodiments, the second marker comprises hygromycin resistance.

[0193] In some embodiments, a homology arm sequence on a nucleic acid molecule corresponds to a sequence located near a template sequence, a target sequence, or a target location. A homology arm that is within 0 bp, 5 base pairs (bp), 10 bp, 15 bp, 20 bp, 30 bp, 40 bp, 50 bp, 70 bp, 80 bp, 90 bp, 100 bp, 120 bp, 140 bp, 160 bp, 180 bp, 200 bp, or 250 bp of a template sequence, target sequence, or target location may be considered to be near said sequence.

[0194] In some embodiments, the length of the nucleic acid molecule homologous sequence corresponding to the template sequence or the target chromosomal sequence is between about 20 bp and 2,000 bp, between about 50 bp and 1,500 bp, between about 100 bp and 1,400 bp, between about 150 bp and 1,300 bp, between about 200 bp and 1,200 bp, between about 300 bp and 1,100 bp, between about 400 bp and 1,000 bp, or between about 500 bp and 900 bp, or between about 600 bp and 800 bp. In some embodiments, the length of the nucleic acid molecule homologous sequence is between about 400 and 1,500 bp. In some embodiments, the length of the nucleic acid molecule homologous sequence is between about 500 and 1,300 bp. In some embodiments, the length of the nucleic acid molecule homologous sequence is between about 600 and 1,000 bp.

[0195] In some embodiments, the nucleic acid molecule comprises a marker suitable for expression in a mammalian cell. In some embodiments, the marker is between homology arms in the nucleic acid molecule, whereby the marker is inserted into the target sequence. In some embodiments, the marker is a selectable marker. Suitable selectable markers include dihydrofolate reductase (DHFR), glutamine synthetase (GS), puromycin acetyltransferase, blasticidin deaminase, histidinol dehydrogenase, hygromycin phosphotransferase (hph), bleomycin resistance gene, aminoglycosidase phosphotransferase (neomycin resistance gene), and are described in further detail in Table 3 below.

[0196] In some embodiments, the marker comprises a detectable marker (or reporter). Detectable markers include, but are not limited to, enzymes that mediate luminescence reactions (luxA, luxB, luxAB, luc, rue, nluc), enzymes that mediate colorimetric reactions (lacZ, HRP), and fluorescent proteins such as green fluorescent protein (GFP), eGFP, yellow fluorescent protein (YFP), red fluorescent protein (RFP), cyan fluorescent protein (CFP), blue fluorescent protein (BFP), dsRed, mCherry, tdTomato, near infrared fluorescent protein, etc. The selection of suitable detectable markers will be known to those skilled in the art.

[0197] The markers may be expressed using any suitable promoter known in the art, including, but not limited to, the cytomegalovirus early (CMV) promoter, the PGK promoter, and the EF1a promoter.

[0198] TIFF2024533683000012.tif110170

[0199] In some embodiments, for example, in those embodiments of the method in which two nucleic acid molecules are used, a first nucleic acid molecule having a first marker and a second nucleic acid molecule having a second marker, the first marker or the second marker comprises a fluorescent protein operably linked to a promoter capable of expressing the fluorescent protein in the cell. In some embodiments, the fluorescent protein comprises green fluorescent protein (GFP). In some embodiments, the first marker further comprises a selectable marker. In some embodiments, the second marker further comprises a selectable marker. In some embodiments, the selectable marker is selected from the group consisting of dihydrofolate reductase (DHFR), glutamine synthetase (GS), puromycin acetyltransferase, blasticidin deaminase, histidinol dehydrogenase, hygromycin phosphotransferase (hph), bleomycin resistance gene, and aminoglycoside phosphotransferase. In some embodiments, the first marker and the second marker are not the same selectable marker. In some embodiments, the first marker comprises GFP and puromycin acetyltransferase operably linked to a promoter capable of expressing GFP in the cell, and the second marker comprises hygromycin phosphotransferase.

[0200] Methods for generating double-strand breaks Provided herein are methods for generating double-stranded breaks in template and target chromosomes. The methods provided herein use repair pathways for double-stranded break repair in a cellular environment to facilitate transfer of large sequences between chromosomes.

[0201] Any method known in the art for generating double stranded breaks in DNA sequences, and the repair pathways that repair those double stranded breaks, is contemplated within the scope of this disclosure.

[0202] In some embodiments, the double-stranded breaks in the template chromosome and the target chromosome are generated using one or more endonucleases. In some embodiments, the endonuclease also cleaves one or more nucleic acid molecules that contain homology arms used in the methods described herein. In some embodiments, the one or more endonucleases are selected from the group consisting of CRISPR / Cas endonucleases and one or more guide nucleic acids (gNAs), one or more zinc finger nucleases (ZFNs), or one or more transcription activator-like effector nucleases (TALENs). In some embodiments, the double-stranded breaks in the template chromosome and the target chromosome are generated using one or more CRE recombinases to generate chromosomal rearrangements.

[0203] Different molecules can introduce double-strand breaks and / or single-strand breaks in genomic nucleic acid.Nucleases of the present disclosure include, but are not limited to, homing endonucleases, restriction enzymes, zinc finger nucleases or zinc finger nickases, meganucleases or meganickases, transcription activator-like effector (TALE) nucleases (Kaids), particularly nucleic acid guided nucleases or nickases, such as RNA guided nucleases, DNA guided nucleases, megaTAL nucleases, BurrH nucleases, their modified or chimeric or mutant forms, and their combinations.RNA guided nucleases or RNA guided nickases are optionally part of CRISPR-based systems.

[0204] Nucleases can cleave phosphodiester bonds between monomers of nucleic acids. Many nucleases are involved in DNA repair by recognizing the damaged site and cleaving it from the surrounding DNA. These enzymes may be part of a complex. Endonucleases are nucleases that act on the central region of the target molecule. Deoxyribonucleases act on DNA. Many nucleases involved in DNA repair are not sequence specific. However, in the context of the present application, sequence specific nucleases are preferred. In some embodiments, the sequence specific nucleases are specific for fairly large strings of nucleotides in the target genome, for example 10 or more nucleotides, or 15, 20, 25, 30, 35, 40, 45, or 50 or more nucleotides, with target sequences in the target genome ranging from 5 to 50, 10 to 50, 15 to 50, 15 to 40, or 15 to 30 nucleotides being preferred. The larger such "recognition sequence" is, the fewer target sequences in the genome, and the more specific the incision that the nuclease makes in the genome, i.e., the more site-specific the incision. Site-specific nucleases generally have fewer than 10, 5, 4, 3, 2, or even only one (1) target site in the genome. Nucleases that are modified to modify genomic nucleic acid, such as to cleave a specific genomic target sequence, are referred to herein as modified nucleases. CRISPR-based systems are one type of modified nuclease. However, such modified nucleases can be based on any nuclease described herein.

[0205] Endonucleases that recognize sequences larger than 12 base pairs are called meganucleases. Meganucleases / nickases are endodeoxyribonucleases characterized by large recognition sites (e.g., double-stranded DNA sequences of 12-40 base pairs, such as 20-40 or 30-40 base pairs).

[0206] "Homing endonucleases" are a type of meganuclease, a double-stranded DNAse with a large asymmetric recognition site and coding sequence, usually integrated into an intron or intein. Homing endonucleases cut at very few, and sometimes singular, positions in the genome because their recognition sites are extremely rare (see also WO2004067736, US Patent No. 8,697,395 B2).

[0207] Zinc finger nucleases / nickases (ZFNs) are artificial restriction enzymes generated by fusing a zinc finger DNA binding domain to a DNA cleavage domain. The zinc finger domain can be engineered to target a specific DNA sequence of interest.

[0208] RNA-guided nucleases / nickases, particularly endonucleases, include, for example, Cas9 and Cpf1. CRISPR systems have already been described in detail. Any CRISPR-based system is part of the present disclosure. When another RNA-guided endonuclease(s) is used, suitable guide RNA, sgRNA or crRNA or other suitable RNA sequence can be used that interacts with the RNA-guided endonuclease and targets it to a genomic target site in genomic nucleic acid.

[0209] As used herein, the term "CRISPR-associated protein" or "CRISPR / Cas" protein refers to a nucleic acid guided (guided) DNA endonuclease associated with the CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) type II adaptive immune system found in certain bacteria, such as Streptococcus pyogenes. CRISPR / Cas proteins, such as Cas9, are not limited to the wild-type (wt) proteins found in bacteria. CRISPR / Cas proteins encompassing mutations to the wild-type CRISPR / Cas sequence or derivatives thereof are contemplated within the scope of this disclosure. The original type II CRISPR system from Streptococcus pyogenes includes the Cas9 protein and a guide RNA consisting of two RNAs: a mature CRISPR RNA (crRNA) and a trans-acting RNA (tracrRNA) that is partially complementary. Cas9 unwinds the foreign DNA and checks for sites complementary to the 20 base pair spacer region of the guide RNA. Cas9 targeting is simplified, and most Cas-based systems are designed to require only one or two chimeric guide RNAs resulting from the fusion of a crRNA and a tracrRNA, or a single guide RNA (chiRNA, often also simply called guide RNA, gRNA, or sgRNA). The spacer region can be engineered as needed.

[0210] As used herein, the term "Cas9 coding sequence" refers to a polynucleotide that can be transcribed and / or translated according to a genetic code functional in a host cell / host mammal to produce a Cas9 protein. The Cas9 coding sequence can be DNA (such as a plasmid) or RNA (such as an mRNA).

[0211] As used herein, the term CRISPR / Cas ribonucleoprotein refers to a protein / nucleic acid complex consisting of a CRISPR / Cas protein and an associated guide nucleic acid. For example, Cas9 ribonucleoprotein refers to Cas9 in complex with an associated guide RNA.

[0212] In some embodiments, the nuclease is an RNA-guided nuclease.Non-limiting examples of RNA-guided nucleases, including nucleic acid-guided nucleases, for use in the present disclosure include CasI, CasIB, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, CasX, CasY, Cas12a (Cpf1), Cas12b, Cas13a, CsyI, Csy2, Csy3, CseI, Cse2, CscI, Csc2, Csa5, Csn2, Csm2, These include, but are not limited to, Csm3, Csm4, Csm5, Csm6, CmrI, Cmr3, Cmr4, Cmr5, Cmr6, CsbI, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, CsfI, Csf2, Csf3, Csf4, Cms1, C2c1, C2c2, C2c3, or homologs, orthologs, or modified forms thereof.

[0213] "megaTAL nuclease / nickase" refers to an engineered nuclease that includes an engineered TALE DNA binding domain and an engineered meganuclease or engineered homing endonuclease. The TALE DNA binding domain can be designed to bind to DNA at almost any locus of a nucleic acid sequence in a genome, and when such a DNA binding domain is fused to an engineered meganuclease, it can cleave the target sequence. Exemplary designs of megaTAL nucleases and TALE DNA binding domains are disclosed, for example, in Boissel et al., (MegaTALs: a rare-cleaving nuclease architecture for therapeutic genome engineering (2013), Nucleic Acids Research 42 (4):2591-2601), and references cited therein, all of which are incorporated herein by reference in their entirety. MegaTAL nucleases optionally include one or more linkers and / or additional functional domains, such as a C-terminal domain (CTD) polypeptide, an N-terminal domain (NTD) polypeptide, an end-processing enzyme domain of an end-processing enzyme exhibiting 5-3' exonuclease or 3-5' exonuclease, or other non-nuclease domains, such as a helicase domain.

[0214] Transcription activator-like effector (TALE) nucleases / nickases are restriction enzymes that can be engineered to cleave specific sequences in DNA. Transcription activator-like effectors (TALEs) can be designed to bind to virtually any desired DNA sequence, and when combined with a DNA cleavage domain, can cleave DNA at a specific location.

[0215] A "TALE DNA-binding domain" is the DNA-binding portion of a transcription activator-like effector (TALE or TAL effector) that mimics plant transcriptional activators to manipulate the plant transcriptome. In some embodiments, contemplated TALE DNA binding domains are designed from de novo or naturally occurring TALEs, such as AvrBs3 from Xanthomonas campestris pv vesicatoria, Xanthomonas gardneri, Xanthomonas translucens, Xanthomonas axonopodis, Xanthomonas perforans, Xanthomonas alfalfa, Xanthomonas citri, Xanthomonas euvesicatoria, and Xanthomonas oryzae, and AvrBs4 from Ralstonia spp. Exemplary TALE proteins for deriving and engineering DNA binding domains include, but are not limited to, brg11 and hpx17 from Ralstonia solanacearum. Exemplary TALE proteins for deriving and engineering DNA binding domains are disclosed in U.S. Patent No. 9,017,967 and the references cited therein, all of which are incorporated herein by reference in their entireties.

[0216] "BurrH-nuclease" refers to a fusion protein with nuclease activity that contains modular base-per-base specific nucleic acid binding domains (MBBBDs). These domains are derived from a protein from the bacterial endosymbiont Burkholderia Rhizoxinica or other similar proteins identified from marine organisms. By combining different modules of these binding domains, modular base-per-base binding domains can be designed with binding properties for specific nucleic acid sequences, such as DNA binding domains. Such designed MBBBDs can cleave DNA at almost any locus of a nucleic acid sequence in a genome by fusing them with a nuclease catalytic domain. Exemplary designs of BurrH-nucleases and MBBBDs are disclosed in WO2014 / 018601 and US2015225465 A1, and references cited therein, all of which are incorporated herein by reference in their entirety.

[0217] A related aspect of the present disclosure provides a nucleic acid molecule, such as a vector, suitable for generating a CRISPR / Cas-mediated double-strand break (DSB) in a cell. In some embodiments, the vector comprises a sequence encoding a CRISPR / Cas protein, such as Cas9, and a guide nucleic acid (Cas9 single guide RNA, or sgRNA), operably linked to a suitable promoter for expression in the cell, and further comprises other vector components, such as an origin or replication and a selectable marker. In some embodiments, the cell is an embryonic stem cell or an embryonic hybrid stem cell as described herein.

[0218] According to the present disclosure, homologous recombination is promoted by the double-strand break (DSB) generated by endonuclease.In some embodiments, the endonuclease comprises CRISPR / Cas9 and one or more single guide RNAs (abbreviated as "sgRNA" or "gRNA").Those skilled in the art can select the flanking sequences of template sequence and target sequence, or the guide RNA with targeting sequence at the target position, as described above for endonuclease site.

[0219] In some embodiments, the enzyme can be introduced by introducing a nucleic acid molecule, such as a vector(s) or coding sequence encoding a CRISPR / Cas protein, and one or more sgRNA(s). In some embodiments, the vector or coding sequence encoding a CRISPR / Cas protein is a CRISPR / Cas mRNA. In some embodiments, the vector or coding sequence encoding a CRISPR / Cas protein is a vector, such as a plasmid, that includes a DNA sequence encoding a CRISPR / Cas protein and a gRNA. In some embodiments, the CRISPR / Cas protein is Cas9.

[0220] In certain embodiments, the CRISPR / Cas protein to be isolated can be directly introduced into cells (e.g., zygotes or ES cells by microinjection or electroporation). The CRISPR / Cas protein can be in the form of a CRISPR / Cas ribonucleoprotein, which is a CRISPR / Cas protein / gNA (guide nucleic acid) complex. Alternatively, the CRISPR / Cas protein can be gNA-free, in which case the CRISPR / Cas protein and one or more gNAs are co-introduced into zygotes or ES cells, allowing the formation of a CRISPR / Cas protein / gNA complex in situ in the cell. In some embodiments, the CRISPR / Cas protein and gNA are encoded by a vector and introduced into the cell by transfection, electroporation or transduction. In some embodiments, the CRISPR / Cas protein is Cas9.

[0221] In order to function as an endonuclease for use in the methods of the present disclosure, it is necessary for the CRISPR / Cas proteins to form a functional complex with the gRNA.

[0222] In some embodiments, multiple gNAs are used, each of which targets a specific CRISPR / Cas cleavage site.For example, a total of four gNAs can be used, two of which have targeting sequences specific to the gNA target sequence on both sides of the template sequence, and two of which have targeting sequences specific to the gNA target sequence on both sides of the target sequence.Alternatively, three gNAs can be used, one of which has targeting sequences specific to the gNA target sequence at the target position, and two of which have targeting sequences specific to the gNA target sequence on both sides of the template sequence.As yet another example, two gNAs can be used, one of which has targeting sequences specific to the gNA target sequence adjacent to the template sequence, and one of which has targeting sequences specific to the gNA target sequence adjacent to the target sequence.

[0223] Preferably, regardless of the number of gNAs used to create a DSB, in one embodiment, each of the gNAs is independently selected based on its proximity to the 5' and 3' ends of the template and target sequences, or to the target position.

[0224] Selection and design of gNAs can be performed using well-known principles or online tools based on user inputs such as target genome and sequence type. Generally, in the case of Cas9, gRNAs are short synthetic RNAs composed of a "scaffold" sequence required for Cas9 binding and a user-defined ~20 nucleotide "spacer" or "targeting" sequence that defines the genomic target to be bound or modified by the targeting sequence. For simplicity, "gRNA targets Cas9 cleavage site" refers to the fact that the spacer or targeting sequence of the gRNA is designed to bind to the genomic target sequence and cleave at the cleavage site.

[0225] Guide nucleic acids, including gRNAs and gDNAs according to the present disclosure, can be anywhere from 10 nucleotides in length, including 10-50 nucleotides, 10-40, 10-30, 10-20, 15-25, 16-24, 17-23, 18-22, 19-21, and 20 nucleotides in length.

[0226] Preferably, the targeting sequence is sufficiently unique so that it theoretically binds to a unique (relative to the rest of the genome) genomic target sequence. The target sequence should be immediately upstream (or 5') of the protospacer adjacent motif (or "PAM" sequence). The PAM sequence is essential for target binding, and the exact sequence varies depending on the type of Cas9. In the most widely used Streptococcus pyogenes Cas9, the PAM sequence is 5'-NGG-3' (where "N" represents any of the four standard nucleotides). Other PAM sequences for additional Cas9s in different species are known in the art. See the exemplary PAM sequences listed in Table 4 below.

[0227] TIFF2024533683000013.tif68170

[0228] The Cas9-gRNA complex binds to the target genomic sequence with a PAM, but Cas9 cleaves the target genomic sequence only if there is sufficient homology between the gRNA spacer and the target genomic sequence. The end result of Cas9-mediated DNA cleavage is a double-stranded break (DSB) within the target genomic sequence, with the cleavage site approximately 3-4 nucleotides upstream of the PAM sequence.

[0229] In some embodiments, double-stranded breaks are generated on both sides of the target sequence or on both sides.For example, in embodiments where the target chromosome comprises a target position, such as the position where the template sequence is inserted with little or no deletion of the target chromosome, double-stranded breaks are generated at the target position.Exemplary target positions include the cleavage sites of any of the nucleases described herein.As a further example, in embodiments where the target chromosome comprises a target sequence, such as the sequence that is replaced or deleted by the insertion of the template sequence, double-stranded breaks are generated on both sides of the target sequence (i.e., both 5' and 3' of the target sequence).

[0230] In certain embodiments, the cleavage site of any selected endonuclease, e.g., a gNA targeting sequence, is within about 10 bp, about 20 bp, about 30 bp, about 50 bp, about 70, about 100 bp, about 200 bp, about 300 bp, about 400 bp, or about 500 of the target sequence or position.

[0231] In certain embodiments, the cleavage site of a selected endonuclease, e.g., a gNA targeting sequence, is within about 100 bp, about 200 bp, about 300 bp, about 400 bp, about 500 bp, about 600 bp, about 700 bp, about 800 bp, about 900 bp, about 1,000 bp, about 1,100 bp, about 1,200 bp, about 1,300 bp, about 1,400 bp, about 1,500 bp, about 1,600 bp, about 1,700 bp, about 1,800 bp, about 1,900 bp, or about 2,000 bp of the template sequence.

[0232] In some embodiments, the double strand break is repaired by at least one DNA repair pathway selected from the group consisting of resection, mismatch repair (MMR), nucleotide excision repair (NER), base excision repair (BER), canonical non-homologous end joining (canonical NHEJ), alternative non-homologous end joining (ALT-NHEJ), canonical homology directed repair (canonical HDR), alternative homology directed repair (ALT-HDR), microhomology-mediated end joining (MMEJ), blunt end joining, synthesis-dependent microhomology-mediated end joining, single strand annealing (SSA), Holliday junction model or double strand break repair (DSBR), synthesis-dependent strand annealing (SDSA), single strand break repair (SSBR), translection synthesis repair (TLS), interstrand crosslink repair (ICL), and DNA / RNA processing.

[0233] Recovery of recombinant chromosomes The present disclosure provides methods for recovering recombinant chromosomes described herein and for introducing the recombinant chromosomes into a cellular environment suitable for downstream applications. In some embodiments, recovery of recombinant chromosomes described herein comprises microcell fusion (MMCT).

[0234] Micronuclear cell fusion (MMCT) is a technique for fusing micronuclear cells prepared from donor cells with recipient cells. This technique allows specific (foreign) DNA (e.g., chromosomes) in donor cells to be introduced into recipient cells. Micronuclear cells are usually prepared by treating donor cells with colcemid, but other methods may be used and are contemplated within the scope of the present disclosure.

[0235] An exemplary MMCT protocol includes culturing cells containing recombinant chromosomes in a cell culture medium containing at least one micronucleus inducer under conditions sufficient to induce micronucleus formation to generate micronucleated cells, and recovering the micronucleated cells. Exemplary micronucleus inducers include, but are not limited to, microtubule polymerization inhibitors, microtubule depolymerization inhibitors, and spindle checkpoint inhibitors. Exemplary micronucleus inducers known in the art include, but are not limited to, colcemid, colchicine, vincristine, or combinations thereof. For example, cells can be treated with 0.05 μg / mL to 0.25 μg / mL to induce micronucleus formation.

[0236] Micronucleated cells can be harvested using any suitable method known in the art, such as centrifugation or filtration.

[0237] Thus, the present disclosure provides a method for recovering recombinant chromosomes comprising exposing cells to colcemid under conditions sufficient to induce micronucleus formation and recovering micronucleated cells using centrifugation.

[0238] In some embodiments, the recombinant chromosome contains one or more markers, such as selectable or detectable markers introduced during recombination of the chromosome with the template sequence, which can be used to track the recombinant chromosome and to select cells containing the recombinant chromosome after fusion with the micronucleated cells described above.

[0239] Thus, the present disclosure provides a method of generating embryonic stem cells, comprising: (a) fusing a micronucleated cell containing a recombinant chromosome produced by the method of the present disclosure with an ES cell, where (i) the ES cell contains a chromosome homologous to the recombinant chromosome, the homologous chromosome containing a first fluorescent protein operably linked to a promoter capable of expressing the fluorescent protein in the ES cell, (ii) at least a subset of the micronucleated cells contains a recombinant chromosome, the recombinant chromosome containing a second fluorescent protein different from the first fluorescent protein, the second fluorescent protein operably linked to a promoter capable of expressing the fluorescent protein in the ES cell; (b) selecting the ES cells that express both the first and second fluorescent proteins; (c) culturing the ES cells selected in step (c) until the homologous chromosome is lost by at least a subset of the ES cells; and (d) selecting the ES cells that express the second fluorescent protein and do not express the first fluorescent protein. In some embodiments, the ES cells are mouse, rat, rabbit, guinea pig, hamster, sheep, goat, donkey, cow, horse, camel, chicken or monkey ES cells. In some embodiments, the ES cells are mouse ES cells. In some embodiments, the ES cells are rat ES cells. In some embodiments, the ES cells are monkey ES cells.

[0240] Although the method for generating embryonic stem cells described above uses two different fluorescent proteins as markers, one skilled in the art will understand that other markers are suitable, so long as the marker on the recombinant chromosome and the marker on the homologous chromosome are different. For example, two different selectable markers as described herein can be used. Also, two different surface molecules that can be recognized by labeled antibodies can be used, or can be attached to a selectable marker such as gold particles that allow selection by centrifugation. As a further example, in addition to fluorescent proteins as markers, puromycin and hygromycin / thymidine kinase (TK) markers can be used for positive-negative selection in this step. When thymidine kinase is expressed in the presence of certain thymidine analogs, these analogs are converted into toxic compounds that kill cells. For example, a puromycin resistance marker and a hygromycin / TK marker can be knocked in to two chromosomes at the same location, and double-positive single clones are selected by culturing in puromycin and hygromycin. After several days of culture, puromycin and thymidine kinase are used to select clones that have lost one copy of the chromosome carrying the hygromycin / TK marker.

[0241] In some embodiments, the method of generating embryonic stem cells includes (a) fusing micronucleated cells containing a recombinant chromosome generated by the method of the present disclosure to ES cells, where (i) the ES cells contain a chromosome homologous to the recombinant chromosome, the homologous chromosome containing a first marker, (ii) at least a subset of the micronucleated cells contain the recombinant chromosome, the recombinant chromosome containing a second marker different from the first marker; (b) selecting ES cells expressing both the first and second markers; (c) culturing the ES cells selected in step (c) until the homologous chromosome is lost by at least a subset of the ES cells; and (d) selecting ES cells expressing the second marker and not expressing the first marker.

[0242] The micronucleated cells can be fused with ES cells using any suitable method, including inter alia electrofusion, virus-induced fusion, and chemically induced fusion, for example by the addition of PEG1000 to the cells.

[0243] Considering the inherent instability of trisomies generated by the above-mentioned recombinant chromosome recovery methods, culturing the cells generated by fusion with micronucleated cells for at least 5 days, at least 7 days, at least 10 days, or at least 14 days will be sufficient to obtain cells that have lost the homologous chromosome corresponding to the recombinant chromosome. Alternatively, a selection scheme may be employed using a negatively selectable marker, e.g., a marker located on the homologous chromosome whose expression kills the cell when exposed to a selective regimen. In some embodiments, the selection of cells in steps (b) and (d) comprises fluorescence-activated cell sorting (FACS). For example, the cells may be FACS-sorted cells that express a second fluorescent protein used to label the recombinant chromosome, but not express the first fluorescent protein used to label the homologous chromosome.

[0244] cell The present disclosure provides cells for use in the methods of the present disclosure. In some embodiments, the cells include embryonic stem (ES) cells, hybrid embryonic stem (EHS) cells, or zygotic cells. The present disclosure also provides cells that include recombinant chromosomes produced by the methods of the present disclosure. The present disclosure provides methods for isolating, fusing, and culturing the cells described herein.

[0245] Thus, the present disclosure provides methods for fusing cells to generate the EHS cells described herein. Cell fusion can be accomplished by chemical, biological, or physical means. Examples of these techniques include polyethylene glycol (PEG) fusion, fusogenic virus fusion, and electrofusion, respectively.

[0246] The ES cells used in the methods of the present disclosure can be obtained from various sources, and can be primary isolated ES cells or artificially or naturally occurring ES cell lines. ES cells can also be first genetically modified to introduce useful traits, such as the expression of one or more markers, before or after cell fusion to generate the EHS cells of the present disclosure, or before or after the methods described herein.

[0247] One commonly used technique is chemical fusion, for example using PEG. This technique has been particularly successful in generating hybridomas. The fusion probability is improved by exposing the cells to a very short, strong electric field. Chemical agents can be used to effectively link and bring into close proximity the desired type of cell pair (i.e., two types of EH cells) in suspension before exposure to the electric field.

[0248] Cell electrofusion involves bringing cells into close proximity and exposing them to an alternating electric field. Under appropriate conditions, the cells are forced together, causing fusion of the cell membranes, and subsequently forming fused or hybrid cells. Cell electrofusion and devices therefor are described, for example, in U.S. Pat. Nos. 4,441,972, 4,578,168, and 5,283,194, and International Patent Application No. PCT / AU92 / 00473. In general, the method involves selecting cells and placing them in a fluid-filled chamber adapted as a cell fusion chamber. Individual cell pairs may be involved in the fusion process, i.e., single cell fusion, or bulk fusion may occur with two populations, each containing two or more cells. Bulk fusion can be mini-bulk fusion, involving about 2 to about 1000 cells, or macro-bulk fusion, involving more than about 1000 cells. Fusion can be promoted by chemical means, such as the presence of PEG, biological means, such as the presence of a fusagenic virus, or electrical means, such as electrofusion. Fusion can also combine these techniques. Cells can also be treated with cytokines, such as interleukin 3 (IL-3), to promote fusion.

[0249] Following cell fusion, what is known as a fusate cell or hybrid cell is obtained, which contains the nuclei of at least two cells from the cells involved in the fusion, enclosed in a fused lipid bilayer. The nuclei of the cells fuse, resulting in a hybrid cell with an abnormal number of chromosomes (it may be tetraploid or have fewer or more chromosomes). Hybrid cells have the ability to divide and grow under appropriate culture conditions.

[0250] In some embodiments, EHS cells are generated by electrofusion.For example, human and mouse, human and rat, or human and monkey ES cells can be fused by electrofusion.In some embodiments, two EHS cells from two species selected from the group consisting of human, mouse, rat, rabbit, guinea pig, hamster, sheep, goat, donkey, cow, horse, camel, chicken and monkey undergo electrofusion to generate EHS cells.

[0251] Generally, once fusion has occurred, the resulting hybrid cells are harvested in a suitable rich medium before being expanded in culture for use in the disclosed method. The harvest medium should contain factors that allow the recovery of the cell fusion after fusion stress. Such supplements may contain a high percentage of fetal calf serum, for example 20%.

[0252] The hybrid cells produced by cell fusion can contain unique cell surface markers that are useful for selecting these cells and for monitoring the fusion event.

[0253] In some embodiments, the cells of the present disclosure comprise one or more genetic modifications, such as the introduction of markers as described herein. Genetic modifications can be carried out by any suitable means known in the art. For example, cells can be modified by transfection, transduction, electroporation, lipofection, etc.

[0254] As used herein, transfection refers to the introduction of naked or purified nucleic acid, or nucleic acid including a vector carrying a particular nucleic acid, into cells, particularly eukaryotic cells including mammalian cells. In the context of this disclosure, known transfection methods can be employed. Some of these methods involve increasing the permeability of biological membranes to allow the nucleic acid to enter the cells. Prominent examples are electroporation, microporation and lipofection. These methods may be used by themselves or may be used to selectively increase the flux of nucleic acid into the host cells using sound waves, electromagnetic waves, thermal energy, chemical permeation enhancers, pressure, etc. Other transfection methods are also within the scope of this disclosure, including, for example, carrier-based transfection (also called transduction) including lipofection or viruses, and chemical-based transfection. However, any method of introducing nucleic acid into cells can be used. Transiently transfected cells retain / express the transfected RNA / DNA for a short period of time and do not inherit it. Stably transfected cells continue to express and inherit the transfected DNA: the foreign nucleic acid has been integrated into the cell's genome.

[0255] Many viruses have been used as gene transfer vectors or as the basis for the preparation of gene transfer vectors, including papovaviruses, adenoviruses, vaccinia viruses, adeno-associated viruses, lentiviruses, Sindbis viruses, Semliki Forest viruses, and retroviruses of avian and human origin.

[0256] There are chemical methods of gene transfer, such as calcium phosphate co-precipitation, mechanical methods, e.g., microinjection, transfer via membrane fusion via liposomes, direct DNA uptake and receptor-mediated DNA transfer. Viral-mediated gene transfer can be combined with direct in vivo gene transfer using liposome delivery to target viral vectors to specific cells. Alternatively, retroviral vector-producing cell lines can be injected into specific tissues. Injecting producer cells provides a continuous source of vector particles.

[0257] The present disclosure provides a method of culturing the cells of the present disclosure. In the embodiments described herein, many stem cell medium culture or growth environments are contemplated, such as defined medium, conditioned medium, feeder-free medium, serum-free medium, etc. As used herein, the term "growth environment" refers to the environment in which undifferentiated or differentiated stem cells (e.g., embryonic stem cells) are grown in vitro. Characteristics of the environment include the medium in which the cells are cultured and the supporting structure, if present (such as a substrate on a solid surface). Methods of culturing or maintaining cells are also described in PCT / US2007 / 062755, U.S. Application No. 11 / 993,399, and U.S. Application No. 11 / 875,057.

[0258] Basal cell culture media are known in the art and commercially available. Exemplary basal cell culture media include, but are not limited to, DMEM, CMRL or RPMI-based media.

[0259] The cell culture medium used in the cell culture method of the present disclosure can contain serum or can be serum-free.Cell culture medium can also contain one or more supplements or other medium components known in the art, such as B27 supplement, insulin, glucose, growth factors such as EGF and FGF, cytokines, etc.

[0260] The term "feeder cells" refers to a culture of cells that grow in vitro and secrete at least one factor into the medium, and can be used to support the growth of another cell of interest in culture. As used herein, the term "feeder cell layer" can be used interchangeably with the term "feeder cells." Feeder cells can comprise a monolayer, where the feeder cells cover the surface of a culture dish in a complete layer and then grow on top of each other, or can comprise clusters of multiple cells. In a preferred embodiment, the feeder cells comprise an adherent monolayer.

[0261] Similarly, embodiments in which ES or EHS cell cultures or aggregate suspension cultures are cultured in a defined condition or culture system without the use of feeder cells are "feeder-free." Note that feeder-free methods are also described in U.S. Pat. No. 6,800,480. In some embodiments, ES or EHS cells can be cultured in a two-dimensional or three-dimensional environment. In U.S. Pat. No. 6,800,480, the extracellular matrix is ​​prepared by culturing fibroblasts, lysing the fibroblasts in situ, and washing what remains after lysis. Alternatively, in U.S. Pat. No. 6,800,480, the extracellular matrix can be prepared from isolated matrix components or combinations of components selected from collagen, placental matrix, fibronectin, laminin, merosin, tenascin, heparin sulfate, chondroitin sulfate, dermatan sulfate, aggrecan, biglycan, thrombospondin, vitronectin, and decorin.

[0262] In some embodiments, the culture method or culture system is free of animal sources, hi other embodiments, the culture method is xeno-free.

[0263] The present disclosure includes differentiating ES cells containing the recombinant chromosomes described herein into various cell types for use in various downstream applications. ES cells can be induced to differentiate into various cell types in vitro using a variety of strategies, including supplementing cell culture media with exogenous biochemical compositions that directly recapitulate endogenous developmental cell signals and induce cell-specific differentiation. Strategies for differentiating ES cells are discussed in Vazin and Freed, Restor Neurol Neurosci (2010) 28(4):589-603, the contents of which are incorporated herein by reference.

[0264] For example, a population of ES or EHS cells can be further cultured in the presence of specific supplemental growth factors to obtain a population of cells that develop or will develop into a different cell lineage, or to selectively reverse the development of a different cell lineage. The term "supplemental growth factor" is used in the broadest sense to refer to a substance that is effective to promote the growth of ES cells, maintain cell survival, stimulate cell differentiation, and / or stimulate the reversal of cell differentiation. Additionally, supplemental growth factors can be substances secreted into the medium by feeder cells. Such substances include, but are not limited to, cytokines, chemokines, small molecules, neutralizing antibodies, proteins, and the like. Growth factors also include intercellular signaling polypeptides that control the development and maintenance of cells and the morphology and function of tissues. In a preferred embodiment, the supplemental growth factors are selected from the group consisting of stem cell factor (SCF), oncostatin M (OSM), ciliary neurotrophic factor (CNTF) and interleukin-6 (IL-6) in combination with soluble interleukin-6 receptor (IL-6 R), fibroblast growth factor (FGF), bone morphogenetic protein (BMP), tumor necrosis factor (TNF), and granulocyte-macrophage colony-stimulating factor (GM-CSF).

[0265] The progression of stem cells to various pluripotent and / or differentiated cells can be monitored by determining the relative expression of genes characteristic of a particular cell type, or gene markers, compared to the expression of a second or control gene, e.g., a housekeeping gene. In some processes, the expression of a particular marker is determined by detecting the presence or absence of the marker. Alternatively, the expression of a particular marker can be determined by measuring the level at which the marker is present in one or more cells of a cell culture. In such processes, the measurement of marker expression can be qualitative or quantitative. One method of quantifying the expression of a marker produced by a marker gene is by using quantitative PCR (Q-PCR). Methods for performing Q-PCR are well known in the art. Other methods known in the art can also be used to quantify the expression of a marker gene. For example, the expression of a marker gene product can be detected using an antibody specific for the marker gene product of interest.

[0266] Transgenic animals The present disclosure provides transgenic animals, eg, transgenic mice, that contain the recombinant chromosomes of the present disclosure, and methods for making the same.

[0267] The selection of an appropriate method for generating transgenic animals from ES cells or zygotic cells containing the recombinant chromosomes described herein will vary from animal to animal, and will be known to one of skill in the art.

[0268] In an exemplary method, ES cells containing the recombinant chromosomes are incorporated into an embryo at the blastocyst stage, which is then implanted (transduced) into a pregnant or pseudopregnant female and allowed to mature, resulting in a chimeric animal. If germ cells are generated from the ES cells, the animal's offspring will be fully transgenic and will inherit the recombinant chromosomes.

[0269] In some embodiments, the transgenic animal is a mouse, rat, rabbit, guinea pig, hamster, sheep, goat, donkey, cow, horse, camel, chicken, or monkey.

[0270] In some embodiments, the transgenic animal is a mouse. In some embodiments, the generation of the transgenic mouse comprises injection of ES cells into diploid blastocysts, nuclear transfer from ES cells into enucleated mouse embryos, or tetraploid embryo complementation.

[0271] In some embodiments, the method further comprises implanting the ES cells or zygotes into pseudopregnant females. In mice, pseudopregnant females are prepared by mating 6-8 week old female mice in natural estrus with vasectomized male mice. Zygotes processed for same-day implantation into pseudopregnant females can be removed from culture, placed in pre-warmed appropriate medium (such as M2 medium), and implanted through the oviduct into pseudopregnant females (e.g., 9-11 week old) 0.5 days after mating.

[0272] Once a recombinant chromosome has been inserted into a host mammal using the methods of the present disclosure, the presence of the chromosome can be verified in the resulting transgenic animal (e.g., mouse) or its progeny. Such verification typically involves one or more of the following: genotyping of animals potentially carrying the recombinant chromosome, polymerase chain reaction amplification of junction sequences, direct sequencing of specific DNA stretches (e.g., template sequences), and genetic mapping. Such techniques are well known in the art.

[0273] The present disclosure provides a transgenic mouse comprising the recombinant chromosome of the present disclosure. In some embodiments, the transgenic mouse comprises one or more humanized genes, such as any one of the genes listed in Tables 1 and 2. In some embodiments, the animal model comprises one or more humanized genes (e.g., 1, 2, 5, 10, 20, 50, 100 or more genes). In some embodiments, the transgenic mouse comprises all or part of a humanized immunoglobulin gene. In some embodiments, the transgenic mouse comprises all or part of a humanized TCR subunit gene.

[0274] In some embodiments of the transgenic mouse of the present disclosure, the mouse chromosome 12 comprises a sequence of a human IGH variable region in place of the mouse Igh variable region. H、 D H and J. H 1-6 gene segments and intervening non-coding sequences. In some embodiments, the human IGH variable region comprises the V H , D H and J. H In some embodiments, the recombinant chromosome is mouse chromosome 6, which contains a human IGK variable region sequence in place of the mouse Igk variable region. In some embodiments, the mouse Igk variable region sequence is a mouse V k , and J k1-5 In some embodiments, the template sequence comprises a human IGK variable region sequence. In some embodiments, the human IGK variable region sequence comprises a human V k , and J k1-5 It includes sequences that code for gene segments, as well as intervening non-coding sequences.

[0275] Purpose Downstream applications of cells and transgenic animals containing the recombinant chromosomes described herein are contemplated as within the scope of this disclosure.

[0276] Exemplary downstream applications include basic and applied research on animal models of human diseases and disorders using animal models (e.g., mice, rats, or monkeys) that are humanized with one or more human genes. Exemplary, but non-limiting, genes that can be humanized by replacing the homologous genes in the model animals with the homologous genes in humans are listed in Tables 1 and 2. Model animals of human diseases associated with chromosomal abnormalities (translocations, inversions, etc.) can also be generated using the methods described herein. For example, animal models that require large-scale chromosomal rearrangements of fragments over 300 kB, such as Duchenne muscular dystrophy (DMD) humanized mouse disease models, and animal models that require large-scale insertion or replacement of arrays of up to hundreds of genes are contemplated within the scope of the present disclosure.

[0277] In some embodiments, for example, where the Igh variable region of the animal is humanized, the transgenic animals of the present disclosure can be used to generate humanized antibodies. For example, such animals can be made to generate specific B cells that have human or humanized antibodies. In some embodiments, for example, where the Igk or Igl variable region of the animal is humanized, the transgenic animals of the present disclosure can be used to produce humanized antibodies.

[0278] In some embodiments, for example, in embodiments where a template sequence comprising an antibody or antigen-binding fragment thereof is inserted into a target chromosome, the transgenic animals of the present disclosure can be used to generate antibodies or antigen-binding fragments. For example, the transgenic animals can be used to generate single chain variable fragments (scFv), nanobodies, bispecific antibodies, multispecific antibodies, etc. Such antibodies can be used for research or therapeutic purposes.

[0279] Exemplary downstream applications include those in which the recombinant chromosomes are not incorporated into a transgenic animal, but instead, by way of example, ES cells containing the recombinant chromosomes can be differentiated into another cell type and used for research or therapeutic purposes.

[0280] kit The present disclosure provides kits comprising the nucleic acid molecules described herein. In some embodiments, the nucleic acid molecule is a vector, such as a plasmid.

[0281] In some embodiments of the kits of the present disclosure, the kits include cells, e.g., cryopreserved EHS cells, for use in the methods described herein. In some embodiments, the kits include the nucleic acid molecule and, optionally, instructions for use of the cells. EXAMPLES

[0282] Working Example Example 1: Establishment of Embryonic Hybrid Stem Cells (EHS) The overall goal of this study was to obtain mice with humanized variable domains of Igh and Igk genes. Humans and mice show high similarity in the sequence and expression of antibody genes, and the genomic organization of the heavy chain is also similar in humans and mice. Thus, by replacing approximately 3 MB of mouse genomic sequence containing all segments of VH, DH, and JH genes with approximately 1 MB of contiguous human genomic sequence containing the equivalent human gene fragments, a humanized version of the mouse Igh or Igk gene variable domains can be obtained (Figure 1).

[0283] The first step to create a humanized mouse Igh gene was to generate mouse embryonic hybrid stem (EHS) cells by fusing mouse embryonic stem (ES) cells with human ES cells, generating cells that carry both mouse and human Igh genes.

[0284] Recombinant mouse cells expressing a neomycin resistance gene under the control of a PGK promoter and recombinant human ES cells expressing an mCherry marker under the control of a CAG promoter were fused by electrofusion according to the standard method provided by the manufacturer of the electrofusion apparatus. The hybrid EHS cells were cultured in mouse ES cell medium containing G418 for 7 days, and viable cells were sorted by fluorescence-activated cell sorting (FACS) according to the expression level of mCherry (Figure 2). The positive cells were continuously cultured in mouse ES cell medium containing G418, and single-cell clones were isolated in separate wells for expansion. Genomic DNA was then extracted for each single-cell clone, and genotype analysis was performed. Specifically, PCR was performed using three sets of primers against the V, D, and J regions of the human immunoglobulin heavy chain (IGH) (Figure 3A) to confirm the presence of the target sequence in the EHS clones (Figure 3B). Only clones that had all three desired regions were kept for further experiments.

[0285] Example 2: Humanized chromosome engineering 2.1. Establishment of EHC by HDR-induced chromosomal rearrangement (HMCR) To obtain mouse embryonic hybrid stem (EHS) cells with a humanized variable domain of the Igh gene, we replaced ∼3 MB of the variable domain of the Igh gene on mouse chromosome 12 with ∼1 MB of the variable domain of the human IGH gene on human chromosome 14 by HDR-mediated chromosome rearrangement (HMCR; Figure 4A).

[0286] Two types of plasmids were designed to mediate the HMCR process and are shown in Figure 4A. The 5'HMCR plasmid was designed to replace the 5' end of the mouse Igh gene with that of human, and the 3'HMCR plasmid was designed to replace the 3' end of the mouse Igh gene with that of human. The 5'HMCR plasmid contained a 5' arm with homology to the 5' end of the mouse Igh gene, a 3' arm with homology to the 5' end of the human IGH gene, and a cassette of CMV-EGFP-polyA-PGK-Puromycin-poly inserted between the two homology arms. Similarly, the 3'HMCR plasmid contained a 5' arm with homology to the 3' end of the human IGH variable locus, a 3' arm with homology to the 3' end of the mouse IGH variable locus, and a cassette of PGK-Hygromycin-polyA inserted between the two homology arms (see Figure 4A). The length of the homology arms ranged from 600 bp to 1000 bp. At the same time, four plasmids containing Cas9 and sgRNA targeting the 5' and 3' ends of mouse and human Igh variable domains were also designed (see the zigzag mark in Figure 4A, sgRNA targeting sequences are listed in Table 7). These six plasmids were co-transfected as circular plasmids into EHS cells obtained in Example 1 by standard methods, and the resulting cells were cultured in mouse ES cell medium containing puromycin and hygromycin for 7 days. Surviving GFP-positive single clones were picked and further cultured.

[0287] Genotyping was performed to identify the desired single clones that had successfully undergone HMCR. For genotyping, four pairs of PCR primers were designed as shown in Figure 5A. In the first primer pair, the forward primer was designed upstream of the 5' homology arm of mouse Igh 5'HMCR plasmid, and the reverse primer was designed within the CMV promoter region (Figure 5A). For the second primer pair, the forward primer was within the puromycin gene of 5'HMCR plasmid, and the reverse primer was downstream of the 5' homology arm of human IGH, within the human IGH sequence (Figure 5A). In the third primer pair, the forward primer was upstream of the 3' homology arm of human IGH variable region, and the reverse primer was in the PGK promoter region of 3'HMCR plasmid (Figure 5A). In the last pair of primers, the forward primer was in the hygromycin gene of 3'HMCR plasmid, and the reverse primer was downstream of the 3' homology of 3'HMCR plasmid, within the mouse Igh variable domain (Figure 5A). PCR amplification was performed with each primer pair for each clone, and only clones whose PCR products were positive in all four genotyping tests were retained for further experiments. Of the 196 clones isolated in this step, 6 clones were identified as positive for all four PCR amplicons (Figure 5B).

[0288] To promote the expression of human IGH gene in EHS cells where HMCR was successful, we deleted the 3' selection marker from the genome of positive clones by homology-directed repair (HDR) (Figure 4A), although non-homologous end joining (NHEJ), microhomology-mediated end joining (MMEJ), and homology-mediated end joining (HMEJ) methods can also be used. Through the above-mentioned process, we successfully established recombinant humanized chromosomes (EHCs) in EHS cells, in which the variable domains encompassing the VH, DH, and JH1-6 gene segments of mouse Igh gene on mouse chromosome 12 were replaced with the equivalent human regions by HMCR.

[0289] The sequences of the plasmids used to mediate the HMCR process are shown in Tables 5 and 6 below.

[0290] TIFF2024533683000014.tif232170TIFF2024533683000015.tif228170TIFF2024533683000016.tif212170

[0291] TIFF2024533683000017.tif231170TIFF2024533683000018.tif226170TIFF2024533683000019.tif100170

[0292] TIFF2024533683000020.tif40170

[0293] sgRNA sequences with a PAM sequence (NGG) located on the non-target strand 3' of the sgRNA targeting sequence are provided in Table 7. The corresponding sgRNA targeting sequences without the PAM are provided as SEQ ID NOs:14-17.

[0294] 2.2.Establishment of EHCs by CRE-Loxp-mediated chromosomal rearrangement (CMCR) To obtain mouse EHS cells with humanized variable domains of the Igh gene, we replaced approximately 3 MB of the variable domain of the Igh gene on mouse chromosome 12 with approximately 1 MB of the variable domain of the IGH gene on human chromosome 14 by CRE-Loxp-mediated chromosomal rearrangement (CMCR; Figure 4B). Four plasmids were designed to mediate the CMCR process. The mouse Igh 5' (pCMV-GFP-BGH PolyA-Loxp) and 3' (BGH polyA-Loxp-511-Hygromycin-BGH polyA-PGK-BSD-BGH PolyA) plasmids were designed to insert into the 5' and 3' ends of the mouse Igh variable locus, respectively. At the same time, we designed to insert human IGH 5' (BGH polyA-Loxp-Puro-BGH PolyA-PGK-Neomycin-BGH PolyA) and 3' (pCMV-BGP-BGH PolyA-PGK-Loxp-511) plasmids at the 5' and 3' ends of the human IGH variable locus, respectively (Figure 5). The transfected EHS cells were cultured in mouse ES cell medium containing BSD and neomycin for 7 days. The surviving GFP- and BFP-double positive cells were picked and further cultured. Genotyping was performed to identify the desired single clones that successfully introduced the above plasmids. Cre was transfected into the EHS cells that had successfully undergone CMCR, and the successfully rearranged cells could survive in medium containing puromycin and hygromycin. The surviving cells were pocked for genotyping. To promote the expression of the human IGH gene in the EHS cells that had successfully undergone CMCR, the 3' selectable marker was then deleted from the genome (Figure 5). Through the above steps, we succeeded in establishing a recombinant humanized chromosome (EHC; a humanized version of the variable domain of the Igh gene on mouse chromosome 12) using CMCR of EHS cells.

[0295] Example 3: Chromosome replacement in mouse embryonic stem cells by microcell fusion After obtaining EHS cells carrying recombinant humanized chromosomes (EHCs) as described in Examples 1 and 2, the EHCs were then introduced into mouse ES cells by microcell fusion (MMCT) to establish mouse ES cells in which the variable domains of the Igh genes were humanized.

[0296] EHC-transduced EHS cells were treated with 0.2 μg / ml colcemid at 37°C for 48 hours. Prolonged mitotic arrest led to the formation of micronucleated cells, which were collected by centrifugation (Figure 6). At the same time, mouse ES cells expressing mCherry fluorescent marker on chromosome 12 were obtained (Figure 6). These cells were obtained by inserting a CMV-mCherry-polyA cassette into one copy of mouse chromosome 12.

[0297] The micronucleated cells were then hybridized with mouse ES cells by electrofusion, and the resulting cells were selected by FACS using GFP+ and mCherry+ markers to obtain mouse ES cells that were GFP+ and mCherry+. GFP+ indicated that the EHC was successfully introduced into the mouse ES cells, and the mCherry+ marker indicated that the cells also possessed mCherry+ chromosome 12. The positive cells were continuously cultured in mouse ES cell medium for 2 weeks, and mCherry- and GFP+ mouse ES cells, i.e., cells that had lost the extra chromosome 12, as indicated by mCherry+, were selected by FACS and cultured for 7 days. Single clones were isolated in separate wells for growth and karyotype analysis, and clones with the correct karyotype were retained. As a result, mouse ES cells with humanized variable regions of the Igh gene were obtained.

[0298] Example 4: Generation of Igh-humanized mice The mouse ES cells with humanized variable regions of the Igh gene obtained in Example 3 were injected into blastocysts of mice of the B6D2F1 (C57BL / 6 x DBA2) strain according to standard procedures. Alternatively, nuclear transfer or tetraploid embryo complementation could be used to generate humanized mice.

[0299] Injected blastocysts were implanted into the uterus of pseudopregnant ICR females at 2.5 days post-insemination (dpc). Igh humanized mice were identified by their GFP expression levels under a fluorescent stereomicroscope, and GFP+ mice were further analyzed.

[0300] Next, a series of PCR experiments were designed to verify the Igh-humanized mice. The first series of PCR experiments was designed to verify the integrity of the human IGH variable region. Five pairs of primers for different regions of the human IGH variable region were designed (Figure 7A, see arrows indicating PCR primers 1-10). In the Igh-humanized mice, PCR products were positive for all five pairs of PCR primers (Figure 7B). Primers were also designed upstream and downstream of the human IGH variable region (Figure 7A), but no products were observed in any PCR experiments in the Igh-humanized mice, whereas the right band of PCR products was observed in HEK293T (Figure 7B).

[0301] Fibroblasts were isolated from the tails of Igh-humanized mice and subjected to fluorescence in situ hybridization (FISH). The FISH results showed that chromosome 12 of Igh-humanized mice contained a fragment of human chromosome 14 (Figure 8A), indicating that the variable domain of the human IGH gene had been inserted in situ into mouse chromosome 12.

[0302] G-banding karyotype analysis was also performed to exclude abnormal chromosomes (Figure 8B).

[0303] Genomic DNA from the Igh-humanized mice was also extracted and subjected to whole-genome sequencing (WGS) analysis. The WGS sequences were mapped to a reference genome that includes all mouse chromosomes and human chromosome 14. The variable domain (V H , D H , J H All of the genomic regions (gene segments) were covered by whole-genome sequence reads. Furthermore, no off-target editing was observed in other genomic regions (Figures 9A-9B).

[0304] Example 5: Generation of Igk-humanized mice Using MASIRT, we obtained mice in which the variable domain of the Igk gene was humanized (Figure 10). Using a similar approach to the Igh gene described above, we also obtained Igk-humanized mice. To verify the Igk-humanized mice, we first performed PCR experiments to verify the integrity of the human IGK variable region. Five pairs of primers were designed at different loci of the human IGK variable region (Figure 11A), and the resulting Igk-humanized mice were positive for PCR products in all five experiments (Figure 11B). We also designed primers upstream and downstream of the human IGK variable region (Figure 11A), but in the resulting Igk-humanized mice, no products were observed in any PCR experiments, whereas the right band of the PCR product was observed in HEK293T (Figure 11B). Finally, the genomic DNA of the Igk-humanized mice was also extracted and whole genome sequencing (WGS) analysis was performed.

[0305] TIFF2024533683000021.tif229170TIFF2024533683000022.tif227170TIFF2024533683000023.tif231170

[0306] TIFF2024533683000024.tif230170TIFF2024533683000025.tif228170TIFF2024533683000026.tif139170

[0307] TIFF2024533683000027.tif38170

[0308] sgRNA sequences with a PAM sequence (NGG) located on the non-target strand 3' of the sgRNA targeting sequence are provided in Table 10. The corresponding sgRNA targeting sequences without the PAM are provided as SEQ ID NOs:28-31.

[0309] The whole genome sequence was mapped to a reference genome containing all mouse chromosomes and human chromosome 2. The results showed that the variable domain (V H and J. HAll of the genomic regions (gene segments) were covered by the whole genome sequence, and no off-target editing was observed in other genomic regions (Figure 12).

Claims

1. 1. A method of producing a recombinant chromosome, comprising: a. providing a cell comprising a target chromosome comprising a target sequence and a template chromosome comprising a template sequence; b. contacting the cells with: i. a first nucleic acid molecule comprising, from 5' to 3', a 5' homology arm comprising a nucleotide sequence upstream of the 5' end of the target sequence, at least a first marker, and a 3' homology arm comprising a nucleotide sequence upstream of the 5' end of the template sequence; and ii. a second nucleic acid molecule comprising, from 5' to 3', a 5' homology arm comprising a nucleotide sequence downstream of the 3' end of the template sequence, at least a second marker, and a 3' homology arm comprising a nucleotide sequence downstream of the 3' end of the target sequence; c. generating double-stranded breaks on either side of the target sequence and at the 5' and 3' ends of the template sequence; thereby inserting the template sequence and the first and second markers into the target chromosome; and d. Selecting one or more cells that express the first and second markers.

2. 2. The method of claim 1, wherein after insertion of the template sequence, the first marker is located at the 5' end of the template sequence and the second marker is located at the 3' end of the template sequence.

3. 2. The method of claim 1, wherein the length of the 5' and 3' homology arms of the first and second nucleic acid molecules is between about 20 bp and 2,000 bp, between about 50 bp and 1,500 bp, between about 100 bp and 1,400 bp, between about 150 bp and 1,300 bp, between about 200 bp and 1,200 bp, between about 300 bp and 1,100 bp, between about 400 bp and 1,000 bp, between about 500 bp and 900 bp, or between about 600 bp and 800 bp.

4. The length of the template sequence is 50KB and 250MB, 50KB and 100MB, 50KB and 50MB, 50KB and 20MB, 50KB and 10MB, 50KB and 5MB, 50KB and 3MB, 50KB and 2MB, 50KB and 1MB, 100KB and 200MB, 100KB and 100MB, 100KB and 50MB, 100KB and 20MB, 100KB and 10MB, 100KB and 5MB, 100KB and 3MB, 100KB and 2MB, 100KB and 1MB, 100KB and 500KB, 200KB and 100MB, 200KB and 50MB, 200KB and 20MB, 200KB and 10MB, 200KB and 5MB, 200KB and 3MB, 200KB and 2MB, 200KB and 1MB , 200KB and 500KB, 500KB and 100MB, 500KB and 50MB, 500KB and 20MB, 500KB and 10MB, 500KB and 5MB, 500KB and 3MB, 500KB and 2MB, 500KB and 1MB, 1MB and 100MB, 1MB and 50MB, 1MB and 20MB, 1MB and 10MB, 1MB and 5MB, 1MB and 3MB, 1MB and 2MB, 3MB and 100MB, 3MB and 50MB, 3MB and 20MB, 3MB and 10MB, 3MB and 5MB, 5MB and 100MB, 5MB and 50MB, 5MB and 20MB, 5MB and 10MB, 10MB and 100MB, 10MB and 50MB, or 10MB and 20MB.

5. 2. The method of claim 1, wherein generating the double-stranded break in (c) comprises inducing the double-stranded break using a CRISPR / Cas endonuclease and one or more guide nucleic acids (gNAs), one or more zinc finger nucleases, one or more transcription activator-like effector nucleases (TALENs), or one or more CRE recombinases.

6. 2. The method of claim 1, wherein the target chromosome comprises, from 5' to 3', the sequence of the 5' homology arm of the first nucleic acid molecule, the target sequence, and the sequence of the 3' homology arm of the second nucleic acid molecule.

7. 2. The method of claim 1, wherein the template chromosome comprises, from 5' to 3', the sequence of the 3' homology arm of the first nucleic acid molecule, the template sequence, and the sequence of the 5' homology arm of the second nucleic acid molecule.

8. 2. The method of claim 1, wherein the targeting sequence comprises one or more genes that are homologous to one or more genes of the template sequence.

9. The method of claim 1, wherein the template sequence comprises a naturally occurring sequence.

10. 10. The method of claim 9, wherein the template sequence contains one or more modifications to a naturally occurring sequence.

11. The method of claim 1, wherein the template sequence comprises an artificial sequence.

12. 12. The method of claim 11, wherein the artificial sequence comprises a sequence encoding one or more antibodies or antigen-binding fragments thereof.

13. The method of claim 1, wherein the insertion of the template sequence results in the deletion of the target sequence.

14. 14. The method of claim 13, wherein: a. the target chromosome comprises, from 5' to 3', the sequence of the 5' homology arm of a first nucleic acid molecule, a first sgRNA target sequence, a target sequence, a second sgRNA target sequence, and the sequence of the 3' homology arm of a second nucleic acid molecule; and b. The template chromosome comprises, from 5' to 3', the third sgRNA target sequence, the sequence of the 3' homology arm of the first nucleic acid molecule, the template sequence, the sequence of the 5' homology arm of the second nucleic acid molecule, and the fourth sgRNA target sequence.

15. 15. The method of claim 14, wherein generating the double-stranded break comprises contacting the cell with a CRISPR / Cas endonuclease and the first, second, third, and fourth sgRNAs.

16. 2. The method of claim 1, wherein the insertion of the template sequence does not involve deletion of little or any of the sequence of the target sequence.

17. 17. The method of claim 16, wherein a. the target chromosome comprises, from 5' to 3', the sequence of the 5' homology arm of a first nucleic acid molecule, the first sgRNA target sequence, and the sequence of the 3' homology arm of a second nucleic acid molecule; and b. The template chromosome comprises, from 5' to 3', the second sgRNA target sequence, the sequence of the 3' homology arm of the first nucleic acid molecule, the template sequence, the sequence of the 5' homology arm of the second nucleic acid molecule, and the third sgRNA target sequence.

18. 18. The method of Claim 17, wherein generating the double-stranded break comprises contacting the cell with a CRISPR / Cas endonuclease and the first, second, and third sgRNAs.

19. 2. The method of claim 1, wherein the first or second marker comprises a fluorescent protein operably linked to a promoter capable of expressing the fluorescent protein in the cell.

20. The method of claim 1, wherein the first marker further comprises a selectable marker.

21. The method of claim 1, wherein the second marker further comprises a selectable marker.

22. 22. The method of claim 20 or 21, wherein the selectable marker is selected from the group consisting of dihydrofolate reductase (DHFR), glutamine synthetase (GS), puromycin acetyltransferase, blastidine deaminase, histidinol dehydrogenase, hygromycin phosphotransferase (hph), bleomycin resistance gene, and aminoglycoside phosphotransferase (neomycin resistance).

23. 2. The method of claim 1, wherein the first marker comprises GFP and puromycin acetyltransferase operably linked to a promoter capable of expressing GFP in the cell, and the second marker comprises hygromycin phosphotransferase.

24. The method of claim 1, further comprising, after step (d), (e) deleting all or part of the first or second marker.

25. The method of claim 1, wherein the cell comprises a hybrid cell, an embryonic hybrid stem (EHS) cell, or a zygote.

26. 26. The method of claim 25, wherein the EHS cells are produced by fusing a human embryonic stem cell with an embryonic stem cell from a non-human species.

27. 27. The method of claim 26, wherein the EHS cells are generated by fusing ES cells from any two different species selected from the group consisting of mouse, rat, rabbit, guinea pig, hamster, sheep, goat, donkey, cow, horse, camel, chicken, and monkey.

28. 28. The method of claim 27, comprising generating a hybrid cell, wherein the method comprises: a. generating micronucleated human cells; and b. Fusing the micronucleated human cells with cells from a non-human species, thereby generating hybrid cells.

29. 2. The method of claim 1, wherein the target chromosome comprises mouse chromosome 12 and the template chromosome comprises human chromosome 14, or the target chromosome comprises mouse chromosome 6 and the template chromosome comprises human chromosome 2.

30. 30. The method of claim 29, wherein the target sequence comprises a mouse Igh variable region sequence, a mouse Igk variable region sequence, and / or a mouse Igl variable region sequence.

31. 31. The method of claim 29 or 30, wherein the template sequence comprises a human IGH variable region sequence, a human IGK variable region sequence, and / or a human IGL variable region sequence.

32. 2. The method of claim 1, further comprising recovering recombinant chromosomes from the cells selected in step (d).

33. A recombinant chromosome produced by the method of claim 1.

34. A cell comprising the recombinant chromosome of claim 33.

35. 1. A method for generating mouse embryonic stem cells, comprising: a. Fusing micronucleated cells containing recombinant chromosomes to mouse ES cells by the method of claim 1, wherein: i. the mouse ES cell comprises a chromosome homologous to the recombinant chromosome, the homologous chromosome comprising a first fluorescent protein operably linked to a promoter capable of expressing the fluorescent protein in the ES cell; and ii. at least a subset of the micronucleated cells comprise the recombinant chromosome, wherein the recombinant chromosome comprises a second fluorescent protein different from the first fluorescent protein, the second fluorescent protein being operably linked to a promoter capable of expressing the fluorescent protein in the ES cells; b. selecting ES cells that express both the first and second fluorescent proteins; c. culturing the ES cells selected in step (c) until the homologous chromosome has disappeared in at least a subset of the ES cells; and d. ES cells that express the second fluorescent protein but do not express the first fluorescent protein are selected.

36. A transgenic mouse produced from mouse ES cells produced by the method of claim 35.

37. A method of producing an antibody, comprising: a. exposing the transgenic mouse of claim 36 to an antigen, thereby causing the transgenic mouse to produce a plurality of antibodies comprising human V, D, and J segments from the human IGH variable region; and b. Isolating antibodies specific to the antigen.