Methods for breaking immunological tolerance using multiple guide rnas

By transfecting non-human animal cells with Cas9 and guide RNAs to target and modify genomic loci, the method addresses immune tolerance and efficient genome editing, producing antigen-binding proteins with enhanced titers and diversity.

JP2025159002APending Publication Date: 2025-10-17REGENERON PHARMACEUTICALS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025127388
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2016-07-29
Filing Date
2025-07-30
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Immunization of non-human animals with 'non-self' proteins to generate specific antigen-binding proteins, such as monoclonal antibodies, is challenging due to immune tolerance to self-antigens with high homology, and traditional genome editing methods struggle with efficient targeting and generating homozygous modifications, especially in eukaryotic cells.

Method used

A method involving transfection of non-human animal pluripotent cells with Cas9 protein and guide RNAs to target and modify genomic loci, followed by immunization to reduce tolerance to foreign antigens, using exogenous repair templates to introduce specific modifications, and obtaining antigen-binding proteins.

Benefits of technology

Generates non-human animals with reduced tolerance to foreign antigens, producing antigen-binding proteins with higher titers and diversity, overcoming immune tolerance and efficient genomic modification challenges.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025159002000001_ABST
    Figure 2025159002000001_ABST
Patent Text Reader

Abstract

To provide compositions and improved methods for producing antigen-binding proteins (e.g., antibodies) that bind an epitope on a foreign target antigen of interest (e.g., a human target antigen of interest) that shares the epitope with a self-antigen or is homologous to the self-antigen.SOLUTION: Methods according to the present invention comprise reducing tolerance of the foreign antigen in non-human animals such as rodents (e.g., mice or rats) (optionally comprising in their germline humanized immunoglobulin heavy and / or light chain loci) by employing two or more guide RNAs (gRNAs) to create paired double-strand breaks at different sites within a single target genomic locus.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Referencing sequence listings submitted as text files via EFS WEB The sequence listing in file 497023SEQLIST.txt, created on May 18, 2017, is 38.3 kilobytes and is incorporated herein by reference. [Background technology]

[0002] Immunization of non-human animals (e.g., rodents such as mice or rats) with "non-self" proteins is a commonly used method for obtaining specific antigen-binding proteins, such as monoclonal antibodies. However, this technique relies on sequence differences between the native protein of the non-human animal and the immunized protein, allowing the non-human animal's immune system to recognize the immunogen as non-self (i.e., foreign). The production of antibodies against antigens that share a high degree of homology with self-antigens can be challenging due to immune tolerance. Because functionally important regions of proteins tend to be conserved across species, immune tolerance to self-antigens often poses problems in generating antibodies against these important epitopes.

[0003] Although progress has been made in targeting various genomic loci, there remain many genomic loci that cannot be efficiently targeted or many genomic modifications that cannot be efficiently achieved using traditional targeting strategies. Although the CRISPR / Cas system provides a new tool for genome editing, problems remain. For example, problems can still arise in some situations when attempting to generate large targeted genomic deletions or other large targeted genetic modifications, particularly in eukaryotic cells and organisms.

[0004] In addition, it can be difficult to efficiently generate cells or animals homozygous for a targeted gene modification without subsequent breeding steps, and some loci may be more difficult to target to generate homozygous targeted modifications than others. For example, traditional targeting strategies may result in F0 generation mice heterozygous for large targeted genomic deletions, but subsequent breeding of these heterozygous mice is required to generate F1 generation mice homozygous for the deletion. These additional breeding steps are expensive and time-consuming. Summary of the Invention [Means for solving the problem]

[0005]

[0003] The present invention provides methods and compositions for generating non-human animals with reduced tolerance to a foreign antigen of interest, and methods and compositions for using such animals to produce antigen-binding proteins that bind to the foreign antigen of interest. In one aspect, the invention provides a method for (a) transfecting the genome of a non-human animal pluripotent cell that is not a one-cell stage embryo with (i) a Cas9 protein, (ii) a first guide RNA that hybridizes to a first guide RNA recognition sequence within a first target genomic locus (the first target genomic locus affects expression of a first self antigen that is homologous to the foreign antigen of interest or shares an epitope of interest with the foreign antigen of interest), and (iii) a second guide RNA that hybridizes to a second guide RNA recognition sequence within the first target genomic locus (the first target genomic locus affects expression of a first self antigen that is homologous to the foreign antigen of interest or shares an epitope of interest with the foreign antigen of interest).

[0009] Provided is a method for producing a non-human animal with suppressed tolerance to a foreign antigen of interest, comprising: (a) contacting a non-human animal with a Cas9 protein, a first guide RNA, and a second guide RNA (a gene encoding a Cas9 protein, ... Optionally, the Cas9 protein is introduced into the non-human animal pluripotent cell in the form of DNA encoding the Cas9 protein, the first guide RNA is introduced into the non-human animal pluripotent cell in the form of DNA encoding the first guide RNA, and the second guide RNA is introduced into the non-human animal pluripotent cell in the form of DNA encoding the second guide RNA.

[0006] In some such methods, the contacting step (a) further comprises contacting the genome with (iv) a third guide RNA that hybridizes to a third guide RNA recognition sequence in the first target genomic locus and / or (v) a fourth guide RNA that hybridizes to a fourth guide RNA recognition sequence in the first target genomic locus. In some such methods, the contacting step (a) further comprises contacting the genome with (iv) a third guide RNA that hybridizes to a third guide RNA recognition sequence in a second target genomic locus (the second target genomic locus affects expression of a first autoantigen or a second autoantigen that is homologous to or shares an epitope of interest with the foreign antigen of interest) and / or (v) a fourth guide RNA that hybridizes to a fourth guide RNA recognition sequence in the second target genomic locus.

[0007] In some such methods, the contacting step (a) further comprises contacting the genome with an exogenous repair template comprising a 5' homology arm that hybridizes to a 5' target sequence at the target genomic locus and a 3' homology arm that hybridizes to a 3' target sequence at the target genomic locus. Optionally, the exogenous repair template further comprises a nucleic acid insert flanked by the 5' homology arm and the 3' homology arm. In some such methods, the nucleic acid insert is homologous or orthologous to the first target genomic locus. In some such methods, the exogenous repair template is about 50 nucleotides to about 1 kb in length. In some such methods, the exogenous repair template is about 80 nucleotides to about 200 nucleotides in length. In some such methods, the exogenous repair template is a single-stranded oligodeoxynucleotide. In some such methods, the foreign repair template is a large targeting vector (LTVEC) at least 10 kb in length, and / or the foreign repair template is an LTVEC in which the combined length of the 5' homology arms and 3' homology arms of the LTVEC is at least 10 kb.

[0008] Some such methods further include (d) immunizing the genetically modified F0 generation non-human animal produced in step (c) with a foreign antigen of interest; (e) maintaining the genetically modified F0 generation non-human animal under conditions sufficient to mount an immune response to the foreign antigen of interest; and (f) obtaining from the genetically modified F0 generation non-human animal a first nucleic acid sequence encoding a human immunoglobulin heavy chain variable domain and / or a second nucleic acid sequence encoding a human immunoglobulin light chain variable domain.

[0009] In some such methods, the antigen binding proteins against the foreign antigen of interest obtained after immunization of the genetically modified F0 generation non-human animal with the foreign antigen of interest have a higher titer than the antigen binding proteins obtained after immunization of a control non-human animal that is wild-type at the first target genomic locus. In some such methods, a more diverse repertoire of antigen binding proteins against the foreign antigen of interest is obtained after immunization of the genetically modified F0 generation non-human animal with the foreign antigen of interest compared to the antigen binding proteins obtained after immunization of a control non-human animal that is wild-type at the first target genomic locus.

[0010] In some such methods, expression of the first autoantigen is eliminated.

[0011] In some such methods, the foreign antigen of interest is an ortholog of the first autoantigen. In some such methods, the foreign antigen of interest comprises, consists essentially of, or consists of all or part of a human protein.

[0012] In some such methods, the first target genomic locus is modified to include an insertion of one or more nucleotides, a deletion of one or more nucleotides, or a substitution of one or more nucleotides. In some such methods, the first target genomic locus is modified to include a deletion of one or more nucleotides. In some such methods, the contacting step (a) includes contacting the genome with an exogenous repair template including a 5' homology arm that hybridizes to a 5' target sequence at the target genomic locus and a 3' homology arm that hybridizes to a 3' target sequence at the target genomic locus, wherein when the genome is present in a one-cell stage embryo, the exogenous repair template is 5 kb in length or less, and the exogenous repair template includes a nucleic acid insert flanked by the 5' homology arm and the 3' homology arm, wherein the nucleic acid insert is homologous or orthologous to the deleted nucleic acid sequence, and the nucleic acid insert replaces the deleted nucleic acid sequence. In some such methods, the deletion is a precise deletion without random insertions and deletions (indels). In some such methods, the contacting step (a) comprises contacting the genome with an exogenous repair template comprising a 5' homology arm that hybridizes to a 5' target sequence at the target genomic locus and a 3' homology arm that hybridizes to a 3' target sequence at the target genomic locus, wherein when the genome is in a one-cell stage embryo, the exogenous repair template is 5 kb in length or less, and the deleted nucleic acid sequence consists of a nucleic acid sequence between the 5' target sequence and the 3' target sequence.

[0013] In some such methods, the first target genomic locus comprises, consists essentially of, or consists of all or a portion of a gene encoding a first autoantigen. In some such methods, the modification comprises, consists essentially of, or consists of a homozygous deletion of all or a portion of a gene encoding a first autoantigen. In some such methods, the modification comprises, consists essentially of, or consists of a homozygous disruption of the start codon of a gene encoding a first autoantigen.

[0014] In some such methods, the first guide RNA recognition sequence includes the start codon of a gene encoding the first autoantigen or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the start codon, and the second guide RNA recognition sequence includes the stop codon of a gene encoding the first autoantigen or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the stop codon. Optionally, the first guide RNA recognition sequence includes the start codon and the second guide RNA recognition sequence includes the stop codon. In some such methods, the first guide RNA recognition sequence comprises a first Cas9 cleavage site, the second guide RNA recognition sequence comprises a second Cas9 cleavage site, and the first target genomic locus is modified to comprise a deletion between the first and second Cas9 cleavage sites. Optionally, the deletion is a precise deletion, and the deleted nucleic acid sequence consists of the nucleic acid sequence between the first and second Cas9 cleavage sites.

[0015] In some such methods, the first and second guide RNA recognition sequences are different, and each of the first and second guide RNA recognition sequences includes the start codon of a gene encoding the first autoantigen or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the start codon. Optionally, each of the first and second guide RNA recognition sequences includes the start codon.

[0016] In some such methods, the first nucleic acid sequence and / or the second nucleic acid sequence is obtained from lymphocytes of a transgenic non-human animal or from hybridomas generated from lymphocytes.

[0017] In some such methods, the non-human animal comprises a humanized immunoglobulin locus. In some such methods, the non-human animal is a rodent. In some such methods, the rodent is a mouse. Optionally, the mouse strain comprises a BALB / c strain. Optionally, the mouse strain comprises a BALB / c strain, a C57BL / 6 strain, and a 129 strain. Optionally, the mouse strain is 50% BALB / c, 25% C57BL / 6, and 25% 129. Optionally, the MHC haplotype of the mouse is MHC b / d is.

[0018] In some such methods, the mouse comprises in its germline a human unrearranged variable region gene segment inserted into an endogenous mouse immunoglobulin locus. Optionally, the human unrearranged variable region gene segment is a heavy chain gene segment, and the mouse immunoglobulin locus is a heavy chain locus. Optionally, the human unrearranged variable region gene segment is a light chain segment, and the mouse immunoglobulin locus is a light chain locus. Optionally, the light chain gene segment is a human kappa or lambda light chain gene segment. In some such methods, the mouse comprises in its germline a human unrearranged variable region gene segment operably linked to a mouse constant region gene, and the mouse lacks a human constant region gene, and the mouse constant region gene is present at the endogenous mouse immunoglobulin locus. In some such methods, the mouse comprises (a) a hybrid heavy chain locus comprising an insertion of human immunoglobulin heavy chain V, D, and J gene segments, wherein the human heavy chain immunoglobulin V, D, and J gene segments are operably linked to a mouse immunoglobulin heavy chain gene, and the mouse immunoglobulin heavy chain gene is present at the endogenous mouse immunoglobulin locus, and (b) a hybrid light chain locus comprising an insertion of human immunoglobulin light chain V and J gene segments, wherein the human V and J gene segments are operably linked to a mouse immunoglobulin light chain constant region gene sequence, wherein (a) rearranges to form a hybrid heavy chain sequence comprising a human variable region operably linked to a mouse constant region, and (b) rearranges to form a hybrid light chain sequence comprising a human variable region operably linked to a mouse constant region, and the mouse is unable to form antibodies comprising the human variable region and the human constant region. In some such methods, the mouse comprises a modification of the immunoglobulin heavy chain locus, which modification suppresses or eliminates endogenous ADAM6 function, and the mouse comprises an ectopic nucleic acid sequence encoding a mouse ADAM6 protein, its ortholog, its homolog, or a fragment thereof, and the ADAM6 protein, its ortholog, its homolog, or a fragment thereof is functional in the male mouse.Optionally, the ectopic nucleic acid sequence encoding the mouse ADAM6 protein, its ortholog, its homolog, or fragment thereof is present at the human heavy chain variable region locus. Optionally, the ectopic nucleic acid sequence encoding the mouse ADAM6 protein, its ortholog, its homolog, or fragment thereof is present at a location other than the human heavy chain variable region locus.

[0019] In some such methods, the mouse comprises in its germline a humanized immunoglobulin light chain variable locus comprising no more than one or no more than two rearranged human light chain V / J sequences operably linked to a light chain constant region. Optionally, the light chain constant region gene is a mouse gene. In some such methods, the mouse further comprises a humanized immunoglobulin heavy chain variable locus comprising at least one unrearranged human V segment, at least one unrearranged human D segment, and at least one unrearranged human J segment operably linked to a heavy chain constant region gene. Optionally, the heavy chain constant region gene is a mouse gene. In some such methods, the mouse comprises a humanized heavy chain immunoglobulin variable locus and a humanized light chain immunoglobulin variable locus, and the mouse expresses a single light chain. In some such methods, the mouse comprises (a) a human V of an immunoglobulin light chain, and (b) a human V of an immunoglobulin light chain. L A single rearranged human immunoglobulin light chain variable region (V L / J L )(Single reconstituted human V L / J L the region is selected from a human Vκ1-39 / J gene segment or a human Vκ3-20 / J gene segment), and (b) an endogenous heavy chain variable (V H ) gene segment with one or more human VH gene segments (human V H The gene segment is an endogenous heavy chain constant (C H ) region gene and functionally linked to the human V HIn some such methods, the mice express a population of antibodies, the germline of the mice contains only a single immunoglobulin kappa light chain variable region gene that is a rearranged human germline kappa light chain variable region gene, and the mice are either heterozygous for the single immunoglobulin kappa light chain variable region gene, in that they contain only one copy, or homozygous for the single immunoglobulin kappa light chain variable region gene, in that they contain two copies, and the mice express a population of antibodies, the population of antibodies, the germline of the mice contains only a single immunoglobulin kappa light chain variable region gene that is a rearranged human germline kappa light chain variable region gene, and the mice express a population of antibodies, the population of antibodies, the mouse germline contains only a single immunoglobulin kappa light chain variable region gene that is a rearranged human germline kappa light chain variable region gene, and the mice are either heterozygous for the single immunoglobulin kappa light chain variable region gene, in that they contain only one copy, or homozygous for the single immunoglobulin kappa light chain variable region gene, in that they contain two copies, and the mice express a population of antibodies, the population of antibodies, the mouse germline contains only a single immunoglobulin kappa light chain variable region gene that is a rearranged human germline kappa light chain variable region gene ... mouse germline contains only a single immunoglobulin kappa light chain variable region gene that is a rearranged human germline kappa light chain variable region gene, and the mice express a population of antibodies, the mouse germline contains only a single immunoglobulin kappa light chain variable region gene, and the mice express a population of antibodies, the mouse germline contains only a single immunoglobulin kappa light chain variable region gene, and the mice express a population of antibodies, the mouse germline contains only a single immunoglobulin kappa light chain variable region gene that is a (ii) the population includes antibodies comprising immunoglobulin kappa light chains whose light chain variable domains are encoded by rearranged human germline kappa light chain variable region genes and antibodies comprising immunoglobulin kappa light chains whose light chain variable domains are encoded by somatic variants thereof; and (iii) the mice are characterized by active affinity maturation to produce a diverse population of somatically mutated high-affinity heavy chains that successfully pair with immunoglobulin kappa light chains to form antibodies in the population. Optionally, the mouse is heterozygous or homozygous within its germline for: (a) an insertion at the endogenous mouse κ immunoglobulin light chain variable region locus of a rearranged Vκ / Jκ sequence comprising: (i) a single human germline Vκ sequence (the single human germline Vκ sequence is present in SEQ ID NO: 148 or SEQ ID NO: 149), and (ii) a single human germline Jκ sequence, wherein the rearranged Vκ / Jκ sequence is operably linked to an endogenous mouse κ constant region; and (b) an insertion at the endogenous mouse immunoglobulin heavy chain variable region locus of a plurality of human immunoglobulin heavy chain variable region gene segments, wherein the human immunoglobulin heavy chain variable region gene segments are operably linked to an endogenous mouse immunoglobulin heavy chain constant region, and wherein the human immunoglobulin heavy chain variable region gene segments are capable of rearranging and forming a rearranged human / mouse chimeric immunoglobulin heavy chain gene.In some such methods, the mouse comprises a modification of the immunoglobulin heavy chain locus, wherein the modification suppresses or eliminates endogenous ADAM6 function, and the mouse comprises an ectopic nucleic acid sequence encoding a mouse ADAM6 protein, its ortholog, its homolog, or a fragment thereof, wherein the ADAM6 protein, its ortholog, its homolog, or a fragment thereof is functional in the male mouse. Optionally, the ectopic nucleic acid sequence encoding the mouse ADAM6 protein, its ortholog, its homolog, or a fragment thereof is present at a human heavy chain variable region locus. Optionally, the ectopic nucleic acid sequence encoding the mouse ADAM6 protein, its ortholog, its homolog, or a fragment thereof is present at a location other than the human heavy chain variable region locus.

[0020] In some such methods, the mouse has a genome comprising a modification of an immunoglobulin heavy chain locus, wherein the modification reduces or eliminates endogenous ADAM6 function, and the mouse further comprises a nucleic acid sequence encoding a non-human animal ADAM6 protein, or an ortholog or homolog thereof, or a functional fragment of the corresponding ADAM6 protein. Optionally, the genome of the mouse comprises (a) an ectopic placement of an ADAM6 gene, and (b) one or more human V(s) into the endogenous non-human animal heavy chain locus. H gene segment, one or more human D H gene segment, and one or more human J H a human immunoglobulin heavy chain variable region locus comprising an insertion of a gene segment, H , D H and J H Because the gene segments are operably linked to the heavy chain constant region gene, the mice are (i) fertile, and (ii) when immunized with an antigen, produce one or more human V gene segments operably linked to the heavy chain constant domains encoded by the heavy chain constant region gene. H gene segment, one or more human D H gene segment, and one or more human J HThe gene segment is characterized by producing an antibody comprising a heavy chain variable domain encoded by the gene segment, the antibody exhibiting specific binding to an antigen.

[0021] In some such methods, the non-human animal is a mouse at least partially derived from the BALB / c strain, the mouse comprises a humanized immunoglobulin locus, the foreign antigen of interest is all or a portion of a human protein orthologous to a first autoantigen, the first target genomic locus comprises all or a portion of a gene encoding the first autoantigen, the first guide RNA recognition site comprises a start codon of the gene encoding the first autoantigen, the second guide RNA recognition site comprises a stop codon of the gene encoding the first autoantigen, and the modification comprises a homozygous deletion of all or a portion of the gene encoding the first autoantigen, thereby eliminating expression of the first autoantigen. Optionally, the mouse comprises: (a) an ectopic nucleic acid sequence encoding a mouse ADAM6 protein, its ortholog, its homolog, or a fragment thereof, wherein the ADAM6 protein, its ortholog, its homolog, or its fragment is functional in the male mouse; (b) a hybrid heavy chain locus comprising an insertion of human immunoglobulin heavy chain V, D, and J gene segments, wherein the human heavy chain immunoglobulin V, D, and J gene segments are operably linked to a mouse immunoglobulin heavy chain gene, and the mouse immunoglobulin heavy chain gene is operably linked to an endogenous mouse immunoglobulin gene; (b) a hybrid heavy chain locus comprising a human variable region operably linked to a mouse constant region; and (c) a hybrid light chain locus comprising an insertion of human immunoglobulin light chain V and J gene segments, wherein the human V and J gene segments are operably linked to a mouse immunoglobulin light chain constant region gene sequence, wherein (b) rearranges to form a hybrid heavy chain sequence comprising a human variable region operably linked to a mouse constant region; and (c) rearranges to form a hybrid light chain sequence comprising a human variable region operably linked to a mouse constant region; and the mouse is unable to form antibodies comprising a human variable region and a human constant region.Optionally, the mouse contains, within its germline, an endogenous mouse κ immunoglobulin light chain capable of rearranged Vκ / Jκ sequences comprising: (a) an ectopic nucleic acid sequence encoding a mouse ADAM6 protein, its ortholog, its homolog, or a fragment thereof, wherein the ADAM6 protein, its ortholog, its homolog, or a fragment thereof is functional in the male mouse; and (b) an endogenous mouse κ immunoglobulin light chain capable of rearranged Vκ / Jκ sequences comprising: (i) a single human germline Vκ sequence, wherein the single human germline Vκ sequence is present in SEQ ID NO: 148 or SEQ ID NO: 149, and (ii) a single human germline Jκ sequence. and (c) heterozygous or homozygous for an insertion at the variable region locus (wherein the rearranged Vκ / Jκ sequence is operably linked to the endogenous mouse κ constant region), and (d) for an insertion of multiple human immunoglobulin heavy chain variable region gene segments at the endogenous mouse immunoglobulin heavy chain variable region locus (wherein the human immunoglobulin heavy chain variable region gene segments are operably linked to the endogenous mouse immunoglobulin heavy chain constant region, and the human immunoglobulin heavy chain variable region gene segments are capable of rearranging and forming a rearranged human / mouse chimeric immunoglobulin heavy chain gene).

[0022] In some such methods, the non-human animal is a mouse at least partially derived from the BALB / c strain, the mouse comprises a humanized immunoglobulin locus, the foreign antigen of interest is all or a portion of a human protein orthologous to a first autoantigen, the first target genomic locus comprises all or a portion of a gene encoding the first autoantigen, the first guide RNA recognition site comprises a start codon of the gene encoding the first autoantigen, the second guide RNA recognition site comprises a stop codon of the gene encoding the first autoantigen, and the modification comprises a homozygous disruption of the start codon of the gene encoding the first autoantigen, thereby eliminating expression of the first autoantigen. Optionally, the mouse comprises: (a) an ectopic nucleic acid sequence encoding a mouse ADAM6 protein, its ortholog, its homolog, or a fragment thereof, wherein the ADAM6 protein, its ortholog, its homolog, or its fragment is functional in the male mouse; (b) a hybrid heavy chain locus comprising an insertion of human immunoglobulin heavy chain V, D, and J gene segments, wherein the human heavy chain immunoglobulin V, D, and J gene segments are operably linked to a mouse immunoglobulin heavy chain gene, and the mouse immunoglobulin heavy chain gene is operably linked to an endogenous mouse immunoglobulin gene; (b) a hybrid heavy chain locus comprising a human variable region operably linked to a mouse constant region; and (c) a hybrid light chain locus comprising an insertion of human immunoglobulin light chain V and J gene segments, wherein the human V and J gene segments are operably linked to a mouse immunoglobulin light chain constant region gene sequence, wherein (b) rearranges to form a hybrid heavy chain sequence comprising a human variable region operably linked to a mouse constant region; and (c) rearranges to form a hybrid light chain sequence comprising a human variable region operably linked to a mouse constant region; and the mouse is unable to form antibodies comprising a human variable region and a human constant region.Optionally, the mouse contains, within its germline, an endogenous mouse κ immunoglobulin light chain capable of rearranged Vκ / Jκ sequences comprising: (a) an ectopic nucleic acid sequence encoding a mouse ADAM6 protein, its ortholog, its homolog, or a fragment thereof, wherein the ADAM6 protein, its ortholog, its homolog, or a fragment thereof is functional in the male mouse; and (b) an endogenous mouse κ immunoglobulin light chain capable of rearranged Vκ / Jκ sequences comprising: (i) a single human germline Vκ sequence, wherein the single human germline Vκ sequence is present in SEQ ID NO: 148 or SEQ ID NO: 149, and (ii) a single human germline Jκ sequence. and (c) heterozygous or homozygous for an insertion at the variable region locus (wherein the rearranged Vκ / Jκ sequence is operably linked to the endogenous mouse κ constant region), and (d) for an insertion of multiple human immunoglobulin heavy chain variable region gene segments at the endogenous mouse immunoglobulin heavy chain variable region locus (wherein the human immunoglobulin heavy chain variable region gene segments are operably linked to the endogenous mouse immunoglobulin heavy chain constant region, and the human immunoglobulin heavy chain variable region gene segments are capable of rearranging and forming a rearranged human / mouse chimeric immunoglobulin heavy chain gene).

[0023] In some methods, the non-human animal pluripotent cell is a hybrid cell, and the method further comprises (a') comparing the sequences of the corresponding first and second chromosomes of a homologous chromosome pair within the first target genomic locus to select a target region within the first target genomic locus, and then performing the contacting step (a) based on the target region having a higher percentage of sequence identity between the corresponding first and second chromosomes of the homologous chromosome pair compared to all or a portion of the remainder of the first target genomic locus. Optionally, the target region has a higher percentage of sequence identity between the corresponding first and second chromosomes of the homologous chromosome pair compared to the remainder of the first target genomic locus. Optionally, the target region has at least 99.9% sequence identity between the corresponding first and second chromosomes, and the remainder of the first target genomic locus has 99.8% or less sequence identity between the corresponding first and second chromosomes. Optionally, the target region is identical in corresponding first and second chromosomes of a homologous chromosome pair. Optionally, the target region is within the longest possible stretch of contiguous allelic sequence identity within the first target genomic locus.

[0024] In some such methods, the target region comprises a first guide RNA recognition sequence and at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb of flanking sequence on the 5', 3', or both sides of the first guide RNA recognition sequence, and and at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb of flanking sequence on the 5', 3', or both sides of the second guide RNA recognition sequence.Optionally, step (a') comprises comparing two or more segments of the first target genomic locus, each segment comprising a different guide RNA recognition sequence that is not present elsewhere in the genome, and at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, or 1,000 bp 5', 3', or both of the different guide RNA recognition sequences. selecting as target regions two segments that have the highest percentage of sequence identity relative to other segments (e.g., comprising, consisting essentially of, or consisting of 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb of contiguous sequence). Optionally, the one or more segments comprise, consist essentially of, or consist of segments that correspond to respective different guide RNA recognition sequences in the first target genomic locus but are not present elsewhere in the genome.

[0025] In some such methods, the target region comprises, consists essentially of, or consists of the region between the first guide RNA recognition sequence and the second guide RNA recognition sequence. Optionally, step (a') comprises comparing two or more segments of the first target genomic locus, each segment comprising, consisting essentially of, or consisting of a region between a different guide RNA recognition sequence pair, where the guide RNA recognition sequence is not present elsewhere in the genome, and selecting as the target region the segment having the highest percentage of sequence identity relative to the other segments. Optionally, the one or more segments comprise, consist essentially of, or consist of segments corresponding to each different guide RNA recognition sequence pair in the first target genomic locus, where the guide RNA recognition sequence is not present elsewhere in the genome.

[0026] In some such methods, the target region is at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, or more of the region between the first guide RNA recognition sequence and the second guide RNA recognition sequence, and at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, or more of the region 5', 3', or both of the genomic region between the first guide RNA recognition sequence and the second guide RNA recognition sequence. bp, 800 bp, 900 bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb or 150 kb of contiguous sequence. Optionally, step (a') comprises comparing two or more segments of the first target genomic locus (each segment being at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 ... selecting as the target region the segment having the highest percentage of sequence identity relative to other segments (e.g., 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150 kb of contiguous sequence, wherein the guide RNA recognition sequence is not present elsewhere in the genome). Optionally, the one or more segments comprise, consist essentially of, or consist of segments corresponding to each different guide RNA recognition sequence pair within the first target genomic locus, wherein the guide RNA recognition sequence is not present elsewhere in the genome.

[0027] In some such methods, the target region comprises, consists essentially of, or consists of at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb of contiguous sequence on the 5', 3', or both sides of the genomic region between the first guide RNA recognition sequence and the second guide RNA recognition sequence. Optionally, step (a') compares two or more non-contiguous segments of the first target genomic locus (each non-contiguous segment being at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 ... selecting as the target region the non-contiguous segment that has the highest percentage of sequence identity compared to other non-contiguous segments (e.g., 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb of contiguous sequence, wherein the guide RNA recognition sequence is not present elsewhere in the genome). Optionally, the one or more non-contiguous segments comprise, consist essentially of, or consist of non-contiguous segments corresponding to each different guide RNA recognition sequence pair within the first target genomic locus, wherein the guide RNA recognition sequence is not present elsewhere in the genome.

[0028] In some such methods, the target region comprises, consists essentially of, or consists of at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb of flanking sequence on either side of the genomic region between the first guide RNA recognition sequence and the second guide RNA recognition sequence. Optionally, step (a') compares two or more non-contiguous segments of the first target genomic locus (each non-contiguous segment being at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 ... selecting as the target region the non-contiguous segment that has the highest percentage of sequence identity compared to other non-contiguous segments (e.g., 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb of contiguous sequence, wherein the guide RNA recognition sequence is not present elsewhere in the genome). Optionally, the one or more non-contiguous segments comprise, consist essentially of, or consist of non-contiguous segments corresponding to each different guide RNA recognition sequence pair within the first target genomic locus, wherein the guide RNA recognition sequence is not present elsewhere in the genome.

[0029] In some such methods, the target region of step (a') comprises, consists essentially of, or consists of a region flanking the 5' target sequence and the 3' target sequence. In some such methods, the target region of step (a') comprises, consists essentially of, or consists of a region flanking and containing the 5' target sequence and the 3' target sequence. In some such methods, the target region of step (a') comprises, consists essentially of, or consists of the 5' target sequence and / or the 3' target sequence. Optionally, the target genomic locus of step (a') comprises, consists essentially of, or consists of the 5' target sequence and the 3' target sequence. In some such methods, the target region of step (a') is at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, or a region between the 5' and 3' target sequences and on the 5', 3', or both sides of the region between the 5' and 3' target sequences. bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb or 150 kb of contiguous sequence. In some such methods, the target region of step (a') is at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102 000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb or 150 kb of contiguous sequence.In some such methods, the target region of step (a') comprises, consists essentially of, or consists of at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb of flanking sequence on the 5', 3', or both sides of the region between the 5' target sequence and the 3' target sequence. In some such methods, the target region of step (a') comprises, consists essentially of, or consists of at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb of flanking sequence on either side of the region between the 5' target sequence and the 3' target sequence.

[0030] In another aspect, the invention provides a method for producing a non-human animal one-cell stage embryo by (a) contacting the genome of the non-human animal one-cell stage embryo with (i) a Cas9 protein, (ii) a first guide RNA that hybridizes to a first guide RNA recognition sequence within a first target genomic locus (the first target genomic locus affects expression of a first self antigen that is homologous to or shares an epitope of interest with a foreign antigen of interest), and (iii) a second guide RNA that hybridizes to a second guide RNA recognition sequence within the first target genomic locus (the first target genomic locus affects expression of a first self antigen that is homologous to or shares an epitope of interest with the foreign antigen of interest). (b) modifying the target genomic locus in a first chromosome and a second chromosome pair to generate a biallelic modification, thereby reducing the expression of the first autoantigen in the modified non-human animal one-cell stage embryo; and (b) implanting the modified non-human animal one-cell stage embryo into a surrogate mother to generate a transgenic F0 generation non-human animal in which the first target genomic locus in the first chromosome and the second chromosome pair is modified such that expression of the first autoantigen is suppressed. Optionally, the contacting comprises introducing the Cas9 protein, the first guide RNA, and the second guide RNA into the non-human animal one-cell stage embryo by nucleofection. Optionally, the Cas9 protein is introduced into the non-human animal one-cell stage embryo in the form of DNA encoding the Cas9 protein, the first guide RNA is introduced into the non-human animal one-cell stage embryo in the form of DNA encoding the first guide RNA, and the second guide RNA is introduced into the non-human animal one-cell stage embryo in the form of DNA encoding the second guide RNA.

[0031] In some such methods, the contacting step (a) further comprises contacting the genome with (iv) a third guide RNA that hybridizes to a third guide RNA recognition sequence in the first target genomic locus and / or (v) a fourth guide RNA that hybridizes to a fourth guide RNA recognition sequence in the first target genomic locus. In some such methods, the contacting step (a) further comprises contacting the genome with (iv) a third guide RNA that hybridizes to a third guide RNA recognition sequence in a second target genomic locus (the second target genomic locus affects expression of a first autoantigen or a second autoantigen that is homologous to or shares an epitope of interest with the foreign antigen of interest) and / or (v) a fourth guide RNA that hybridizes to a fourth guide RNA recognition sequence in the second target genomic locus.

[0032] In some such methods, the contacting step (a) further comprises contacting the genome with an exogenous repair template comprising a 5' homology arm that hybridizes to a 5' target sequence at the target genomic locus and a 3' homology arm that hybridizes to a 3' target sequence at the target genomic locus, wherein the exogenous repair template is about 50 nucleotides to about 5 kb in length. Optionally, the exogenous repair template further comprises a nucleic acid insert flanked by the 5' homology arm and the 3' homology arm. In some such methods, the nucleic acid insert is homologous or orthologous to the first target genomic locus. In some such methods, the exogenous repair template is about 50 nucleotides to about 1 kb in length. In some such methods, the exogenous repair template is about 80 nucleotides to about 200 nucleotides in length. In some such methods, the exogenous repair template is a single-stranded oligodeoxynucleotide.

[0033] Some such methods further include (c) immunizing the genetically modified F0 generation non-human animal produced in step (b) with a foreign antigen of interest; (d) maintaining the genetically modified F0 generation non-human animal under conditions sufficient to mount an immune response to the foreign antigen of interest; and (e) obtaining from the genetically modified F0 generation non-human animal a first nucleic acid sequence encoding a human immunoglobulin heavy chain variable domain and / or a second nucleic acid sequence encoding a human immunoglobulin light chain variable domain.

[0034] In some such methods, the antigen binding proteins against the foreign antigen of interest obtained after immunization of the genetically modified F0 generation non-human animal with the foreign antigen of interest have a higher titer than the antigen binding proteins obtained after immunization of a control non-human animal that is wild-type at the first target genomic locus. In some such methods, a more diverse repertoire of antigen binding proteins against the foreign antigen of interest is obtained after immunization of the genetically modified F0 generation non-human animal with the foreign antigen of interest compared to the antigen binding proteins obtained after immunization of a control non-human animal that is wild-type at the first target genomic locus.

[0035] In some such methods, expression of the first autoantigen is eliminated.

[0036] In some such methods, the foreign antigen of interest is an ortholog of the first autoantigen. In some such methods, the foreign antigen of interest comprises, consists essentially of, or consists of all or part of a human protein.

[0037] In some such methods, the first target genomic locus is modified to include an insertion of one or more nucleotides, a deletion of one or more nucleotides, or a substitution of one or more nucleotides. In some such methods, the first target genomic locus is modified to include a deletion of one or more nucleotides. In some such methods, the contacting step (a) includes contacting the genome with an exogenous repair template including a 5' homology arm that hybridizes to a 5' target sequence at the target genomic locus and a 3' homology arm that hybridizes to a 3' target sequence at the target genomic locus, wherein when the genome is present in a one-cell stage embryo, the exogenous repair template is 5 kb in length or less, and the exogenous repair template includes a nucleic acid insert flanked by the 5' homology arm and the 3' homology arm, wherein the nucleic acid insert is homologous or orthologous to the deleted nucleic acid sequence, and the nucleic acid insert replaces the deleted nucleic acid sequence. In some such methods, the deletion is a precise deletion without random insertions and deletions (indels). In some such methods, the contacting step (a) comprises contacting the genome with an exogenous repair template comprising a 5' homology arm that hybridizes to a 5' target sequence at the target genomic locus and a 3' homology arm that hybridizes to a 3' target sequence at the target genomic locus, wherein when the genome is in a one-cell stage embryo, the exogenous repair template is 5 kb in length or less, and the deleted nucleic acid sequence consists of a nucleic acid sequence between the 5' target sequence and the 3' target sequence.

[0038] In some such methods, the first target genomic locus comprises, consists essentially of, or consists of all or a portion of a gene encoding a first autoantigen. In some such methods, the modification comprises, consists essentially of, or consists of a homozygous deletion of all or a portion of a gene encoding a first autoantigen. In some such methods, the modification comprises, consists essentially of, or consists of a homozygous disruption of the start codon of a gene encoding a first autoantigen.

[0039] In some such methods, the first guide RNA recognition sequence includes the start codon of a gene encoding the first autoantigen or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the start codon, and the second guide RNA recognition sequence includes the stop codon of a gene encoding the first autoantigen or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the stop codon. Optionally, the first guide RNA recognition sequence includes the start codon and the second guide RNA recognition sequence includes the stop codon. In some such methods, the first guide RNA recognition sequence comprises a first Cas9 cleavage site, the second guide RNA recognition sequence comprises a second Cas9 cleavage site, and the first target genomic locus is modified to comprise a deletion between the first and second Cas9 cleavage sites. Optionally, the deletion is a precise deletion, and the deleted nucleic acid sequence consists of the nucleic acid sequence between the first and second Cas9 cleavage sites.

[0040] In some such methods, the first and second guide RNA recognition sequences are different, and each of the first and second guide RNA recognition sequences includes the start codon of a gene encoding the first autoantigen or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the start codon. Optionally, each of the first and second guide RNA recognition sequences includes the start codon.

[0041] In some such methods, the first nucleic acid sequence and / or the second nucleic acid sequence is obtained from lymphocytes of a transgenic non-human animal or from hybridomas generated from lymphocytes.

[0042] In some such methods, the non-human animal comprises a humanized immunoglobulin locus. In some such methods, the non-human animal is a rodent. In some such methods, the rodent is a mouse. Optionally, the mouse strain comprises a BALB / c strain. Optionally, the mouse strain comprises a BALB / c strain, a C57BL / 6 strain, and a 129 strain. Optionally, the mouse strain is 50% BALB / c, 25% C57BL / 6, and 25% 129. Optionally, the MHC haplotype of the mouse is MHC b / d is.

[0043] In some such methods, the mouse comprises in its germline a human unrearranged variable region gene segment inserted into an endogenous mouse immunoglobulin locus. Optionally, the human unrearranged variable region gene segment is a heavy chain gene segment, and the mouse immunoglobulin locus is a heavy chain locus. Optionally, the human unrearranged variable region gene segment is a light chain segment, and the mouse immunoglobulin locus is a light chain locus. Optionally, the light chain gene segment is a human kappa or lambda light chain gene segment. In some such methods, the mouse comprises in its germline a human unrearranged variable region gene segment operably linked to a mouse constant region gene, and the mouse lacks a human constant region gene, and the mouse constant region gene is present at the endogenous mouse immunoglobulin locus. In some such methods, the mouse comprises (a) a hybrid heavy chain locus comprising an insertion of human immunoglobulin heavy chain V, D, and J gene segments, wherein the human heavy chain immunoglobulin V, D, and J gene segments are operably linked to a mouse immunoglobulin heavy chain gene, and the mouse immunoglobulin heavy chain gene is present at the endogenous mouse immunoglobulin locus, and (b) a hybrid light chain locus comprising an insertion of human immunoglobulin light chain V and J gene segments, wherein the human V and J gene segments are operably linked to a mouse immunoglobulin light chain constant region gene sequence, wherein (a) rearranges to form a hybrid heavy chain sequence comprising a human variable region operably linked to a mouse constant region, and (b) rearranges to form a hybrid light chain sequence comprising a human variable region operably linked to a mouse constant region, and the mouse is unable to form antibodies comprising the human variable region and the human constant region. In some such methods, the mouse comprises a modification of the immunoglobulin heavy chain locus, which modification suppresses or eliminates endogenous ADAM6 function, and the mouse comprises an ectopic nucleic acid sequence encoding a mouse ADAM6 protein, its ortholog, its homolog, or a fragment thereof, and the ADAM6 protein, its ortholog, its homolog, or a fragment thereof is functional in the male mouse.Optionally, the ectopic nucleic acid sequence encoding the mouse ADAM6 protein, its ortholog, its homolog, or fragment thereof is present at the human heavy chain variable region locus. Optionally, the ectopic nucleic acid sequence encoding the mouse ADAM6 protein, its ortholog, its homolog, or fragment thereof is present at a location other than the human heavy chain variable region locus.

[0044] In some such methods, the mouse comprises in its germline a humanized immunoglobulin light chain variable locus comprising no more than one or no more than two rearranged human light chain V / J sequences operably linked to a light chain constant region. Optionally, the light chain constant region gene is a mouse gene. In some such methods, the mouse further comprises a humanized immunoglobulin heavy chain variable locus comprising at least one unrearranged human V segment, at least one unrearranged human D segment, and at least one unrearranged human J segment operably linked to a heavy chain constant region gene. Optionally, the heavy chain constant region gene is a mouse gene. In some such methods, the mouse comprises a humanized heavy chain immunoglobulin variable locus and a humanized light chain immunoglobulin variable locus, and the mouse expresses a single light chain. In some such methods, the mouse comprises (a) a human V of an immunoglobulin light chain, and (b) a human V of an immunoglobulin light chain. L A single rearranged human immunoglobulin light chain variable region (V L / J L )(Single reconstituted human V L / J L the region is selected from a human Vκ1-39 / J gene segment or a human Vκ3-20 / J gene segment), and (b) an endogenous heavy chain variable (V H ) gene segment with one or more human VH gene segments (human V H The gene segment is an endogenous heavy chain constant (C H ) region gene and functionally linked to the human V HIn some such methods, the mice express a population of antibodies, the germline of the mice contains only a single immunoglobulin kappa light chain variable region gene that is a rearranged human germline kappa light chain variable region gene, and the mice are either heterozygous for the single immunoglobulin kappa light chain variable region gene, in that they contain only one copy, or homozygous for the single immunoglobulin kappa light chain variable region gene, in that they contain two copies, and the mice express a population of antibodies, the population of antibodies, the germline of the mice contains only a single immunoglobulin kappa light chain variable region gene that is a rearranged human germline kappa light chain variable region gene, and the mice express a population of antibodies, the population of antibodies, the mouse germline contains only a single immunoglobulin kappa light chain variable region gene that is a rearranged human germline kappa light chain variable region gene, and the mice are either heterozygous for the single immunoglobulin kappa light chain variable region gene, in that they contain only one copy, or homozygous for the single immunoglobulin kappa light chain variable region gene, in that they contain two copies, and the mice express a population of antibodies, the population of antibodies, the mouse germline contains only a single immunoglobulin kappa light chain variable region gene that is a rearranged human germline kappa light chain variable region gene ... mouse germline contains only a single immunoglobulin kappa light chain variable region gene that is a rearranged human germline kappa light chain variable region gene, and the mice express a population of antibodies, the mouse germline contains only a single immunoglobulin kappa light chain variable region gene, and the mice express a population of antibodies, the mouse germline contains only a single immunoglobulin kappa light chain variable region gene, and the mice express a population of antibodies, the mouse germline contains only a single immunoglobulin kappa light chain variable region gene that is a (ii) the population includes antibodies comprising immunoglobulin kappa light chains whose light chain variable domains are encoded by rearranged human germline kappa light chain variable region genes and antibodies comprising immunoglobulin kappa light chains whose light chain variable domains are encoded by somatic variants thereof; and (iii) the mice are characterized by active affinity maturation to produce a diverse population of somatically mutated high-affinity heavy chains that successfully pair with immunoglobulin kappa light chains to form antibodies in the population. Optionally, the mouse is heterozygous or homozygous within its germline for: (a) an insertion at the endogenous mouse κ immunoglobulin light chain variable region locus of a rearranged Vκ / Jκ sequence comprising: (i) a single human germline Vκ sequence (the single human germline Vκ sequence is present in SEQ ID NO: 148 or SEQ ID NO: 149), and (ii) a single human germline Jκ sequence, wherein the rearranged Vκ / Jκ sequence is operably linked to an endogenous mouse κ constant region; and (b) an insertion at the endogenous mouse immunoglobulin heavy chain variable region locus of a plurality of human immunoglobulin heavy chain variable region gene segments, wherein the human immunoglobulin heavy chain variable region gene segments are operably linked to an endogenous mouse immunoglobulin heavy chain constant region, and wherein the human immunoglobulin heavy chain variable region gene segments are capable of rearranging and forming a rearranged human / mouse chimeric immunoglobulin heavy chain gene.In some such methods, the mouse comprises a modification of the immunoglobulin heavy chain locus, wherein the modification suppresses or eliminates endogenous ADAM6 function, and the mouse comprises an ectopic nucleic acid sequence encoding a mouse ADAM6 protein, its ortholog, its homolog, or a fragment thereof, wherein the ADAM6 protein, its ortholog, its homolog, or a fragment thereof is functional in the male mouse. Optionally, the ectopic nucleic acid sequence encoding the mouse ADAM6 protein, its ortholog, its homolog, or a fragment thereof is present at a human heavy chain variable region locus. Optionally, the ectopic nucleic acid sequence encoding the mouse ADAM6 protein, its ortholog, its homolog, or a fragment thereof is present at a location other than the human heavy chain variable region locus.

[0045] In some such methods, the mouse has a genome comprising a modification of an immunoglobulin heavy chain locus, wherein the modification reduces or eliminates endogenous ADAM6 function, and the mouse further comprises a nucleic acid sequence encoding a non-human animal ADAM6 protein, or an ortholog or homolog thereof, or a functional fragment of the corresponding ADAM6 protein. Optionally, the genome of the mouse comprises (a) an ectopic placement of an ADAM6 gene, and (b) one or more human V(s) into the endogenous non-human animal heavy chain locus. H gene segment, one or more human D H gene segment, and one or more human J H a human immunoglobulin heavy chain variable region locus comprising an insertion of a gene segment, H , D H and J H Because the gene segments are operably linked to the heavy chain constant region gene, the mice are (i) fertile, and (ii) when immunized with an antigen, produce one or more human V gene segments operably linked to the heavy chain constant domains encoded by the heavy chain constant region gene. H gene segment, one or more human D H gene segment, and one or more human J HThe gene segment is characterized by producing an antibody comprising a heavy chain variable domain encoded by the gene segment, the antibody exhibiting specific binding to an antigen.

[0046] In some such methods, the non-human animal is a mouse at least partially derived from the BALB / c strain, the mouse comprises a humanized immunoglobulin locus, the foreign antigen of interest is all or a portion of a human protein orthologous to a first autoantigen, the first target genomic locus comprises all or a portion of a gene encoding the first autoantigen, the first guide RNA recognition site comprises a start codon of the gene encoding the first autoantigen, the second guide RNA recognition site comprises a stop codon of the gene encoding the first autoantigen, and the modification comprises a homozygous deletion of all or a portion of the gene encoding the first autoantigen, thereby eliminating expression of the first autoantigen. Optionally, the mouse comprises: (a) an ectopic nucleic acid sequence encoding a mouse ADAM6 protein, its ortholog, its homolog, or a fragment thereof, wherein the ADAM6 protein, its ortholog, its homolog, or its fragment is functional in the male mouse; (b) a hybrid heavy chain locus comprising an insertion of human immunoglobulin heavy chain V, D, and J gene segments, wherein the human heavy chain immunoglobulin V, D, and J gene segments are operably linked to a mouse immunoglobulin heavy chain gene, and the mouse immunoglobulin heavy chain gene is operably linked to an endogenous mouse immunoglobulin gene; (b) a hybrid heavy chain locus comprising a human variable region operably linked to a mouse constant region; and (c) a hybrid light chain locus comprising an insertion of human immunoglobulin light chain V and J gene segments, wherein the human V and J gene segments are operably linked to a mouse immunoglobulin light chain constant region gene sequence, wherein (b) rearranges to form a hybrid heavy chain sequence comprising a human variable region operably linked to a mouse constant region; and (c) rearranges to form a hybrid light chain sequence comprising a human variable region operably linked to a mouse constant region; and the mouse is unable to form antibodies comprising a human variable region and a human constant region.Optionally, the mouse contains, within its germline, an endogenous mouse κ immunoglobulin light chain capable of rearranged Vκ / Jκ sequences comprising: (a) an ectopic nucleic acid sequence encoding a mouse ADAM6 protein, its ortholog, its homolog, or a fragment thereof, wherein the ADAM6 protein, its ortholog, its homolog, or a fragment thereof is functional in the male mouse; and (b) an endogenous mouse κ immunoglobulin light chain capable of rearranged Vκ / Jκ sequences comprising: (i) a single human germline Vκ sequence, wherein the single human germline Vκ sequence is present in SEQ ID NO: 148 or SEQ ID NO: 149, and (ii) a single human germline Jκ sequence. and (c) heterozygous or homozygous for an insertion at the variable region locus (wherein the rearranged Vκ / Jκ sequence is operably linked to the endogenous mouse κ constant region), and (d) for an insertion of multiple human immunoglobulin heavy chain variable region gene segments at the endogenous mouse immunoglobulin heavy chain variable region locus (wherein the human immunoglobulin heavy chain variable region gene segments are operably linked to the endogenous mouse immunoglobulin heavy chain constant region, and the human immunoglobulin heavy chain variable region gene segments are capable of rearranging and forming a rearranged human / mouse chimeric immunoglobulin heavy chain gene).

[0047] In some such methods, the non-human animal is a mouse at least partially derived from the BALB / c strain, the mouse comprises a humanized immunoglobulin locus, the foreign antigen of interest is all or a portion of a human protein orthologous to a first autoantigen, the first target genomic locus comprises all or a portion of a gene encoding the first autoantigen, the first guide RNA recognition site comprises a start codon of the gene encoding the first autoantigen, the second guide RNA recognition site comprises a stop codon of the gene encoding the first autoantigen, and the modification comprises a homozygous disruption of the start codon of the gene encoding the first autoantigen, thereby eliminating expression of the first autoantigen. Optionally, the mouse comprises: (a) an ectopic nucleic acid sequence encoding a mouse ADAM6 protein, its ortholog, its homolog, or a fragment thereof, wherein the ADAM6 protein, its ortholog, its homolog, or its fragment is functional in the male mouse; (b) a hybrid heavy chain locus comprising an insertion of human immunoglobulin heavy chain V, D, and J gene segments, wherein the human heavy chain immunoglobulin V, D, and J gene segments are operably linked to a mouse immunoglobulin heavy chain gene, and the mouse immunoglobulin heavy chain gene is operably linked to an endogenous mouse immunoglobulin gene; (b) a hybrid heavy chain locus comprising a human variable region operably linked to a mouse constant region; and (c) a hybrid light chain locus comprising an insertion of human immunoglobulin light chain V and J gene segments, wherein the human V and J gene segments are operably linked to a mouse immunoglobulin light chain constant region gene sequence, wherein (b) rearranges to form a hybrid heavy chain sequence comprising a human variable region operably linked to a mouse constant region; and (c) rearranges to form a hybrid light chain sequence comprising a human variable region operably linked to a mouse constant region; and the mouse is unable to form antibodies comprising a human variable region and a human constant region.Optionally, the mouse contains, within its germline, an endogenous mouse κ immunoglobulin light chain capable of rearranged Vκ / Jκ sequences comprising: (a) an ectopic nucleic acid sequence encoding a mouse ADAM6 protein, its ortholog, its homolog, or a fragment thereof, wherein the ADAM6 protein, its ortholog, its homolog, or a fragment thereof is functional in the male mouse; and (b) an endogenous mouse κ immunoglobulin light chain capable of rearranged Vκ / Jκ sequences comprising: (i) a single human germline Vκ sequence, wherein the single human germline Vκ sequence is present in SEQ ID NO: 148 or SEQ ID NO: 149, and (ii) a single human germline Jκ sequence. and (c) heterozygous or homozygous for an insertion at the variable region locus (wherein the rearranged Vκ / Jκ sequence is operably linked to the endogenous mouse κ constant region), and (d) for an insertion of multiple human immunoglobulin heavy chain variable region gene segments at the endogenous mouse immunoglobulin heavy chain variable region locus (wherein the human immunoglobulin heavy chain variable region gene segments are operably linked to the endogenous mouse immunoglobulin heavy chain constant region, and the human immunoglobulin heavy chain variable region gene segments are capable of rearranging and forming a rearranged human / mouse chimeric immunoglobulin heavy chain gene).

[0048] In some methods, the non-human animal one-cell stage embryo is a hybrid one-cell stage embryo, and the method further comprises (a') comparing the sequences of the corresponding first and second chromosomes of a homologous chromosome pair within the first target genomic locus to select a target region within the first target genomic locus, and then performing the contacting step (a) based on the target region having a higher percentage of sequence identity between the corresponding first and second chromosomes of the homologous chromosome pair compared to all or a portion of the remainder of the first target genomic locus. Optionally, the target region has a higher percentage of sequence identity between the corresponding first and second chromosomes of the homologous chromosome pair compared to the remainder of the first target genomic locus. Optionally, the target region has at least 99.9% sequence identity between the corresponding first and second chromosomes, and the remainder of the first target genomic locus has 99.8% or less sequence identity between the corresponding first and second chromosomes. Optionally, the target region is identical in corresponding first and second chromosomes of a homologous chromosome pair. Optionally, the target region is within the longest possible stretch of contiguous allelic sequence identity within the first target genomic locus.

[0049] In some such methods, the target region comprises a first guide RNA recognition sequence and at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb of flanking sequence on the 5', 3', or both sides of the first guide RNA recognition sequence, and and at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb of flanking sequence on the 5', 3', or both sides of the second guide RNA recognition sequence.Optionally, step (a') comprises comparing two or more segments of the first target genomic locus, each segment comprising a different guide RNA recognition sequence that is not present elsewhere in the genome, and at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, or 1,000 bp 5', 3', or both of the different guide RNA recognition sequences. selecting as target regions two segments that have the highest percentage of sequence identity relative to other segments (e.g., comprising, consisting essentially of, or consisting of 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb of contiguous sequence). Optionally, the one or more segments comprise, consist essentially of, or consist of segments that correspond to respective different guide RNA recognition sequences in the first target genomic locus but are not present elsewhere in the genome.

[0050] In some such methods, the target region comprises, consists essentially of, or consists of the region between the first guide RNA recognition sequence and the second guide RNA recognition sequence. Optionally, step (a') comprises comparing two or more segments of the first target genomic locus, each segment comprising, consisting essentially of, or consisting of a region between a different guide RNA recognition sequence pair, where the guide RNA recognition sequence is not present elsewhere in the genome, and selecting as the target region the segment having the highest percentage of sequence identity relative to the other segments. Optionally, the one or more segments comprise, consist essentially of, or consist of segments corresponding to each different guide RNA recognition sequence pair in the first target genomic locus, where the guide RNA recognition sequence is not present elsewhere in the genome.

[0051] In some such methods, the target region is at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, or more of the region between the first guide RNA recognition sequence and the second guide RNA recognition sequence, and at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, or more of the region 5', 3', or both of the genomic region between the first guide RNA recognition sequence and the second guide RNA recognition sequence. bp, 800 bp, 900 bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb or 150 kb of contiguous sequence. Optionally, step (a') comprises comparing two or more segments of the first target genomic locus (each segment being at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 ... selecting as the target region the segment having the highest percentage of sequence identity relative to other segments (e.g., 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150 kb of contiguous sequence, wherein the guide RNA recognition sequence is not present elsewhere in the genome). Optionally, the one or more segments comprise, consist essentially of, or consist of segments corresponding to each different guide RNA recognition sequence pair within the first target genomic locus, wherein the guide RNA recognition sequence is not present elsewhere in the genome.

[0052] In some such methods, the target region comprises, consists essentially of, or consists of at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb of contiguous sequence on the 5', 3', or both sides of the genomic region between the first guide RNA recognition sequence and the second guide RNA recognition sequence. Optionally, step (a') compares two or more non-contiguous segments of the first target genomic locus (each non-contiguous segment being at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 ... selecting as the target region the non-contiguous segment that has the highest percentage of sequence identity compared to other non-contiguous segments (e.g., 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb of contiguous sequence, wherein the guide RNA recognition sequence is not present elsewhere in the genome). Optionally, the one or more non-contiguous segments comprise, consist essentially of, or consist of non-contiguous segments corresponding to each different guide RNA recognition sequence pair within the first target genomic locus, wherein the guide RNA recognition sequence is not present elsewhere in the genome.

[0053] In some such methods, the target region comprises, consists essentially of, or consists of at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb of flanking sequence on either side of the genomic region between the first guide RNA recognition sequence and the second guide RNA recognition sequence. Optionally, step (a') compares two or more non-contiguous segments of the first target genomic locus (each non-contiguous segment being at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 ... selecting as the target region the non-contiguous segment that has the highest percentage of sequence identity compared to other non-contiguous segments (e.g., 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb of contiguous sequence, wherein the guide RNA recognition sequence is not present elsewhere in the genome). Optionally, the one or more non-contiguous segments comprise, consist essentially of, or consist of non-contiguous segments corresponding to each different guide RNA recognition sequence pair within the first target genomic locus, wherein the guide RNA recognition sequence is not present elsewhere in the genome.

[0054] In some such methods, the target region of step (a') comprises, consists essentially of, or consists of a region flanking the 5' target sequence and the 3' target sequence. In some such methods, the target region of step (a') comprises, consists essentially of, or consists of a region flanking and containing the 5' target sequence and the 3' target sequence. In some such methods, the target region of step (a') comprises, consists essentially of, or consists of the 5' target sequence and / or the 3' target sequence. Optionally, the target genomic locus of step (a') comprises, consists essentially of, or consists of the 5' target sequence and the 3' target sequence. In some such methods, the target region of step (a') is at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, or a region between the 5' and 3' target sequences and on the 5', 3', or both sides of the region between the 5' and 3' target sequences. bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb or 150 kb of contiguous sequence. In some such methods, the target region of step (a') is at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102 000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb or 150 kb of contiguous sequence.In some such methods, the target region of step (a') comprises, consists essentially of, or consists of at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb of flanking sequence on the 5', 3', or both sides of the region between the 5' target sequence and the 3' target sequence. In some such methods, the target region of step (a') comprises, consists essentially of, or consists of at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb of flanking sequence on either side of the region between the 5' target sequence and the 3' target sequence.

[0055] In another aspect, (a) (i) introducing into a non-human animal one-cell stage embryo or into a non-human animal pluripotent cell that is not a one-cell stage embryo: (I) Cas9 protein; (II) a first guide RNA that hybridizes to a first guide RNA recognition sequence within a target genomic locus, where the target genomic locus includes all or a portion of a gene encoding an autoantigen that is homologous to or shares an epitope of interest with the autoantigen of interest; and (III) a second guide RNA that hybridizes to a second guide RNA recognition sequence within the target genomic locus, where the target genomic locus is modified within a corresponding pair of first and second chromosomes to produce a modified non-human animal one-cell stage embryo or a modified non-human animal pluripotent cell having a biallelic modification wherein expression of the autoantigen is eliminated; and (ii) introducing into a modified non-human animal one-cell stage embryo or into a non-human animal pluripotent cell that is not a one-cell stage embryo: (I) Cas9 protein; (II) a first guide RNA that hybridizes to a first guide RNA recognition sequence within a target genomic locus, where the target genomic locus includes all or a portion of a gene encoding an autoantigen that is homologous to or shares an epitope of interest with the autoantigen of interest; and (III) a second guide RNA that hybridizes to a second guide RNA recognition sequence within the target genomic locus, where the target genomic locus is modified within a corresponding pair of first and second chromosomes to produce a modified non-human animal one-cell stage embryo or a modified non-human animal pluripotent cell having a biallelic modification wherein expression of the autoantigen is eliminated. (b) immunizing the genetically modified F0 generation non-human animal produced in step (a) with the foreign antigen of interest; and (c) maintaining the genetically modified F0 generation non-human animal under conditions sufficient to mount an immune response to the foreign antigen of interest, wherein the genetically modified F0 generation non-human animal produces an antigen-binding protein against the foreign antigen of interest.

[0056] In some methods, the cells in step (a)(i) are pluripotent stem cells of a non-human animal, and producing the transgenic F0 generation non-human animal in step (a)(ii) comprises (I) introducing the modified non-human animal pluripotent cells into a host embryo, and (II) implanting the host embryo into a surrogate mother to produce a transgenic F0 generation non-human animal in which the target genomic locus has been modified within a corresponding pair of first and second chromosomes to eliminate expression of an autoantigen. Optionally, the pluripotent cells are embryonic stem (ES) cells. In some methods, the cells in step (a)(i) are non-human animal one-cell stage embryos, and producing the transgenic F0 generation non-human animal in step (a)(ii) comprises implanting the modified non-human animal one-cell stage embryo into a surrogate mother to produce a transgenic F0 generation non-human animal in which the target genomic locus has been modified within a corresponding pair of first and second chromosomes to eliminate expression of an autoantigen.

[0057] Some such methods further include generating hybridomas from B cells isolated from the immunized transgenic F0 generation non-human animal. Some such methods further include obtaining from the immunized transgenic F0 generation non-human animal a first nucleic acid sequence encoding an immunoglobulin heavy chain variable domain of one of the antigen binding proteins for the foreign antigen of interest and / or a second nucleic acid sequence encoding an immunoglobulin light chain variable domain of one of the antigen binding proteins for the foreign antigen of interest. Optionally, the first nucleic acid sequence and / or the second nucleic acid sequence are obtained from lymphocytes (e.g., B cells) of the transgenic F0 generation non-human animal or from hybridomas generated from the lymphocytes. Optionally, the transgenic F0 generation non-human animal comprises a humanized immunoglobulin locus, wherein the first nucleic acid sequence encodes a human immunoglobulin heavy chain variable domain and the second nucleic acid sequence encodes a human immunoglobulin light chain variable domain.

[0058] In some such methods, the antigen-binding proteins against the foreign antigen of interest produced by the genetically modified F0 generation non-human animals have higher titers than the antigen-binding proteins produced by a control non-human animal that is wild-type at the target genomic locus after immunization with the foreign antigen of interest. In some such methods, a more diverse repertoire of antigen-binding proteins against the foreign antigen of interest is produced by the genetically modified F0 generation non-human animal after immunization with the foreign antigen of interest compared to the antigen-binding proteins produced by the control non-human animal that is wild-type at the target genomic locus after immunization with the foreign antigen of interest. In some such methods, the antigen-binding proteins against the foreign antigen of interest produced by the genetically modified F0 generation non-human animals use a greater diversity of heavy chain V gene segments and / or light chain V gene segments than the antigen-binding proteins produced by the control non-human animal that is wild-type at the target genomic locus after immunization with the foreign antigen of interest. In some such methods, some of the antigen binding proteins produced by the transgenic F0 generation non-human animals against the foreign antigen of interest cross-react with self-antigens.

[0059] In some such methods, the first guide RNA recognition sequence is 5' to the second guide RNA recognition sequence within the target genomic locus, and step (a)(i) further comprises performing a retention assay to confirm that the copy number within about 1 kb of the 5' region and the first guide RNA recognition sequence and / or within about 1 kb of the 3' region and the second guide RNA recognition sequence is 2.

[0060] In some such methods, the foreign antigen of interest is an ortholog of a self antigen. In some such methods, the foreign antigen of interest comprises all or a portion of a human protein.

[0061] In some such methods, the target genomic locus is modified to include an insertion of one or more nucleotides, a deletion of one or more nucleotides, or a substitution of one or more nucleotides. Optionally, the deletion is a precise deletion without random insertions and deletions (indels).

[0062] In some such methods, the first guide RNA recognition sequence includes the start codon of the gene encoding the autoantigen or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the start codon, and the second guide RNA recognition sequence includes the stop codon of the gene encoding the autoantigen or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the stop codon. In some such methods, the first and second guide RNA recognition sequences are different, and each of the first and second guide RNA recognition sequences includes the start codon of the gene encoding the autoantigen or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the start codon.

[0063] In some such methods, the target genomic locus is modified to include a biallelic deletion of about 0.1 kb to about 200 kb. In some such methods, the modification includes a biallelic deletion of all or part of the gene encoding the autoantigen. In some such methods, the modification includes a biallelic disruption of the start codon of the gene encoding the autoantigen.

[0064] In some such methods, the introducing step (a)(i) further comprises introducing into the non-human animal pluripotent cell or non-human animal one-cell stage embryo: (iv) a third guide RNA that hybridizes to a third guide RNA recognition sequence in the target genomic locus, and / or (v) a fourth guide RNA that hybridizes to a fourth guide RNA recognition sequence in the target genomic locus.

[0065] In some such methods, the cell in step (a)(i) is a non-human animal pluripotent stem cell, and the Cas9 protein, first guide RNA, and second guide RNA are each introduced into the non-human animal pluripotent stem cell in the form of DNA. In some such methods, the cell in step (a)(i) is a non-human animal pluripotent stem cell, and the Cas9 protein, first guide RNA, and second guide RNA are each introduced into the non-human animal pluripotent stem cell by electroporation or nucleofection. In some such methods, the cell in step (a)(i) is a non-human animal one-cell embryo, and the Cas9 protein, first guide RNA, and second guide RNA are each introduced into the non-human animal one-cell embryo in the form of RNA. In some such methods, the cell in step (a)(i) is a non-human animal one-cell embryo, and the Cas9 protein, first guide RNA, and second guide RNA are introduced into the non-human animal one-cell embryo by pronuclear injection or cytoplasmic injection.

[0066] In some such methods, no exogenous repair template is introduced in step (a)(i). In some such methods, the introducing step (a)(i) further comprises, when the cell of step (a)(i) is a non-human animal one-cell stage embryo and the length of the exogenous repair template is about 5 kb or less, introducing into the non-human animal pluripotent cell or the non-human animal one-cell stage embryo an exogenous repair template comprising a 5' homology arm that hybridizes to a 5' target sequence in the target genomic locus and a 3' homology arm that hybridizes to a 3' target sequence in the target genomic locus. Optionally, the exogenous repair template further comprises a nucleic acid insert flanked by the 5' homology arm and the 3' homology arm. Optionally, the nucleic acid insert is homologous or orthologous to the target genomic locus. Optionally, the exogenous repair template is about 50 nucleotides to about 1 kb in length. Optionally, the exogenous repair template is about 80 nucleotides to about 200 nucleotides in length. Optionally, the foreign repair template is a single-stranded oligodeoxynucleotide. Optionally, the cell in step (a)(i) is a non-human animal pluripotent cell, and (a) the foreign repair template is a large targeting vector (LTVEC) at least 10 kb in length, or (b) the foreign repair template is an LTVEC (the total length of the 5' homology arm and 3' homology arm of the LTVEC is at least 10 kb). Optionally, the target genomic locus is modified to include a deletion of one or more nucleotides, and the deleted nucleic acid sequence consists of a nucleic acid sequence between the 5' target sequence and the 3' target sequence. Optionally, the foreign repair template includes a nucleic acid insert flanked by the 5' homology arm and the 3' homology arm, and the nucleic acid insert is homologous or orthologous to the deleted nucleic acid sequence, and the target genomic locus is modified to include a deletion of one or more nucleotides, and the nucleic acid insert replaces the deleted nucleic acid sequence.

[0067] In some such methods, the non-human animal comprises humanized immunoglobulin loci. In some such methods, the non-human animal is a rodent. Optionally, the rodent is a mouse. Optionally, the mouse strain comprises a BALB / c strain. Optionally, the mouse strain comprises a BALB / c strain, a C57BL / 6 strain, and a 129 strain. Optionally, the mouse strain is 50% BALB / c, 25% C57BL / 6, and 25% 129. Optionally, the MHC haplotype of the mouse is MHC b / d is.

[0068] In some such methods, the mouse comprises in its germline a human unrearranged variable region gene segment inserted into an endogenous mouse immunoglobulin locus. Optionally, the human unrearranged variable region gene segment is a heavy chain gene segment, and the mouse immunoglobulin locus is a heavy chain locus, and / or the human unrearranged variable region gene segment is a kappa or lambda light chain segment, and the mouse immunoglobulin locus is a light chain locus. Optionally, the mouse comprises in its germline a human unrearranged variable region gene segment operably linked to a mouse constant region gene, and the mouse lacks a human constant region gene, and the mouse constant region gene is present at the endogenous mouse immunoglobulin locus. Optionally, the mouse comprises (a) a hybrid heavy chain locus comprising an insertion of human immunoglobulin heavy chain V, D, and J gene segments, wherein the human heavy chain immunoglobulin V, D, and J gene segments are operably linked to a mouse immunoglobulin heavy chain gene, and the mouse immunoglobulin heavy chain gene is present at the endogenous mouse immunoglobulin locus, and (b) a hybrid light chain locus comprising an insertion of human immunoglobulin light chain V and J gene segments, wherein the human V and J gene segments are operably linked to a mouse immunoglobulin light chain constant region gene sequence, wherein (a) rearranges to form a hybrid heavy chain sequence comprising a human variable region operably linked to a mouse constant region, and (b) rearranges to form a hybrid light chain sequence comprising a human variable region operably linked to a mouse constant region, and the mouse is unable to form antibodies comprising the human variable region and the human constant region.

[0069] In some such methods, the mouse comprises in its germline a humanized immunoglobulin light chain variable locus comprising no more than one or no more than two rearranged human light chain V / J sequences operably linked to a mouse light chain constant region, and the mouse further comprises a humanized immunoglobulin heavy chain variable locus comprising at least one unrearranged human V segment, at least one unrearranged human D segment, and at least one unrearranged human J segment operably linked to a mouse heavy chain constant region gene. Optionally, the mouse comprises a humanized heavy chain immunoglobulin variable locus and a humanized light chain immunoglobulin variable locus, and the mouse expresses a single light chain. Optionally, the mouse comprises: (a) a human V of an immunoglobulin light chain; L A single rearranged human immunoglobulin light chain variable region (V L / J L )(Single reconstituted human V L / J L The region is selected from a human Vκ1-39 / Jκ5 gene segment or a human Vκ3-20 / Jκ1 gene segment), and (b) an endogenous heavy chain variable (V H ) gene segment with one or more human VH gene segments (human V H The gene segment is an endogenous heavy chain constant (C H ) region gene and functionally linked to the human V HOptionally, the mouse expresses a population of antibodies, and the germline of the mouse comprises only a single immunoglobulin kappa light chain variable region gene that is a rearranged human germline kappa light chain variable region gene, and the mouse is either heterozygous for the single immunoglobulin kappa light chain variable region gene, in that it contains only one copy, or homozygous for the single immunoglobulin kappa light chain variable region gene, in that it contains two copies, and the mouse is further characterized by: (i) each immunoglobulin kappa light chain in the population is a rearranged human germline kappa light chain variable region gene; (ii) the population includes antibodies comprising immunoglobulin kappa light chains whose light chain variable domains are encoded by rearranged human germline kappa light chain variable region genes, and antibodies comprising immunoglobulin kappa light chains whose light chain variable domains are encoded by somatic variants thereof; and (iii) the mice are characterized by active affinity maturation to produce a diverse population of somatically mutated high-affinity heavy chains that successfully pair with immunoglobulin kappa light chains to form antibodies in the population. Optionally, the mouse is heterozygous or homozygous within its germline for: (a) an insertion at the endogenous mouse κ immunoglobulin light chain variable region locus of a rearranged Vκ / Jκ sequence comprising: (i) a single human germline Vκ sequence (the single human germline Vκ sequence is present in SEQ ID NO: 148 or SEQ ID NO: 149), and (ii) a single human germline Jκ sequence, wherein the rearranged Vκ / Jκ sequence is operably linked to an endogenous mouse κ constant region; and (b) an insertion at the endogenous mouse immunoglobulin heavy chain variable region locus of a plurality of human immunoglobulin heavy chain variable region gene segments, wherein the human immunoglobulin heavy chain variable region gene segments are operably linked to an endogenous mouse immunoglobulin heavy chain constant region, and wherein the human immunoglobulin heavy chain variable region gene segments are capable of rearranging and forming a rearranged human / mouse chimeric immunoglobulin heavy chain gene.

[0070] In some such methods, the mouse comprises a modification of the immunoglobulin heavy chain locus, the modification suppressing or eliminating endogenous ADAM6 function, the mouse comprises an ectopic nucleic acid sequence encoding the mouse ADAM6 protein, its ortholog, its homolog, or a fragment thereof, the ADAM6 protein, its ortholog, its homolog, or a fragment thereof is functional in the male mouse, and the ectopic nucleic acid sequence encoding the mouse ADAM6 protein, its ortholog, its homolog, or a fragment thereof is present in a human heavy chain variable region locus.

[0071] In some such methods, the non-human animal is a mouse at least partially derived from the BALB / c strain, the mouse comprises humanized immunoglobulin loci, the foreign antigen of interest is all or a portion of a human protein orthologous to an autoantigen, the first guide RNA recognition sequence includes the start codon of a gene encoding the autoantigen or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the start codon, the second guide RNA recognition sequence includes the stop codon of a gene encoding the autoantigen or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the stop codon, and the modification comprises a biallelic deletion of all or a portion of the gene encoding the autoantigen, thereby eliminating expression of the autoantigen. In some such methods, the non-human animal is a mouse at least partially derived from the BALB / c strain, the mouse comprises a humanized immunoglobulin locus, the foreign antigen of interest is all or a portion of a human protein orthologous to an autoantigen, the first guide RNA recognition sequence comprises the start codon of a gene encoding the autoantigen, the second guide RNA recognition sequence comprises the stop codon of the gene encoding the autoantigen or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the start codon, and the modification comprises a biallelic disruption of the start codon of the gene encoding the autoantigen, thereby eliminating expression of the autoantigen.Optionally, the mouse comprises: (a) an ectopic nucleic acid sequence encoding a mouse ADAM6 protein, its ortholog, its homolog, or a fragment thereof, wherein the ADAM6 protein, its ortholog, its homolog, or its fragment is functional in the male mouse; (b) a hybrid heavy chain locus comprising an insertion of human immunoglobulin heavy chain V, D, and J gene segments, wherein the human heavy chain immunoglobulin V, D, and J gene segments are operably linked to a mouse immunoglobulin heavy chain gene, and the mouse immunoglobulin heavy chain gene is operably linked to an endogenous mouse immunoglobulin gene; (b) a hybrid heavy chain locus comprising a human variable region operably linked to a mouse constant region; and (c) a hybrid light chain locus comprising an insertion of human immunoglobulin light chain V and J gene segments, wherein the human V and J gene segments are operably linked to a mouse immunoglobulin light chain constant region gene sequence, wherein (b) rearranges to form a hybrid heavy chain sequence comprising a human variable region operably linked to a mouse constant region; and (c) rearranges to form a hybrid light chain sequence comprising a human variable region operably linked to a mouse constant region; and the mouse is unable to form antibodies comprising a human variable region and a human constant region. Optionally, the mouse contains, within its germline, an endogenous mouse κ immunoglobulin light chain capable of rearranged Vκ / Jκ sequences comprising: (a) an ectopic nucleic acid sequence encoding a mouse ADAM6 protein, its ortholog, its homolog, or a fragment thereof, wherein the ADAM6 protein, its ortholog, its homolog, or a fragment thereof is functional in the male mouse; and (b) an endogenous mouse κ immunoglobulin light chain capable of rearranged Vκ / Jκ sequences comprising: (i) a single human germline Vκ sequence, wherein the single human germline Vκ sequence is present in SEQ ID NO: 148 or SEQ ID NO: 149, and (ii) a single human germline Jκ sequence. and (c) heterozygous or homozygous for an insertion at the variable region locus (wherein the rearranged Vκ / Jκ sequence is operably linked to the endogenous mouse κ constant region), and (d) for an insertion of multiple human immunoglobulin heavy chain variable region gene segments at the endogenous mouse immunoglobulin heavy chain variable region locus (wherein the human immunoglobulin heavy chain variable region gene segments are operably linked to the endogenous mouse immunoglobulin heavy chain constant region, and the human immunoglobulin heavy chain variable region gene segments are capable of rearranging and forming a rearranged human / mouse chimeric immunoglobulin heavy chain gene).

[0072] In some such methods, the non-human animal pluripotent cell is a hybrid cell or the non-human mammalian one-cell stage embryo is a hybrid one-cell stage embryo, and the method further comprises (a') comparing sequences of a pair of corresponding first and second chromosomes within the target genomic locus to select a target region within the target genomic locus, and then performing the contacting step (a) based on a target region having a higher percentage of sequence identity between the pair of corresponding first and second chromosomes compared to all or a portion of the remainder of the target genomic locus, wherein the target region comprises a first guide RNA recognition sequence and at least 10 bp, 20 bp, or 3'-bp sequences on either the 5', 3', or both sides of the first guide RNA recognition sequence. 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb or 10 kb of flanking sequence and / or a second guide RNA recognition sequence and a second guide RNA recognition sequence. and at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, or 10 kb of flanking sequence on the 5', 3', or both sides of the sequence. Optionally, the target region has a higher percentage of sequence identity between the corresponding first and second pair compared to the remainder of the target genomic locus. Optionally, the target region has at least 99.9% sequence identity between the corresponding first and second chromosome pair, and the remainder of the target genomic locus has 99.8% or less sequence identity between the corresponding first and second chromosome pair.

[0073] In another aspect, (a) a non-human animal one-cell stage embryo or a non-human animal pluripotent cell that is not a one-cell stage embryo is transfected with (i) a Cas9 protein, (ii) a first guide RNA that hybridizes to a first guide RNA recognition sequence within a target genomic locus (the target genomic locus includes all or a portion of a gene encoding an autoantigen that is homologous to a foreign antigen of interest or that shares an epitope of interest with the foreign antigen of interest), and (iii) a second guide RNA that hybridizes to a second guide RNA recognition sequence within the target genomic locus (the target genomic locus is a pair of corresponding first and second chromosomes). and (b) introducing a target genomic locus into a genetically modified non-human animal to produce a modified non-human animal one-cell stage embryo or a modified non-human animal pluripotent cell having a biallelic modification, wherein expression of the self-antigen is eliminated; and (b) generating a genetically modified F0 generation non-human animal from the modified non-human animal one-cell stage embryo or the modified non-human animal pluripotent cell, wherein the target genomic locus is modified within a corresponding pair of first and second chromosomes in the genetically modified F0 generation non-human animal to eliminate expression of the self-antigen.

[0074] Such methods may include, for example, any of the variations disclosed above as methods for producing an antigen-binding protein against a foreign antigen of interest. For example, in some such methods, the cells in step (a) are pluripotent stem cells of a non-human animal, and producing a genetically modified F0 generation non-human animal in step (b) includes (I) introducing the modified non-human animal pluripotent cells into a host embryo, and (II) implanting the host embryo into a surrogate mother to produce a genetically modified F0 generation non-human animal in which the target genomic locus has been modified within a pair of corresponding first and second chromosomes so as to eliminate expression of the self-antigen. Optionally, the pluripotent cells are embryonic stem (ES) cells. In some such methods, the cell in step (a) is a non-human animal one-cell stage embryo, and producing the genetically modified F0 generation non-human animal in step (b) comprises implanting the modified non-human animal one-cell stage embryo into a surrogate mother to produce a genetically modified F0 generation non-human animal in which the target genomic locus has been modified within a corresponding pair of first and second chromosomes such that expression of the self-antigen is eliminated. In some such methods, the foreign antigen of interest is an ortholog of the self-antigen. In some such methods, the first guide RNA recognition sequence includes the start codon of the gene encoding the autoantigen or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the start codon, and the second guide RNA recognition sequence includes the stop codon of the gene encoding the autoantigen or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the stop codon. In some such methods, the first and second guide RNA recognition sequences are different, and each of the first and second guide RNA recognition sequences includes the start codon of the gene encoding the autoantigen or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the start codon.In some such methods, the first guide RNA recognition sequence is 5' to the second guide RNA recognition sequence within the target genomic locus, and step (a)(i) further comprises performing a retention assay to confirm that the copy number within about 1 kb of the 5' region and the first guide RNA recognition sequence and / or within about 1 kb of the 3' region and the second guide RNA recognition sequence is 2. In some such methods, the modification comprises a biallelic deletion of all or part of the gene encoding the autoantigen. In some such methods, the modification comprises a biallelic disruption of the start codon of the gene encoding the autoantigen. In some such methods, the non-human animal is a mouse. [Brief explanation of the drawings]

[0075] [Figure 1] This figure shows a conventional approach to breaking immune tolerance in VELOCIMMUNE® mice (VI-3; homozygous humanized for both IgH and Igκ). In the conventional approach, a heterozygous knockout (null) allele of a gene encoding a self-antigen homologous to a foreign target antigen of interest is generated in F1H4 embryonic stem (ES) cells. The time from designing the targeting vector to generating F0 mice heterozygous for the knockout takes approximately five months. The VI-3 mice are then bred into F0 mice carrying a heterozygous knockout mutation in the endogenous gene encoding the self-antigen homologous to the foreign target antigen of interest. Two additional generations of breeding are required to generate triple homozygous mice suitable for immunization (homozygous null for the target of interest and homozygous humanized for both IgH and Igκ). The entire process from designing the targeting vector to generating triple homozygous mice takes approximately 15–16 months. [Figure 2]

[0023] Figure 1 shows an accelerated process for breaking immune tolerance in VELOCIMMUNE® (VI-3) or universal light chain (ULC or common light chain) mice. In this process, ES cells derived from VI-3 or ULC mice are targeted to generate heterozygous null alleles of endogenous genes encoding self-antigens homologous to the foreign target antigen of interest. Successive targeting steps are required to obtain homozygous null VI-3 or ULC ES cell clones. [Figure 3]

[0033] Figure 1 shows a further accelerated process for breaking tolerance in VELOCIMMUNE® (VI-3) or universal light chain (or common light chain) (ULC) mice. In this process, VI-3 or ULC ES cells are targeted with CRISPR / Cas9 and paired guide RNAs to generate, in a single step, homozygous disruption of an endogenous gene encoding a self-antigen homologous to a foreign target antigen of interest. TAQMAN® screening may include, for example, both allelic loss and retention assays. [Figure 4] Figure 1 shows a general outline of the simultaneous deletion of a mouse gene encoding a self-antigen homologous to a foreign target antigen of interest and its replacement with a neomycin-selectable marker using a large targeting vector (LTVEC) and paired upstream and downstream guide RNAs (gU and gD). The locations of the Cas9 cleavage sites guided by the two guide RNAs are indicated by arrows below the mouse gene sequence. TAQMAN® assay probes are indicated by horizontal lines and include retention assay probes and loss of allele (LOA) assay probes for the upstream, middle, and downstream alleles. The bottom of the diagram shows the expected target allele types. [Figure 5]

[0023] Figure 1 shows a general outline of the simultaneous deletion of a mouse gene encoding a self-antigen homologous to a foreign target antigen of interest and its replacement with a loxP-flanked neomycin selectable marker and lacZ using a large targeting vector (LTVEC) and three overlapping guide RNAs, each targeting the mouse ATG start codon. The guide RNAs are indicated by horizontal arrows, and the TAQMAN® assay probe is indicated by a circled horizontal line. The bottom of the figure shows the predicted target allele type. [Figure 6] FIG. 1 shows antibody titer data for human target antigen (target 8) in wild-type universal light chain (ULC 1-39) mice and ULC 1-39 mice homozygous null for the endogenous gene encoding the autoantigen orthologous to target 8 (autoantigen 8). [Figure 7] FIG. 1 shows the matings performed to generate hybrid VGF1(F1H4) ES cells (C57BL6(XB6) / 129S6(Y129)). [Figure 8] Figure 1 shows a schematic diagram of simultaneous deletion and replacement of mouse genes or portions of mouse genes with their human counterparts using LTVEC and either one or two 5', central, and 3' gRNAs. The LTVEC is shown at the top of the diagram, and the mouse locus is shown at the bottom. The locations of Cas9 cleavage sites guided by the eight guide RNAs are indicated by vertical arrows below the mouse gene sequence. [Figure 9]Figure 9A shows a general outline of simultaneous deletion of a mouse gene and replacement with the corresponding human form using LTVEC and two guide RNAs (guide RNAs A and B). The LTVEC is shown at the top of Figure 9A, and the mouse locus is shown at the bottom of Figure 9A. The location of the Cas9 cleavage site guided by the two guide RNAs is indicated by arrows below the mouse gene sequence. Figures 9B-E show unique biallelic modifications (allelic forms) that occur more frequently when using two guide RNAs. The thick, diagonally shaded line indicates the mouse gene, the dotted line indicates a deletion within the mouse gene, and the thick black line indicates an insertion into the human gene. Figure 9B shows a homozygous disrupted allele (large CRISPR-induced deletion). Figure 9C shows a homozygous targeted allele. Figure 9D shows a hemizygous targeted allele. Figure 9E shows a compound heterozygous allele. [Figure 10]Figures 10A and 10B show PCR assays to confirm the genotype of selected clones. Figure 10A shows the results of a long-range PCR assay of selected ES cell clones using primers m-lr-f and m-5'-r, which create a junction between the human insert and sequences outside the sequence homologous to the 5' homology arm, thereby achieving precise targeting. Figure 10B shows the results of PCR assays using 5'Del J, 5'Ins J, Del A + F, and Del A + E2. 5'Del J shows the PCR product using primers m-5'-f and m-5-r, which amplify the wild-type sequence surrounding the gRNA A cleavage site to confirm retention or loss of this sequence. 5'Ins J shows the PCR product using primers m-5'-f and h-5'-r, which create a junction between the human insert and the mouse genome. The assay shows positive results in both targeted and randomly integrated clones. Del A + F indicates the expected amplicon size (359 bp) and the actual band due to the large deletion associated with dual gRNA A and F cleavage in clones BO-F10 and AW-A8. Del A + E2 indicates a similar event in clone BA-A7. NT indicates no template, + / + indicates the parental VGF1 hybrid ES cell wild-type control, H / + indicates the heterozygous humanized genotype, H / Δ indicates the hemizygous humanized genotype, H / H indicates the homozygous humanized genotype, and Δ / Δ indicates the homozygous deletion genotype. [Figure 11]Figures 11A-11C show fluorescence in situ hybridization (FISH) analyses of mouse ES cell clone AW-D9 (Figure 11A) and mouse ES cell clone BA-D5 (Figure 11C) targeted by Lrp5-humanized LTVEC in combination with Cas9 and two gRNAs, and mouse ES cell clone BS-C4 (Figure 11B) targeted by LTVEC alone. Arrows indicate the location of the hybridization signal on band B of chromosome 19. Red signals indicate hybridization with the mouse probe alone (dashed arrow, Figure 11B). Yellow mixed-color signals indicate hybridization with both the red mouse probe and the green human probe. One chromosome 19 B band with a red signal (dashed arrow) and the other with a yellow signal (solid arrow) confirmed targeting to the correct locus and the heterozygous genotype in the BS-C4 clone (Figure 11B). Both chromosome 19 B bands with yellow signals (solid arrows, Figures 11A and 11C) confirmed targeting to the correct locus and the homozygous genotype in the AW-D9 and BS-C4 clones. [Figure 12] Figure 1 shows a schematic of chromosome 19 with an assay designed to confirm gene conversion or mitotic recombination events involving two guide RNAs by analyzing loss of heterozygosity (LOH) in VGF1 hybrid ES cells. The approximate locations of the TAQMAN® qPCR chromosome copy number (CCN) probes are indicated by arrows. The approximate locations of structural variant (SV) polymorphism PCR probes are indicated by chevrons (top) indicating their distance (Mb) from the Lrp5 locus. The approximate locations of single nucleotide polymorphism (SNV) TAQMAN® allele-discriminating probes are indicated by arrowheads (bottom) indicating their distance (Mb) from the Lrp5 locus. The locations of gRNA recognition sequences for F, E2, D, B2, and A are indicated by diagonal arrows above the Lrp5 gene designation. [Figure 13]Figures 13A and 13B show fluorescence in situ hybridization (FISH) analyses of mouse ES cell clone Q-E9 (Figure 13A) and mouse ES cell clone O-E3 (Figure 13B) targeted by Hc-humanized LTVEC in combination with Cas9 and two gRNAs. The arrows indicate the location of the hybridization signal on chromosome 2 band B. The red signal indicates hybridization with the mouse probe only (dashed arrow, Figure 13A). The yellow mixed-color signal indicates hybridization with both the red mouse probe and the green human probe (solid arrow). One chromosome 2 band B with a red signal (dashed arrow) and the other with a yellow signal (solid arrow) confirmed the correct locus targeting and heterozygous genotype in the Q-E9 clone (Figure 13A). B bands on both chromosomes 2 with yellow signals (solid arrows, FIG. 13B) confirmed targeting to the correct locus and the homozygous genotype in the O-E3 clone. [Figure 14] Figure 1 shows a schematic of the mouse C5 gene-containing chromosome with an assay designed to confirm gene conversion or mitotic recombination events involving two guide RNAs by analyzing loss of heterozygosity (LOH) in VGF1 hybrid ES cells. The approximate locations of structural variant (SV) polymorphism PCR probes are indicated by horizontal arrows (top) indicating their distance (Mb) from the C5 locus. The locations of gRNA recognition sequences for E2 and A are indicated by diagonal arrows above the C5 locus designation. [Figure 15] Figures 15A-15E show the results of structural polymorphism (SV) assays of clones BR-B4, BP-G7, BO-G11, BO-F10, B0-A8, and BC-H9, using VGF1 (F1H4), 129, and B6 DNA as controls. Assays were performed at the following distances telomeric to the Lrp5 locus: 13.7 Mb (Figure 15A), 20.0 Mb (Figure 15B), 36.9 Mb (Figure 15C), 48.3 Mb (Figure 15D), and 56.7 Mb (Figure 15E). The positions of the PCR products for the B6 and 129 alleles are indicated by arrows. [Figure 16] Figures 16A-16C show allelic discrimination plots for the centromeric 0.32 Mb of Lrp5 (Figure 16A), the telomeric 1.2 Mb of Lrp5 (Figure 16B), and the telomeric 57.2 Mb of Lrp5 (Figure 16C). Values ​​on each axis represent relative fluorescence intensity. Plots are shown for quadruplicates of each sample and are indicated by solid dots (B6 allele), empty dots (129 allele), and hatched dots (both B6 and 129 alleles). [Figure 17] Figures 17A-17C are schematics illustrating the mitotic recombination mechanisms that can occur during the G2 phase of the cell cycle, resulting in homozygous events and widespread gene conversion, as detected by loss of heterozygosity. Figure 17A shows duplicated homologous chromosomes showing two chromatids in hybrid 129 / B6 ES cells heterozygous for targeted humanization on the 129 homolog. The double arrows indicate potential double-strand breaks generated by dual gRNA-guided Cas9 cleavage, which promotes reciprocal exchange between chromatids on the homologous chromosomes by homologous recombination, as shown as crossovers on the centromeric side of the targeted allele, resulting in the hybrid chromatid shown in Figure 17B. Figure 17C shows the aftermath of mitosis and cell division, allowing for four types of chromosome segregation into daughter cells. Two chromosomes that maintained heterozygosity, a parental heterozygote (Hum / +, top left) and a balanced exchange heterozygote (Hum / +, top right), are indistinguishable by LOH assays. The other two, a humanized homozygote with loss of the telomeric B6 allele (Hum / Hum, e.g., clone BO-A8, bottom left) and a wild-type homozygote with loss of the telomeric 129 allele (+ / +, bottom right), show loss of heterozygosity, the latter type being lost due to the lack of the drug resistance cassette of the humanized allele. [Figure 18A]Figure 1 shows a possible mechanism explaining the observations, including loss of heterozygosity (LOH), in CRISPR / Cas9-assisted humanization experiments using F1 hybrid mouse ES cells harboring one haploid chromosome set derived from the 129S6 / SvEvTac mouse strain and one haploid chromosome set derived from the C57BL / 6NTac (B6) mouse strain. Heterozygous modification occurs on chromosome 129 before or after genome duplication, followed by interchromatid exchange by mitotic crossing over, which results in gene conversion between sister chromatids. [Figure 18B] Figure 1 shows a possible mechanism explaining the observations, including loss of heterozygosity (LOH), in CRISPR / Cas9-assisted humanization experiments using F1 hybrid mouse ES cells harboring one haploid chromosome set derived from the 129S6 / SvEvTac mouse strain and one haploid chromosome set derived from the C57BL / 6NTac (B6) mouse strain. Interchromatid exchange by mitotic crossing over is shown, in which a single 129 chromatid is altered after genome duplication. [Figure 18C] Figure 1 shows a possible mechanism explaining the observations, including loss of heterozygosity (LOH), in CRISPR / Cas9-assisted humanization experiments using F1 hybrid mouse ES cells harboring one haploid chromosome set from the 129S6 / SvEvTac mouse strain and one haploid chromosome set from the C57BL / 6NTac (B6) mouse strain. LTVEC targeting did not occur, but Cas9 cleavage occurred on either chromosome 129 or chromosome B6 (indicating B6 cleavage), indicating interchromatid exchange by mitotic crossing over. [Figure 18D]Figure 1 shows a possible mechanism explaining the observations, including loss of heterozygosity (LOH), in CRISPR / Cas9-assisted humanization experiments using F1 hybrid mouse ES cells harboring one haploid chromosome set derived from the 129S6 / SvEvTac mouse strain and one haploid chromosome set derived from the C57BL / 6NTac (B6) mouse strain. Chromatid duplication by break-induced duplication is shown, in which heterozygous modification occurs on chromosome 129 before or after genome duplication, followed by gene conversion between sister chromatids. [Figure 18E] Figure 1 shows a possible mechanism explaining the observations, including loss of heterozygosity (LOH), in CRISPR / Cas9-assisted humanization experiments using F1 hybrid mouse ES cells harboring one haploid chromosome set derived from the 129S6 / SvEvTac mouse strain and one haploid chromosome set derived from the C57BL / 6NTac (B6) mouse strain. Chromatid duplication by break-induced duplication is shown, in which a single 129 chromatid is altered after genome duplication. [Figure 18F] Figure 1 shows a possible mechanism explaining the observations, including loss of heterozygosity (LOH), in CRISPR / Cas9-assisted humanization experiments using F1 hybrid mouse ES cells harboring one haploid chromosome set from the 129S6 / SvEvTac mouse strain and one haploid chromosome set from the C57BL / 6NTac (B6) mouse strain, demonstrating chromatid duplication by break-induced replication, where LTVEC targeting has not occurred but Cas9 cleavage has occurred on either the 129 or B6 chromosome (indicating B6 cleavage). [Figure 19]Figure 1 shows a schematic diagram of the mouse Lrp5 locus, which was targeted for deletion and replacement with the corresponding human LRP5 locus using LTVEC and one or more gRNAs in VGF1 hybrid ES cells. The region within the vertical dotted line is the target region (regions within the 5' and 3' target sequences of LTVEC). The reference sequence for identifying single nucleotide polymorphisms was the genomic sequence of the C57BL / 6J mouse strain (Jackson Laboratory). This reference sequence was compared with the 129S6 / SvEv strain (Taconic Biosciences), the C57BL / 6N strain (Taconic Biosciences), and VGF1 hybrid cell lines generated from the 129S6 / SvEv and C57BL / 6N strains (listed in the bottom three rows of the figure). The vertical lines in each of the three rows indicate single nucleotide polymorphisms compared to the reference sequence. [Figure 20] Figure 1 shows a schematic diagram of the mouse Hc locus targeted for deletion and replacement with the corresponding humanized form using LTVEC and one or more gRNAs in VGF1 hybrid ES cells. The region within the vertical dotted line is the target region (regions within the 5' and 3' target sequences of LTVEC). The reference sequence for identifying single nucleotide polymorphisms was the genomic sequence of the C57BL / 6J mouse strain (Jackson Laboratory). This reference sequence was compared with the 129S6 / SvEv strain (Taconic Biosciences), the C57BL / 6N strain (Taconic Biosciences), and VGF1 hybrid cell lines generated from the 129S6 / SvEv and C57BL / 6N strains (listed in the bottom three rows of the figure). The vertical lines in each of the three rows indicate single nucleotide polymorphisms compared to the reference sequence. [Figure 21]Figure 1 shows a schematic diagram of the mouse Trpa1 locus, which was targeted for deletion and replacement with the corresponding human form using LTVEC and one or more gRNAs in VGF1 hybrid ES cells. The region within the vertical dotted line is the target region (regions within the 5' and 3' target sequences of LTVEC). The reference sequence for identifying single nucleotide polymorphisms was the genomic sequence of the C57BL / 6J mouse strain (Jackson Laboratory). This reference sequence was compared with the 129S6 / SvEv strain (Taconic Biosciences), the C57BL / 6N strain (Taconic Biosciences), and VGF1 hybrid cell lines generated from the 129S6 / SvEv and C57BL / 6N strains (listed in the bottom three rows of the figure). The vertical lines in each of the three rows indicate single nucleotide polymorphisms compared to the reference sequence. [Figure 22] Figure 1 shows a schematic diagram of the mouse Adamts5 locus, which was targeted for deletion and replacement with the corresponding human form using LTVEC and one or more gRNAs in VGF1 hybrid ES cells. The region within the vertical dotted line is the target region (regions within the 5' and 3' target sequences of LTVEC). The reference sequence for identifying single nucleotide polymorphisms was the genomic sequence of the C57BL / 6J mouse strain (Jackson Laboratory). This reference sequence was compared with the 129S6 / SvEv strain (Taconic Biosciences), the C57BL / 6N strain (Taconic Biosciences), and VGF1 hybrid cell lines generated from the 129S6 / SvEv and C57BL / 6N strains (listed in the bottom three rows of the figure). The vertical lines in each of the three rows indicate single nucleotide polymorphisms compared to the reference sequence. [Figure 23]Figure 1 shows a schematic diagram of the mouse Folh1 locus, which was targeted for deletion and replacement with the corresponding human form using LTVEC and one or more gRNAs in VGF1 hybrid ES cells. The region within the vertical dotted line is the target region (regions within the 5' and 3' target sequences of LTVEC). The reference sequence for identifying single nucleotide polymorphisms was the genomic sequence of the C57BL / 6J mouse strain (Jackson Laboratory). This reference sequence was compared with the 129S6 / SvEv strain (Taconic Biosciences), the C57BL / 6N strain (Taconic Biosciences), and VGF1 hybrid cell lines generated from the 129S6 / SvEv and C57BL / 6N strains (listed in the bottom three rows of the figure). The vertical lines in each of the three rows indicate single nucleotide polymorphisms compared to the reference sequence. [Figure 24] Figure 1 shows a schematic diagram of the mouse Dpp4 locus targeted for deletion and replacement with the corresponding human form in VGF1 hybrid ES cells using LTVEC and one or more gRNAs. The region within the vertical dotted line is the target region (regions within the 5' and 3' target sequences of LTVEC). The reference sequence for identifying single nucleotide polymorphisms was the genomic sequence of the C57BL / 6J mouse strain (Jackson Laboratory). This reference sequence was compared with the 129S6 / SvEv strain (Taconic Biosciences), the C57BL / 6N strain (Taconic Biosciences), and VGF1 hybrid cell lines generated from the 129S6 / SvEv and C57BL / 6N strains (listed in the bottom three rows of the figure). The vertical lines in each of the three rows indicate single nucleotide polymorphisms compared to the reference sequence. [Figure 25]Figure 1 shows a schematic diagram of the mouse Ror1 locus, which was targeted for deletion and replacement with the corresponding human form using LTVEC and one or more gRNAs in VGF1 hybrid ES cells. The region within the vertical dotted line is the target region (regions within the 5' and 3' target sequences of LTVEC). The reference sequence for identifying single nucleotide polymorphisms was the genomic sequence of the C57BL / 6J mouse strain (Jackson Laboratory). This reference sequence was compared with the 129S6 / SvEv strain (Taconic Biosciences), the C57BL / 6N strain (Taconic Biosciences), and VGF1 hybrid cell lines generated from the 129S6 / SvEv and C57BL / 6N strains (listed in the bottom three rows of the figure). The vertical lines in each of the three rows indicate single nucleotide polymorphisms compared to the reference sequence. [Figure 26] This figure shows a schematic of a mouse locus containing a gene encoding a transmembrane protein. The mouse locus is targeted for deletion and replacement with the corresponding humanized form using LTVEC and one or more gRNAs in VGF1 hybrid ES cells. The rectangles indicate different genes within the targeted genomic region. The area within the vertical dotted line is the target region (the area within the 5' and 3' target sequences of LTVEC). The reference sequence for identifying single nucleotide polymorphisms was the genomic sequence of the C57BL / 6J mouse strain (Jackson Laboratory). This reference sequence was compared with the 129S6 / SvEv strain MP mutant (Taconic Biosciences), the C57BL / 6N strain RGC mutant (Taconic Biosciences), and VGF1 hybrid cell lines generated from the 129S6 / SvEv and C57BL / 6N strains (listed in the bottom three rows of the figure). The MP mutant and RGC mutant are different mice from the same strain. The vertical lines in each of the three rows indicate single nucleotide polymorphisms compared to the reference sequence. [Figure 27]Figures 27A-27C are schematics illustrating mitotic recombination mechanisms that can occur during the G2 phase of the cell cycle, resulting in homozygous events and gene conversions detected by localized loss of heterozygosity. Figure 27A shows duplicated homologous chromosomes representing two chromatids in hybrid 129 / B6 ES cells heterozygous for targeted humanization on the 129 homolog. Heterozygous modification on the 129 homolog occurs before genome duplication, or a single 129 chromatid is modified after genome duplication, followed by interchromatid gene conversion. The double arrows indicate potential double-strand breaks generated by dual gRNA-guided Cas9 cleavage, which promotes dual-strand invasion and synthesis-guided repair. The diagonal dashed arrow indicates a hybrid chromatid generated by a gene conversion event that duplicates a small portion of one modified chromatid, as shown in Figure 27B. Figure 27C shows after mitosis and cell division, two types of chromosome segregation into daughter cells are possible: one that maintains heterozygosity without loss of heterozygosity (parental heterozygote (Hum / +, top)), and one that has localized loss of heterozygosity surrounding the targeted modification (Hum / Hum, bottom, retaining the 129 allele). [Figure 28] FIG. 1 shows the efficiency of CRISPR / Cas9-mediated deletion of different autoantigen targets of different sizes in VI-3 and ULC 1-39 embryonic stem (ES) cells using paired guide RNAs targeting the start and stop codon regions of autoantigen-encoding genes, either alone or together with large targeting vectors. [Figure 29] FIG. 1 shows the percentage of mouse pups generated with disrupted alleles after targeting of VI-3 and ULC 1-39 one-cell embryos with CRISPR / Cas9 to delete different autoantigen targets of different sizes using paired guide RNAs targeted to the start and stop codon regions of the autoantigen-encoding gene. [Figure 30]Figures 30A and 30B show antibody titer data for human target antigen (target 9) in wild-type VI-3-Adam6 mice (Figure 30B) and VI3-Adam6 mice that are homozygous null for the endogenous gene encoding the autoantigen orthologous to target 9 (autoantigen 9) (Figure 30A) after immunization with target 9 full-length DNA from parental VI-3T3 cells and VI-3T3 cells genetically modified to express target 9. [Figure 31] Figures 31A and 31B show antibody titer data for a human target antigen (target 4) and the corresponding orthologous murine autoantigen (autoantigen 4). Figure 31A shows antibody titer data for human target 4 and murine autoantigen 4 in VI3-Adam6 mice, which are homozygous null for the endogenous gene encoding autoantigen 4. Figure 31B shows antibody titer data for the combination of human target 4 and murine autoantigen 4 in ULC 1-39 mice, which are homozygous null for the endogenous gene encoding autoantigen 4. [Figure 32] Figure 1 shows schematic diagrams of the immunoglobulin heavy chain loci (top) and immunoglobulin light chain loci (bottom) in VI3-Adam6 and ULC 1-39 mice, which have 50% BALB / cTac, 25% C57BL / 6NTac, and 25% 129S6 / SvEvTac genetic backgrounds, respectively. In VI3-Adam6 mice, the endogenous mouse immunoglobulin heavy and light chain variable regions have been replaced with the corresponding human DNA in addition to the reinserted mouse Adam6 gene (Adam6b and Adam6a, shown as trapezoids). In universal light chain (ULC 1-39) mice, the endogenous mouse immunoglobulin heavy chain variable region has been replaced with the corresponding human DNA in addition to the reinserted mouse Adam6 gene, and the immunoglobulin light chain variable region contains a single rearranged human immunoglobulin light chain nucleotide sequence (Vκ1-39 / Jκ5) operably linked to the hVκ3-15 promoter. Human segments are shown in black, mouse segments are shown in diagonal lines. DETAILED DESCRIPTION OF THE INVENTION

[0076] definition The terms "protein," "polypeptide," and "peptide" are used interchangeably herein and include polymeric forms of amino acids of any length, including coded and non-coded amino acids, and amino acids that are chemically or biochemically modified or derivatized. The terms also include modified polymers, such as polypeptides having modified peptide backbones.

[0077] Proteins are said to have an "N-terminus" and a "C-terminus." The term "N-terminus" refers to the beginning of a protein or polypeptide, which ends with an amino acid having a free amine group (-NH2). The term "C-terminus" refers to the terminal end of an amino acid chain (protein or polypeptide), which ends with a free carboxyl group (-COOH).

[0078] The terms "nucleic acid" and "polynucleotide" are used interchangeably herein and include polymeric forms of nucleotides of any length, containing ribonucleotides, deoxyribonucleotides, or analogs or modifications thereof, including single-, double-, and multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, and polymers containing purine bases, pyrimidine bases, or other natural, chemically modified, biochemically modified, non-natural, or derivatized nucleotide bases.

[0079] Nucleic acids are said to have a "5' end" and a "3' end" because mononucleotides react to form oligonucleotides in such a way that the 5' phosphate of one mononucleotide pentose ring is unidirectionally linked to the 3' oxygen of an adjacent mononucleotide pentose ring via a phosphodiester bond. An end of an oligonucleotide is called a "5' end" if its 5' phosphate is not linked to the 3' oxygen of the mononucleotide pentose ring. An end of an oligonucleotide is called a "3' end" if its 3' oxygen is not linked to the 5' phosphate of another mononucleotide pentose ring. Even within a larger oligonucleotide, a nucleic acid sequence may be said to have a 5' end and a 3' end. In either a linear or circular DNA molecule, individual elements are referred to as "upstream" elements, 5' of "downstream" elements, or 3' elements.

[0080] The term "wild-type" includes components having a structure and / or activity found in a normal (as opposed to mutated, abnormal, altered, etc.) state or condition. Wild-type genes and polypeptides often exist in multiple alternative forms (e.g., alleles).

[0081] The term "isolated," with respect to proteins and nucleic acids, includes proteins and nucleic acids that are relatively purified relative to other bacterial, viral, or cellular components that may normally be present in situ, and up to and including substantially pure protein and polynucleotide preparations. The term "isolated" also includes proteins and nucleic acids that are chemically synthesized free of natural counterparts and are therefore substantially free of contaminating other proteins or nucleic acids, or proteins and nucleic acids that have been separated or purified from most of the other cellular components (e.g., other cellular proteins, polynucleotides, or cellular components) with which they naturally occur.

[0082] A "foreign" molecule or sequence includes a molecule or sequence that is not normally present in a cell in that form. Normal presence includes presence in a cell at a particular developmental stage and environmental conditions. A foreign molecule or sequence may include, for example, a mutant version of a corresponding endogenous sequence in the cell (e.g., a humanized version of an endogenous sequence), or may include a sequence that corresponds to an endogenous sequence in the cell but in a different form (i.e., not in a chromosome). In contrast, an endogenous molecule or sequence includes a molecule or sequence that is normally present in that form in a particular cell at a particular developmental stage under particular environmental conditions.

[0083] "Codon optimization" generally refers to the process of modifying a nucleic acid sequence to enhance expression, particularly in a host cell, by replacing at least one codon in the native sequence with a codon that is more frequently or most frequently utilized in the host cell's genes while maintaining the native amino acid sequence. For example, a polynucleotide encoding a Cas9 protein may be modified to replace a codon that is more frequently utilized in any prokaryotic or eukaryotic cell (including bacterial cells, yeast, human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, hamster cells, or any other host cell) compared to the native nucleic acid sequence. Codon usage tables are readily available, for example, in the "Codon Usage Database." These tables may be applied in a variety of ways. See Nakamura et al. (2000) Nucleic Acids Research 28:292 (incorporated herein by reference in its entirety for all purposes). Computer algorithms are also available for performing codon optimization in a particular sequence for expression in a particular host (see, for example, Gene Forge).

[0084] The term "locus" refers to a specific location in a gene (or key sequence), DNA sequence, polypeptide-coding sequence, or chromosomal location in the genome of an organism. For example, "Lrp5 locus" can refer to a specific location in the Lrp5 gene, Lrp5 DNA sequence, LRP5 coding sequence, or the location of Lrp5 on a chromosome in the genome of an organism in which such a sequence has been identified. The "Lrp5 locus" may include regulatory elements of the Lrp5 gene, including, for example, an enhancer, promoter, 5'UTR and / or 3'UTR, and combinations thereof.

[0085] The term "gene" refers to a DNA sequence within a chromosome that encodes a product (e.g., an RNA product and / or a polypeptide product) and includes a coding region interposed by non-coding introns, and sequences located adjacent to the coding region on both the 5' and 3' ends such that the gene represents the full-length mRNA (including 5' and 3' untranslated sequences). The term "gene" also includes other non-coding sequences, including regulatory sequences (e.g., promoters, enhancers, and transcription factor binding sites), polyadenylation signals, internal ribosome entry sites, silencers, insulator sequences, and matrix attachment regions. These sequences may be located adjacent to (e.g., within 10 kb) or distant from) the coding region of a gene and influence the level or rate of transcription and translation of the gene.

[0086] The term "allele" refers to a variant form of a gene. Some genes have several different forms that are located at the same position, or locus, on a chromosome. Diploid organisms have two alleles at each locus. Each pair of alleles represents a genotype at a particular locus. A genotype is described as homozygous when two identical alleles are present at a particular locus and heterozygous when the two alleles are different.

[0087] A "promoter" is a regulatory region of DNA, typically containing a TATA box that can direct RNA polymerase II to initiate RNA synthesis at the appropriate transcription start site of a particular polynucleotide sequence. A promoter may additionally contain other regions that affect the rate of transcription initiation. The promoter sequences disclosed herein regulate the transcription of an operably linked polynucleotide. The promoter may be active in one or more of the cell types disclosed herein (e.g., eukaryotic cells, non-human mammalian cells, human cells, rodent cells, pluripotent cells, one-cell stage embryos, differentiated cells, or combinations thereof). The promoter may be, for example, a constitutively active promoter, a conditional promoter, an inducible promoter, a temporally restricted promoter (e.g., a developmentally regulated promoter), or a spatially restricted promoter (e.g., a cell-specific or tissue-specific promoter). Examples of promoters can be found, for example, in WO2013 / 176772 (incorporated herein by reference in its entirety).

[0088] Examples of inducible promoters include chemically regulated promoters and physically regulated promoters. Chemically regulated promoters include alcohol-regulated promoters (e.g., alcohol dehydrogenase (alcA) gene promoter), tetracycline-regulated promoters (e.g., tetracycline-responsive promoters, tetracycline operator sequence (tetO), tet-On promoters, or tet-Off promoters), steroid-regulated promoters (e.g., rat glucocorticoid receptor, estrogen receptor promoter, or ecdysone receptor promoter), or metal-regulated promoters (e.g., metalloprotein promoters). Physically regulated promoters include temperature-regulated promoters (e.g., heat shock promoters) and light-regulated promoters (e.g., light-inducible promoters or light-repressible promoters).

[0089] The tissue-specific promoter may be, for example, a neuron-specific promoter, a glial-specific promoter, a muscle cell-specific promoter, a cardiac cell-specific promoter, a kidney cell-specific promoter, a bone cell-specific promoter, an endothelial cell-specific promoter, or an immune cell-specific promoter (e.g., a B cell promoter or a T cell promoter).

[0090] Developmentally-regulated promoters include, for example, promoters that are active only during embryonic development or promoters that are active only in mature cells.

[0091] "Operably linked" or "operably linked" includes juxtaposition of two or more components (e.g., a promoter and another sequence element) such that both components function normally to enable at least one of the components to mediate the function of at least one of the other components. For example, a promoter may be operably linked to a coding sequence if the promoter regulates the level of transcription of the coding sequence depending on the presence or absence of one or more transcriptional regulatory factors. Operably linked sequences may include sequences that are contiguous with each other or that act in trans (e.g., regulatory sequences may act at a distance to regulate transcription of the coding sequence). In another example, a nucleic acid sequence of an immunoglobulin variable region (or V(D)J segment) may be operably linked to a nucleic acid sequence of an immunoglobulin constant region for proper inter-sequence recombination into an immunoglobulin heavy or light chain sequence.

[0092] "Complementarity" of nucleic acids means that the nucleotide sequence of one strand of a nucleic acid forms hydrogen bonds with another sequence on an opposing nucleic acid strand due to the orientation of its nucleobase groups. Complementary bases in DNA are typically A and T and C and G. Complementary bases in RNA are typically C and G and U and A. Complementarity may be complete or substantial / sufficient. Complete complementarity between two nucleic acids means that the two nucleic acids can form a duplex in which all bases in the duplex are bonded to complementary bases by Watson-Crick pairing. "Substantially" or "sufficiently" complementary means that the sequence of one strand is not completely and / or perfectly complementary to the sequence of the opposing strand under defined hybridization conditions (e.g., salt concentration and temperature), but sufficient binding occurs between the bases on the two strands to form a stable hybrid complex. Such conditions can be predicted by using standard mathematical calculations to predict the sequence and Tm (melting temperature) of hybridized strands, or by experimentally determining the Tm using conventional methods. The Tm comprises the temperature at which a population of hybridization complexes formed between two nucleic acid strands becomes 50% denatured (i.e., the population of double-stranded nucleic acid molecules becomes half-dissociated into single strands). Preferably, hybridization complexes form at temperatures below the Tm, whereas preferably, the strands of the hybridization complex melt or separate at temperatures above the Tm. While taking into account other known Tm calculations for nucleic acid structural properties, for example, Tm = 81.5 + 0.41 (% G + C) may be used to estimate the Tm of a nucleic acid with a known G + C content in 1 M aqueous NaCl solution.

[0093] "Hybridization conditions" include the cumulative environment in which one nucleic acid strand binds to a second nucleic acid strand through complementary interactions and hydrogen bonds to form a hybridization complex. Such conditions include the chemical components and concentrations (e.g., salts, chelating agents, formamide) of the aqueous or organic solution containing the nucleic acid, as well as the temperature of the mixture. Other factors, such as the length of incubation time or reaction chamber dimensions, can contribute to the environment. See, e.g., Sambrook et al., Molecular Cloning, A Laboratory Manual, 2nd ed., pp. 1.90-1.91, 9.47-9.51, 11.47-11.57 (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989), the entire contents of which are incorporated herein by reference for all purposes.

[0094] Hybridization requires two nucleic acids containing complementary sequences, although mismatches between bases may occur. Suitable conditions for hybridization between two nucleic acids depend on the length of the nucleic acids and the degree of complementarity, variables well known in the art. The higher the degree of complementarity between two nucleotide sequences, the higher the melting temperature (Tm) of hybrids of nucleic acids having those sequences. For hybridization between nucleic acids with short stretches of complementarity (e.g., 35 or less, 30 or less, 25 or less, 22 or less, 20 or less, or more than 18 or less nucleotides), the position of mismatches becomes important (see Sambrook et al., supra, 11.7-11.8). Typically, the length of a hybridizable nucleic acid is at least about 10 nucleotides. Exemplary minimum lengths for a hybridizable nucleic acid include at least about 15 nucleotides, at least about 20 nucleotides, at least about 22 nucleotides, at least about 25 nucleotides, and at least about 30 nucleotides. Furthermore, the temperature and salt concentration of the wash solution may be adjusted as needed depending on factors such as the length of the complementary region and the degree of complementarity.

[0095] To be specifically hybridizable, the sequence of a polynucleotide need not be 100% complementary to the sequence of its target nucleic acid. Furthermore, polynucleotides may hybridize across one or more segments (e.g., loop or hairpin structures) such that the hybridization event does not involve intervening or adjacent segments. Polynucleotides (e.g., gRNAs) may contain at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence complementarity to their targets within the target nucleic acid sequence to which they are targeted. For example, a gRNA that specifically hybridizes because 18 of 20 nucleotides are complementary to its target would exhibit 90% complementarity. In this example, the remaining non-complementary nucleotides may be clustered or interspersed among complementary nucleotides and need not be contiguous to each other or to complementary nucleotides.

[0096] The percent complementarity between individual stretches of nucleic acid sequences within a nucleic acid can be determined using the BLAST program (basic local alignment search tool) and PowerBLAST program (Altschul et al., 2001) well known in the art. al. (1990) J. Mol. Biol. 215:403-410; Zhang and Madden (1997) Genome Res. 7:649-656), or by using the Gap program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, Madison Wis.) with default settings using the algorithm of Smith and Waterman (Adv. Appl. Math., 1981, 2, 482-489).

[0097] The methods and compositions provided herein utilize a variety of different components. Throughout this specification, it is understood that some components may have active variants and active fragments. Such components include, for example, Cas9 protein, CRISPR RNA, tracrRNA, and guide RNA. The biological activities of each of these components are described elsewhere herein.

[0098] "Sequence identity" or "identity," in the context of two polynucleotide or polypeptide sequences, refers to residues in the two sequences that are identical when aligned for maximum correspondence over a specified comparison window. When percentage sequence identity is used to refer to proteins, it should be understood that non-identical residue positions often differ by conservative amino acid substitutions, in which an amino acid residue is replaced with another amino acid residue having similar chemical properties (e.g., charge or hydrophobicity), thereby resulting in no change in the functional properties of the molecule. When sequences differ by conservative substitutions, the percent sequence identity may be adjusted upward to correct for the conservatism of the substitution. Sequences that differ by such conservative substitutions are said to have "sequence similarity" or "similarity." Methods for making this correction are well known to those skilled in the art. Typically, this method involves increasing the percent sequence identity by scoring conservative substitutions as partial rather than complete mismatches. Thus, for example, conservative substitutions are given a score between zero and one, where identical amino acids are given a score of one and non-conservative substitutions are given a score of zero. For example, conservative substitution scores are calculated by running the program PC / GENE (Intelligenetics, Mountain View, California).

[0099] "Percentage sequence identity" includes values ​​calculated by comparing two optimally aligned sequences over a comparison window, where the portion of the polynucleotide sequence within the comparison window may contain additions or deletions (i.e., gaps) compared to a reference sequence (which does not contain additions or deletions) for optimally aligning the two sequences. The percentage is calculated by calculating the number of positions where the same nucleic acid base or amino acid residue occurs in both sequences to obtain the number of matched positions, dividing the number of matched positions by the total number of positions within the comparison window, and multiplying the result by 100 to obtain the percentage sequence identity.

[0100] Unless otherwise specified, sequence identity / similarity values ​​include values ​​obtained using GAP Version 10 or any equivalent program with the following parameters: % identity and % similarity for nucleotide sequences using a GAP Weight of 50, a Length Weight of 3, and the nwsgapdna.cmp scoring matrix; % identity and % similarity for amino acid sequences using a GAP Weight of 8, a Length Weight of 2, and the BLOSUM62 scoring matrix. "Equivalent program" refers to the program that runs GAP Version 10 or any equivalent program with the following parameters for any two sequences in question: Any sequence comparison program that produces alignments that have identical nucleotide or amino acid residue matches and identical percent sequence identity when compared to the corresponding alignment produced in 10.

[0101] As used herein, the term "substantial identity" refers to shared epitopes comprising sequences containing identical residues at corresponding positions. For example, two sequences can be considered substantially identical if at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more of their corresponding residues are identical over the relevant stretch. The relevant stretch can be, for example, the complete sequence, or at least 5, 10, 15, or more residues.

[0102] The term "conservative amino acid substitution" refers to the substitution of an amino acid normally occurring in a sequence with a different amino acid of similar size, charge, or polarity. Examples of conservative substitutions include the substitution of a non-polar (hydrophobic) residue, such as isoleucine, valine, or leucine, for another non-polar residue. Similarly, examples of conservative substitutions include the substitution of one polar (hydrophilic) residue for another, such as between arginine and lysine, between glutamine and asparagine, or between glycine and serine. Furthermore, the substitution of one basic residue (such as lysine, arginine, or histidine) for another, or the substitution of one acidic residue (such as aspartic acid or glutamic acid) for another, are additional examples of conservative substitutions. Examples of non-conservative substitutions include the substitution of a non-polar (hydrophobic) amino acid residue, such as isoleucine, valine, leucine, alanine, or methionine, for a polar (hydrophilic) residue, such as cysteine, glutamine, glutamic acid, or lysine, and / or the substitution of a polar residue for a non-polar residue. Representative amino acid classes are summarized below. [Table A]

[0103] The term "germline" with respect to immunoglobulin nucleic acid sequences includes nucleic acid sequences that are passed on to progeny.

[0104] The term "antigen-binding protein" includes any protein that binds to an antigen. Examples of antigen-binding proteins include antibodies, antigen-binding fragments of antibodies, multispecific antibodies (e.g., bispecific antibodies), scFvs, bis-scFVs, diabodies, triabodies, tetrabodies, V-NARs, VHHs, VLs, F(ab)s, F(ab)2s, DVDs (dual variable domain antigen-binding proteins), SVDs (single variable domain antigen-binding proteins), bispecific T-cell engagers (BiTEs), or Davisbodies (U.S. Patent No. 8,586,713, the entire contents of which are incorporated herein by reference for all purposes).

[0105] The term "antigen" refers to a substance, whether an entire molecule or a domain within a molecule, that is capable of inducing the production of antibodies with binding specificity to that substance. The term "antigen" also includes substances that do not elicit antibody production through self-recognition in a wild-type host organism, but that can induce such a response in a host animal that has been appropriately genetically engineered to disrupt immune tolerance.

[0106] The term "epitope" refers to a site on an antigen to which an antigen-binding protein (e.g., an antibody) binds. Epitopes can be formed from contiguous amino acids or noncontiguous amino acids arranged by tertiary folding of one or more proteins. Epitopes formed from contiguous amino acids (also known as linear epitopes) typically retain their structure upon exposure to denaturing solvents, while epitopes formed from tertiary folding (also known as conformational epitopes) typically are lost upon treatment with denaturing solvents. Epitopes typically comprise at least three, more commonly at least five, or 8-10 amino acids in a unique conformation. Methods for identifying the conformation of an epitope include, for example, X-ray crystallography and two-dimensional nuclear magnetic resonance. See, for example, "Epitope Mapping Protocols," in "Methods in Molecular Biology," Vol. 66, Glenn E. Morris, Ed. (1996), incorporated herein by reference in its entirety for all purposes.

[0107] The term "self" when used in conjunction with an antigen or epitope refers to an antigen or epitope that would not be recognized or poorly recognized by the B cell receptors of wild-type members of the host species because it is contained in a substance normally biosynthesized by the host species or to which the host species is normally exposed. Such a substance induces tolerance in the host immune system. The term "foreign" when used in conjunction with an antigen or epitope refers to an antigen or epitope that is not a self-antigen or self-epitope. A foreign antigen is any antigen not normally produced by the host species.

[0108] The term "antibody" includes an immunoglobulin molecule comprising four polypeptide chains, two heavy (H) chains and two light (L) chains linked together by disulfide bonds. Each heavy chain contains a heavy chain variable domain and a heavy chain constant region (C H The heavy chain constant region contains three domains: C H 1. C H 2 and C H3. Each light chain comprises a light chain variable domain and a light chain constant region (C L ). The heavy and light chain variable domains can be further subdivided into regions of hypervariability called complementarity determining regions (CDRs) interspersed with more conserved regions called framework regions (FRs). Each heavy and light chain variable domain contains three CDRs and four FRs, arranged from amino-terminus to carboxy-terminus in the following order: FR1, CDR1, FR2, CDR2, FR3, CDR3, FR4 (heavy chain CDRs may be abbreviated as HCDR1, HCDR2, and HCDR3, and light chain CDRs may be abbreviated as LCDR1, LCDR2, and LCDR3). The term "high affinity" antibody refers to an antibody that binds to its target epitope with a high affinity of about 10 -9 M or less (e.g., about 1 × 10 -9 M, 1 x 10 -10 M, 1 x 10 -11 M, or approximately 1 x 10 -12 M)'s K D In one embodiment, the antibody has the formula: D is measured by surface plasmon resonance, e.g., BIACORE™, and in another embodiment, K D is measured by ELISA.

[0109] The term "heavy chain" or "immunoglobulin heavy chain" includes immunoglobulin heavy chain sequences, including immunoglobulin heavy chain constant region sequences, from any organism. A heavy chain variable domain, unless otherwise specified, includes three heavy chain CDR regions and four FR regions. Fragments of heavy chains include CDRs, CDRs and FRs, and combinations thereof. A typical heavy chain contains the following variable domains (in order from N-terminus to C-terminus): C H 1 domain, hinge, C H 2 domain, and C H A functional fragment of a heavy chain has three domains capable of specifically recognizing an epitope (e.g., K in the micromolar, nanomolar, or picomolar range). DThe heavy chain variable domain is encoded by a variable region nucleotide sequence, and typically includes a V domain present in the germline. H , D H and J H V derived from segmental repertoire H , D H and J H The sequences, locations and nomenclature of V, D and J heavy chain segments in various organisms can be found in the IMGT database, accessible via the internet on the world wide web (www) at the URL "imgt.org."

[0110] The term "light chain" includes immunoglobulin light chain sequences from any organism, and includes human kappa (κ) and lambda (λ) light chains and VpreB, as well as surrogate light chains, unless otherwise specified. A light chain variable domain typically includes three light chain CDRs and four framework (FR) regions, unless otherwise specified. A full-length light chain typically includes a variable domain comprising, from the amino terminus to the carboxyl terminus, FR1-CDR1-FR2-CDR2-FR3-CDR3-FR4, and a light chain constant region amino acid sequence. The light chain variable domain is encoded by a light chain variable region nucleotide sequence, and typically includes a light chain V derived from a repertoire of light chain V and J gene segments present in the germline. L and light chain J LLight chains include gene segments. The sequences, locations, and nomenclature of light chain V and J gene segments in various organisms can be found in the IMGT database, accessible via the internet at the world wide web (www) URL "imgt.org." Light chains include, for example, those that do not selectively bind to either the first or second epitope to which the epitope-binding protein selectively binds, when the first or second epitope is present. Light chains also include those that, when the epitope is present, bind to and recognize one or more epitopes to which the epitope-binding protein selectively binds, or assist the heavy chain in binding and recognizing one or more epitopes to which the epitope-binding protein selectively binds.

[0111] As used herein, the term "complementarity-determining region" or "CDR" includes amino acid sequences encoded by the nucleic acid sequence of an organism's immunoglobulin genes that are normally (i.e., in a wild-type animal) located between two framework regions within the variable region of the light or heavy chain of an immunoglobulin molecule (e.g., an antibody or T-cell receptor). CDRs may be encoded, for example, by germline or rearranged sequences, and may be encoded, for example, by naive or mature B or T cells. CDRs may be somatically mutated (e.g., different from the sequence encoded by the animal's germline), humanized, and / or modified by amino acid substitution, addition, or deletion. In some situations (e.g., CDR3), CDRs may be encoded by two or more sequences (e.g., germline sequences) that are not contiguous (e.g., within the unrearranged nucleic acid sequence) but are contiguous within the B-cell nucleic acid sequence, for example, as a result of recombination of sequences (e.g., VDJ) by splicing or joining to form the heavy chain CDR3.

[0112] The term "unrearranged" includes a state of an immunoglobulin locus in which the V and J gene segments (and in the heavy chain, the D gene segments) are maintained separately but can combine to form rearranged V(D)J genes comprising a single V, (D), and J of the V(D)J repertoire.

[0113] The term "heavy chain variable region locus" refers to a wild-type heavy chain variable (V H ), heavy chain diversity (D H ), and heavy chain binding (J H ) region. This includes the location on a chromosome (e.g., a mouse chromosome) where a DNA sequence has been found.

[0114] The term "kappa light chain variable region locus" includes the location on a chromosome (eg, a mouse chromosome) where wild-type kappa variable (Vκ) and kappa joining (Jκ) region DNA sequences are found.

[0115] The term "lambda light chain variable region locus" includes the location on a chromosome (eg, a mouse chromosome) at which wild-type lambda variable (Vλ) and lambda joining (Jλ) region DNA sequences are found.

[0116] A "homologous" sequence (e.g., a nucleic acid sequence) includes a sequence that is either identical or substantially similar to a known reference sequence, e.g., at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the known reference sequence. Homologous sequences can include, for example, orthologous and paralogous sequences. For example, homologous genes are typically descended from a common ancestral DNA sequence by either speciation events (orthologous genes) or gene duplication events (paralogous genes). "Orthologous" genes include genes from different species that have evolved from a common ancestral gene by speciation. Orthologs typically maintain the same function during evolution. "Paralogous" genes include genes associated with intragenomic duplication. Paralogs can develop new functions during evolution.

[0117] The term "in vitro" includes an artificial environment and processes or reactions that occur within an artificial environment (e.g., a test tube). The term "in vivo" includes a natural environment (e.g., a cell, organism, or living organism) and processes or reactions that occur within a natural environment. The term "ex vivo" includes cells removed from an individual's organism and processes or reactions that occur within such cells.

[0118] The term "hybrid" includes cells or strains that have one or more sequence variations (e.g., allelic variations) at one or more target genomic loci between a first and a second chromosome within a homologous chromosome pair. For example, a hybrid cell may be derived from the progeny of a mating between two genetically distinct parents (i.e., parents that differ in one or more genes). As an example, a hybrid may be created by mating two different inbred strains (i.e., genetically homogeneous breeding strains). All humans are considered hybrids.

[0119] A composition or method that "comprising" or "including" one or more recited elements may include other elements not expressly recited. For example, a composition that "comprises" or "includes" a protein may include the protein alone or in combination with other components.

[0120] Any reference to a range of values ​​includes every integer within or defining that range, and every sub-range defined by the integers within that range.

[0121] Unless otherwise clear from the context, the term "about" includes values ​​within the standard error of measurement (eg, SEM) of the stated value.

[0122] The singular articles "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. For example, the term "a Cas9 protein" or "at least one Cas9 protein" can include multiple Cas9 proteins, including combinations thereof.

[0123] Statistically significant means p≦0.05.

[0124] I. Overview Provided herein are compositions and improved methods for generating antigen-binding proteins (e.g., antibodies) that bind to epitopes on a foreign target antigen of interest (e.g., a human target antigen of interest) that share an epitope with or are homologous to a self-antigen. Such methods include suppressing tolerance to a foreign antigen in a non-human animal, such as a rodent (e.g., a mouse or rat), (optionally comprising a humanized immunoglobulin heavy and / or light chain locus in its germline) by using two or more guide RNAs (gRNAs) to generate paired double-stranded breaks at different sites within a single target genomic locus. Optionally, the cell comprising the target genomic locus is a hybrid cell, and the method further comprises selecting a target region within the target genomic locus and performing the targeted gene modification such that the target region has a higher degree of sequence identity between the corresponding first and second chromosomes of a homologous chromosome pair compared to all or a portion of the remainder of the target genomic locus. Such paired double-strand breaks affect the expression of the self-antigen, reducing or eliminating expression of the self-antigen or reducing or eliminating expression of an epitope from the self-antigen that is shared with the foreign antigen. Thus, such a genetically modified non-human animal comprising humanized immunoglobulin heavy and light chain loci and harboring such a mutation within the targeted genomic locus may be immunized with the foreign antigen, the non-human animal may be maintained under conditions sufficient for the non-human animal to generate an immune response against the foreign antigen, and an antigen-binding protein that binds to the foreign antigen may be obtained from the non-human animal or a cell derived from the non-human animal.

[0125] Mice used to generate antibodies against human antigens, such as those containing humanized immunoglobulin heavy and / or light chain loci in their germline, are typically derived from a combination of strains including BALB / c, due to the BALB / c strain's increased ability to produce a diverse repertoire of antibodies compared to other mouse strains. However, compared to embryonic stem (ES) cells typically used to generate targeted gene modifications in mice (e.g., F1H4(VGF1) cells described herein), ES cells derived from antibody-producing mice of this strain typically have a reduced ability to be targeted in culture and / or to generate F0 generation mice that carry the targeted gene modification and inherit the targeted modification through their germline. As a result, traditional methods for generating targeted knockout mice and overcoming tolerance involve multiple rounds of breeding and / or serial targeting, and the entire process takes approximately 15–16 months to produce mice homozygous for null alleles at the target of interest and ready for immunization.

[0126] The methods described herein advantageously shorten this time period to approximately 4-5 months (pups homozygous for a null allele in the target of interest can be born in approximately 3 months). In addition to the shorter time period, the methods described herein streamline the process by reducing the number of rounds of electroporation required to generate homozygous modifications, reducing the number of passages required in culture, shortening the time required for culture, reducing the number of cells required, and simplifying screening by eliminating the need for targeting vectors. The methods described herein advantageously result in increased antibody diversity after immunization with a foreign antigen of interest due to increased use of heavy and light chain V gene segments compared to mice in which expression of the self-antigen is not silenced. Additionally, the methods described herein produce antibodies that cross-react with the corresponding self-antigen (i.e., antibodies that bind to epitopes that overlap between the self-antigen and the foreign antigen of interest), resulting in a greater diversity of antibodies produced against epitopes after immunization with the foreign antigen of interest, thereby enabling the generation of a larger pool of antibodies against the foreign antigen of interest.

[0127] II. Methods for modifying targeted genomic loci to disrupt tolerance Immunization of non-human animals (e.g., rodents such as mice or rats) containing humanized immunoglobulin heavy and / or light chain loci in their germline with "non-self" proteins is a commonly used method for obtaining specific antigen-binding proteins, such as monoclonal antibodies. This immunization approach is attractive because it has the potential to yield high-affinity antigen-binding proteins that are matured in vivo and can be both cost-effective and time-effective. However, this approach relies on sequence differences between the native protein of the non-human animal and the immunized protein such that the non-human animal's immune system can recognize the immunogen as non-self (i.e., foreign).

[0128] B cell receptors are assembled from ordered gene segments (e.g., V, D, and J) through a series of recombination events, but this assembly of gene segments is known to be imprecise, producing receptors with affinities for a variety of antigens, including self-antigens. Despite this ability to produce B cell receptors that bind to self-molecules, the immune system has several self-tolerance mechanisms to prevent the development and proliferation of such autoreactive B cell receptors and to avoid autoimmunity by distinguishing self from non-self. See, e.g., Shlomchik (2008) Immunity 28:18-28, and Kumar See, e.g., Wang and Mohan (2008) 40(3):208-23, the entire contents of each of which are incorporated herein by reference for all purposes. Therefore, the production of human antibodies in non-human animals with humanized immunoglobulin loci against human antigens that share a high degree of homology (e.g., structural or sequence homology) with the non-human animal's self-antigens can be a challenging task due to immune tolerance. Because functionally important regions of proteins tend to be conserved across species, immune tolerance to self-antigens often poses a problem in the production of antibodies against these important epitopes. Immunizing non-human animals (e.g., rodents such as mice or rats) with highly similar or "homologous" foreign (e.g., human) antigens results in weak or no antibody responses, thereby making it uncertain to obtain antigen-binding proteins (e.g., antibodies) with binding affinity toward such human antigens. For example, the amount of sequence identity shared between an endogenous protein (self-antigen) and a foreign target antigen may be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity, such that the immune system does not recognize the target antigen as foreign. For example, shared epitopes between a foreign antigen and a self-antigen in a non-human animal may preclude the formation of an effective immune response to the foreign antigen in the non-human animal, because immune tolerance reduces and / or eliminates B cells that express neutralizing antibodies to the foreign antigen. To overcome this tolerance and obtain monoclonal antibodies that bind to self-antigens or their homologs (e.g., human homologs) in non-human animals, specific genetically modified or knockout non-human animals may be engineered to remove genes encoding non-human animal proteins that share significant homology (or shared epitopes of interest) and / or genes that are highly conserved in their human counterparts that encode the antigens used for immunization (or shared epitopes of interest).See, for example, U.S. Patent No. 7,119,248, the entire contents of which are incorporated by reference for all purposes. However, the generation of such non-human animals can be both expensive and time-consuming.

[0129] Conventional methods for generating targeted knockout mice and overcoming tolerance include multiple rounds of breeding and / or serial targeting. Mice used to generate antibodies against human antigens, such as mice containing humanized immunoglobulin heavy and / or light chain loci in their germline (e.g., VELOCIMMUNE® mice homozygously humanized at both the IgH and Igκ loci), are typically derived from a combination of strains including BALB / c, due to the BALB / c strain's greater ability to produce a diverse repertoire of antibodies compared to other mouse strains. However, compared to embryonic stem (ES) cells typically used to generate targeted gene modifications in mice (e.g., F1H4 (VGF1) cells described herein, which are composed of 50% 129SvS6 and 50% C57BL / 6N strains), ES cells derived from antibody-producing mice of such strains typically have reduced ability to be targeted in culture and / or to generate F0 generation mice that carry targeted gene modifications and inherit the targeted modifications through their germline. Therefore, a conventional approach to breaking immune tolerance in antibody-producing mice, such as VELOCIMMUNE® mice, involves first targeting a gene encoding an autoantigen in a susceptible ES cell line (e.g., F1H4) and then inheriting the targeted modification through the germline. In this approach, a large targeting vector (LTVEC) is designed to create a knockout (null) allele in F1H4 ES cells, generating F0 mice with a heterozygous knockout mutation in the target of interest (typically over a 5-month period). VELOCIMMUNE® mice are then bred into F0 mice with a heterozygous knockout mutation in the target of interest. Two additional generations of breeding are required to generate triple homozygous mice suitable for immunization (homozygous null for the target of interest and homozygously humanized for both IgH and Igκ). The entire process takes approximately 15-16 months (see, e.g., Figure 1) and is more effective compared to the serial targeting approach described below (see, e.g., Figure 2).

[0130] Alternatively, large targeting vectors (LTVECs) can be designed and constructed and then electroporated into embryonic stem (ES) cells derived from antibody-producing mice (e.g., VELOCIMMUNE® mice, or VELOCIMMUNE® mice containing a functional ectopic mouse Adam6 gene ("VI-3 mice")) to generate heterozygous modifications in endogenous genes encoding autoantigens that are homologous to the target antigen or share an epitope of interest with the target antigen. A second round of targeting is then performed to generate homozygous modifications. While less time-consuming than the breeding methods described above, this process can still be time-consuming, taking approximately 9-10 months to generate F0 mice ready for immunization with the target antigen (see, e.g., Figure 2). In addition, such methods require multiple rounds of electroporation and longer culture periods with more passages, all of which result in reduced pluripotency and a reduced ability to generate F0 mice for antigen-binding protein production. See, e.g., Buehr et al. (2008) Cell 135:1287-1298; Li et al. (2008) Cell 135(7):1299-1310; and Liu et al. (1997) Dev. Dyn. 209:85-91, the entire contents of each of which are incorporated herein by reference for all purposes.

[0131] The methods described herein advantageously shorten this time period to approximately 4-5 months (see, e.g., Figure 3 (pups homozygous for a null allele in the target of interest can be born in approximately 3 months, but then grown for 4-5 weeks before immunization)). In addition to the shorter time period, the methods described herein reduce the number of rounds of electroporation required to generate homozygous modifications, reduce the number of passages required in culture, shorten the time required in culture, and reduce the number of cells required. Screening is simplified and streamlined, for example, by eliminating the need for allelic probe recovery and copy number calibration. The methods described herein also result in increased antibody diversity after immunization with a foreign antigen of interest due to increased usage of heavy and light chain V gene segments compared to mice in which self-antigen expression is not silenced. Additionally, the methods described herein can result in antibodies being produced against a greater variety of epitopes after immunization with a foreign antigen of interest by producing antibodies that cross-react with the corresponding self-antigen (i.e., antibodies that bind to epitopes that overlap between the self-antigen and the foreign antigen of interest), thereby allowing for the production of a larger pool of antibodies against the foreign antigen of interest.

[0132] Various methods for modifying a target genomic locus to disrupt tolerance are provided herein. These methods can be performed ex vivo or in vivo, and utilize two or more guide RNAs (e.g., two gRNAs, three guide RNAs, or four guide RNAs) targeting different regions within a single target genomic locus that affect the expression of an autoantigen that is homologous to or shares an epitope of interest with the foreign antigen of interest, forming two or more complexes with Cas proteins to cleave the target nucleic acid. When the cell is a one-cell embryo, two or more guide RNAs may be used alone or in combination with an exogenous repair template, for example, the exogenous repair template may be less than 5 kb in length. These methods facilitate the generation of biallelic genetic modifications at the target locus and may include genome disruption or other targeted modifications, such as simultaneous deletion of a genomic nucleic acid sequence and its replacement with an exogenous nucleic acid sequence. Compared to targeting with one gRNA (which generates biallelic modifications at low frequencies), targeting with two or more gRNAs results in a significantly higher rate of biallelic modifications (e.g., compound heterozygously targeted cells, including homozygously targeted cells, homozygously deleted cells, and hemizygously targeted cells).

[0133] Repair of double-strand breaks (DSBs) occurs primarily through two conserved DNA repair pathways: non-homologous end joining (NHEJ) and homologous recombination (HR). See, & Humphrey (2011) Seminars in Cell & Dev. Biol. 22:886-897 (the entire contents of which are incorporated herein by reference for all purposes). NHEJ involves the repair of double-strand breaks in nucleic acids by direct ligation of the broken ends to each other or to a foreign sequence without the need for a homologous template. Ligation of non-contiguous sequences by NHEJ can often result in deletions, insertions, or translocations near the double-strand break site.

[0134] Repair of target nucleic acids involving exogenous repair templates may include any genetic information exchange process between two polynucleotides. For example, NHEJ can also result in the targeted introduction of exogenous repair templates by directly ligating the cleaved ends to the ends of the exogenous repair template (i.e., NHEJ-based complementation). Such NHEJ-mediated targeted introduction may be preferable for the insertion of exogenous repair templates when the homology-directed repair (HDR) pathway is not readily available (e.g., in non-dividing cells, primary cells, and cells in which homology-based DNA repair is poorly performed). In addition, in contrast to homology-directed repair, information about the large region of sequence identity (beyond the overhang generated by Cas-mediated cleavage) around the cleavage site is not required (which may be useful when attempting targeted insertion into an organism with a genome with limited genomic sequence information). Introduction can proceed by blunt-end ligation between the exogenous repair template and the cleaved genomic sequence, using an exogenous repair template flanked by overhangs compatible with the overhangs in the cleaved genomic sequence generated by the Cas protein, or by cohesive-end ligation (i.e., with 5' or 3' overhangs). See, e.g., US2011 / 020722, WO2014 / 033644, WO2014 / 089290, and Maresca et al. (2013) Genome Res. 23(3):539-546 (the entire contents of each are incorporated by reference herein for all purposes). When blunt-end ligation is used, cleavage of the target and / or donor may be required to generate the microhomology regions necessary for fragment joining, which may result in unwanted modifications in the target sequence.

[0135] Repair can also be performed by homology-directed repair (HDR) or homologous recombination (HR). HDR or HR involves a form of nucleic acid repair that may require nucleotide sequence homology, using a "donor" molecule as a template to repair a "target" molecule (i.e., a molecule that has undergone a double-strand break), resulting in the transfer of genetic information from the donor to the target. Without being limited to any particular theory, such transfer may involve mismatch repair of heteroduplex DNA formed between the cut target and the donor, and / or synthesis-dependent strand annealing, which uses the donor to resynthesize the genetic information that will become part of the target, and / or related processes. In some instances, a donor polynucleotide, a portion of a donor polynucleotide, a copy of a donor polynucleotide, or a portion of a copy of a donor polynucleotide is introduced into the target DNA. See Wang et al. (2013) Cell 153:910-918; Mandalos et al. (2012) PLOS ONE 7:e45768:1-9; and Wang et al. (2013) Nat Biotechnol. 31:530-532, the entire contents of each of which are incorporated herein by reference for all purposes.

[0136] To produce a non-human animal with suppressed tolerance to a target foreign antigen of interest, one or more target genomic loci that affect the expression of an autoantigen that is homologous to the target foreign antigen or shares an epitope with the target foreign antigen can be targeted to suppress the expression of the autoantigen.Preferably, the expression of the autoantigen is eliminated.When the autoantigen is no longer expressed (for example, when the autoantigen is a protein, the protein is no longer expressed, or when the autoantigen is a specific epitope on a protein, the protein that contains the epitope is no longer expressed), the expression of the autoantigen is considered to be eliminated.

[0137] In one example, the genome of a non-human animal pluripotent cell that is not a one-cell embryo (e.g., an embryonic stem (ES) cell) may be contacted with a Cas protein, a first guide RNA that hybridizes to a first guide RNA recognition sequence within the target genomic locus, and a second guide RNA that hybridizes to a second guide RNA recognition sequence within the target genomic locus. In another example, the genome of a non-human animal one-cell embryo may be contacted with a Cas protein, a first guide RNA that hybridizes to a first guide RNA recognition sequence within the target genomic locus, and a second guide RNA that hybridizes to a second guide RNA recognition sequence within the target genomic locus.

[0138] In some methods provided herein, the targeted cell is a hybrid cell as defined elsewhere herein. Such methods may also include selecting a target region within the target genome locus as described elsewhere herein. The target region may be selected so that the target region has a higher percentage of sequence identity between the corresponding first and second chromosomes of a homologous chromosome pair compared to other segments in the target genome locus or the remainder of the target genome locus. For example, selecting a target region may include comparing the sequences of the corresponding first and second chromosomes of a homologous chromosome pair within the target genome locus, and selecting a target region that has a higher percentage of sequence identity between the corresponding first and second chromosomes of a homologous chromosome pair compared to all or part of the remainder of the target genome locus. Methods for selecting a target region are described in more detail elsewhere herein.

[0139] Optionally, the genome may further be contacted with another guide RNA that hybridizes to the guide RNA recognition sequence within the target genomic locus (or within a second target genomic locus that affects expression of the autoantigen, or that affects expression of a second autoantigen that is homologous to or shares an epitope of interest with the foreign antigen of interest), such as a third guide RNA that hybridizes to a third guide RNA recognition sequence within the target genomic locus, or a third guide RNA and a fourth guide RNA that hybridize to a fourth guide RNA recognition sequence within the target genomic locus. Contacting may include introducing a Cas protein and guide RNA into the cell in any form and by any method, as described in more detail elsewhere herein. The guide RNA forms a complex with the Cas protein, directing the Cas protein to the guide RNA recognition sequence at the target genomic locus, and the Cas protein cleaves the target genomic locus at the Cas protein cleavage site within the guide RNA recognition sequence. Cleavage with a Cas protein can generate a double-stranded break or a single-stranded break (e.g., when the Cas protein is a nickase). Examples and variations of Cas proteins and guide RNAs that can be used in the method are described elsewhere herein. Cleavage of a target genomic locus with a Cas protein can modify the target genomic locus within a pair of first and second chromosomes to generate a biallelic modification that silences expression of an autoantigen.

[0140] The foreign antigen of interest may be any foreign antigen for which an antigen-binding protein is desired. For example, the foreign antigen of interest may comprise, consist essentially of, or consist of all or a portion of a viral protein, bacterial protein, mammalian protein, simian protein, canine protein, feline protein, equine protein, bovine protein, rodent protein (e.g., rat or mouse), or human protein. For example, the foreign antigen of interest may comprise, consist essentially of, or consist of a human protein with one or more mutations or alterations. The foreign antigen of interest and the autoantigen may be homologous. For example, the foreign antigen of interest and the autoantigen may be orthologous or paralogous. Alternatively or additionally, the foreign antigen of interest and the autoantigen may comprise, consist essentially of, or consist of a shared epitope. The shared epitope may exist between homologous proteins or between different, non-homologous proteins. Epitopes may share linear amino acid sequence and / or structural compatibility (e.g., similar antigen surfaces even in the absence of primary sequence homology). For example, shared epitopes include substantially identical epitopes. When an epitope is shared between two antigens, antibodies to the epitope on the first antigen will typically also bind to the epitope on the second antigen.

[0141] Contacting can be carried out in the absence or presence of exogenous repair template, which is recombined with target genome locus to generate target gene modification.For example, cell can be 1-cell stage embryo, and exogenous repair template can be less than 5 kb in length.Examples of exogenous repair template are described elsewhere herein.

[0142] In some such methods, repair of the target nucleic acid with an exogenous repair template is achieved by homology-directed repair (HDR). Homologous recombination repair can be achieved when a Cas protein cleaves both strands of DNA at the target genomic locus to generate a double-stranded break, when the Cas protein is a nickase that cleaves one strand of DNA at the target genomic locus to generate a single-stranded break, or when a pair of Cas nickases is used to generate a double-stranded break formed by two offset nicks. In such methods, the exogenous repair template includes 5' and 3' homology arms corresponding to the 5' and 3' target sequences of the target genomic locus. The guide RNA recognition sequence or cleavage site may be adjacent to the 5' target sequence, adjacent to the 3' target sequence, adjacent to both the 5' and 3' target sequences, or adjacent to neither the 5' nor the 3' target sequence. The adjacent sequences include sequences within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of each other. Optionally, the foreign repair template may further include a nucleic acid insert flanked by the 5' and 3' homology arms, wherein the nucleic acid insert is inserted between the 5' and 3' target sequences. In the absence of the nucleic acid insert, the foreign repair template may function to remove genomic sequences between the 5' and 3' target sequences.

[0143] Alternatively, repair of the target nucleic acid using the exogenous repair template can be performed by non-homologous end joining (NHEJ)-mediated ligation. In such methods, at least one end of the exogenous repair template contains a short single-stranded region complementary to at least one overhang generated by Cas-mediated cleavage at the target genomic locus. The complementary ends in the exogenous repair template can be adjacent to the nucleic acid insert. For example, each end of the exogenous repair template can contain a short single-stranded region complementary to an overhang generated by Cas-mediated cleavage at the target genomic locus, and these complementary regions in the exogenous repair template can be adjacent to the nucleic acid insert. Overhangs (i.e., sticky ends) can be generated by cleaving the blunt ends of the double-stranded breaks generated by Cas-mediated cleavage. While such cleavage can generate microhomology regions necessary for fragment joining, it can also generate unwanted or uncontrollable modifications in the target nucleic acid. Alternatively, such overhangs can be generated using paired Cas nickases. For example, if the Cas protein is a nickase, the target genome locus can be contacted with a first and a second guide RNA that target opposite strands of DNA, thereby modifying the genome through double nicking. This can be achieved by contacting the target genome locus with two guide RNAs that hybridize to different guide RNA recognition sequences within the target genome locus. The two guide RNAs form two complexes with Cas nickase, and Cas nickase nicks the first strand of the target genome locus within one guide RNA recognition sequence and nicks the second strand of the target genome locus within the other guide RNA recognition sequence. Then, the exogenous repair template recombines with the target genome locus to generate the target gene modification.

[0144] In some methods, the nucleic acid insert comprises a sequence homologous or orthologous to all or part of the gene encoding the autoantigen. This can be useful, for example, when knocking out the autoantigen may result in embryonic lethality. The nucleic acid insert can be in any form described herein (e.g., targeting vector, LTVEC, ssODN, etc.) within the foreign repair template, and the nucleic acid insert can further comprise a selection cassette (e.g., a self-deletion selection cassette), or can lack a selection cassette. In such methods, for example, all or part of the gene encoding the autoantigen can be deleted and replaced with a corresponding homologous or orthologous sequence. For example, the entire gene encoding the autoantigen can be deleted and replaced with a corresponding homologous or orthologous sequence, or a part of the gene encoding a specific motif or region of the autoantigen can be deleted and replaced with a corresponding homologous or orthologous sequence. Optionally, the corresponding homologous or orthologous sequence can be from another species. For example, if the autoantigen is a mouse antigen, the corresponding homologous or orthologous sequence may be, for example, a homologous or orthologous rat, hamster, cat, dog, turtle, lemur, or human sequence. Alternatively or additionally, the homologous or orthologous sequence may contain one or more point mutations (e.g., 1, 2, 3, 4, 5, or more) compared to the substituted sequence. Such point mutations may, for example, function to eliminate the expression of one or more epitopes within the autoantigen. Such epitopes may be epitopes shared with the foreign antigen of interest. Optionally, such point mutations may result in conservative amino acid substitutions within the encoded polypeptide (e.g., substitution of aspartic acid [Asp, D] with glutamic acid [Glu, E]). Such amino acid substitutions may result in the expression of an autoantigen that maintains the function of the wild-type autoantigen but lacks epitopes present on the foreign antigen of interest and shared with the wild-type autoantigen.Similarly, deleting all or part of the gene encoding an autoantigen and replacing it with a corresponding homologous or orthologous sequence that lacks the epitope shared between the foreign antigen of interest and the autoantigen can result in the expression of a homolog or ortholog of the autoantigen that retains the function of the wild-type autoantigen but lacks the epitope present on the foreign antigen of interest and shared with the wild-type autoantigen. Antigen binding proteins against those epitopes can then be generated.

[0145] The modified non-human animal pluripotent cells may then be used to generate transgenic non-human animals using methods described elsewhere herein. For example, the modified non-human animal pluripotent cells may be introduced into a host embryo, and the host embryo may be implanted into a surrogate mother to generate a transgenic F0 generation non-human animal with a biallelic modification of the targeted genomic locus within a first and second chromosome pair to suppress or eliminate expression of the self-antigen. In the example of one-cell stage embryos, transgenic embryos may be selected and then implanted into a surrogate mother to generate a transgenic F0 generation non-human animal with a biallelic modification of the targeted genomic locus within a first and second chromosome pair to suppress or eliminate expression of the self-antigen. The F0 generation non-human animal may then be used to generate an antigen-binding protein against a foreign antigen of interest using methods described elsewhere herein.

[0146] A. Selection of Target Region Targeted gene modification via homologous recombination between an exogenous repair template (e.g., a targeting vector) and a target genomic locus can be extremely inefficient, especially in cell types other than rodent embryonic stem cells. Induction of one or more double-stranded DNA breaks by CRISPR / Cas9-induced cleavage can promote homozygous targeted gene recombination via homologous recombination (HR) between an exogenous repair template (e.g., a targeting vector) and a target genomic locus. CRISPR / Cas9 can also promote homozygous insertion or deletion mutations (i.e., identical biallelic modifications) via the non-homologous end joining (NHEJ) repair mechanism. For gene modification involving very large-scale humanization, combining a targeting vector with a CRISPR / Cas9 nuclease system guided by two guide RNAs targeting a single target genomic locus can further improve targeting efficiency beyond that achieved with a single guide RNA. Compared with targeting with one guide RNA (which generates biallelic alterations at low frequencies or not at all), targeting with two guide RNAs results in a significantly higher proportion of homozygously targeted cells, homozygously deleted cells, and compound heterozygously targeted cells (including hemizygously targeted cells). However, at some genomic loci, it may still be difficult to obtain homozygously targeted or homozygously deleted cells.

[0147] Unlike inbred mouse and rat strains commonly used in laboratory environments (which are homozygous at virtually all of their genomic loci), the sequences of the two alleles at the target genomic locus in hybrid cells (e.g., all human) are usually not 100% identical. However, as shown in the examples provided herein, the frequency of homozygous genome modification depends on the degree of sequence similarity between the two alleles at the target genomic locus, regardless of whether the initial CRISPR / Cas9-induced modification is generated by HR or NHEJ. This observation indicates that CRISPR / Cas9-induced homozygous gene modification is a homology-dependent phenomenon. Supporting this, as shown in the examples herein, CRISPR / Cas9-induced homozygous modification is often accompanied by loss of heterozygosity (LOH) in the allele sequence and structural polymorphisms (single nucleotide polymorphisms (SNVs) or structural polymorphisms (SVs)) linked to the target genomic locus on the same chromosome. LOH can involve either a local gene conversion mechanism of a variant at one of the target genomic loci or a global gene conversion (polar gene conversion) involving all variants on the telomeric side of the target genomic locus. Such gene conversion events must be mediated by homology-driven mitotic recombination mechanisms.

[0148] This finding provides guidance for designing CRISPR / Cas9-assisted homozygous targeting experiments. Selecting a target region where the two alleles share a high degree of sequence identity maximizes the chances of success. CRISPR / Cas9-assisted homozygous targeting in target regions with a high degree of sequence divergence between the two alleles appears to be largely unsuccessful. Even for loci with a high density of SNVs and SVs, the success rate can be improved by using guide RNAs or nuclease reagents that recognize sequences within the longest possible stretch of continuous allelic sequence identity within the target genomic locus, or sequences within the stretch of the target genomic locus with maximized allelic sequence identity.

[0149] The methods described herein may include selecting a target region so that sequence identity can be maximized across all or part of the target region between the corresponding first and second chromosomes of a homologous chromosome pair. In hybrid cells, the sequence on one copy of a homologous chromosome pair usually has some differences (e.g., single nucleotide polymorphisms) compared to the other copy of the chromosome pair. Therefore, such methods may include comparing the sequences of the corresponding first and second chromosomes of a homologous chromosome pair within the target genomic locus (e.g., human cells have 23 homologous chromosome pairs), and then selecting a target region within the target genomic locus so that sequence identity can be maximized across all or part of the target region between the corresponding first and second chromosomes of the homologous chromosome pair. If no sequences are available, such methods may further include sequencing the target genomic locus on each single chromosome within the homologous chromosome pair before comparing the sequences.

[0150] A target region may, for example, comprise, consist essentially of, or consist of any segment or region targeted by one of two or more guide RNAs or one or more exogenous repair templates in the methods disclosed herein, or any segment or region adjacent to a segment or region targeted by one of two or more guide RNAs or one or more exogenous repair templates in the methods disclosed herein. A target region may be a contiguous or non-contiguous genomic sequence. For example, a target region may comprise, consist essentially of, or consist of a genomic segment or region targeted for deletion, replacement, or insertion in accordance with the methods disclosed herein, and / or may comprise, consist essentially of, or consist of 5' and / or 3' sequences adjacent to a genomic segment or region targeted for deletion, replacement, or insertion in accordance with the methods disclosed herein. Preferably, the target region comprises, consists essentially of, or consists of the sequence immediately upstream and / or immediately downstream of the region targeted for deletion, substitution, or insertion by the methods disclosed herein (e.g., the sequence upstream and / or downstream of the region between two guide RNA recognition sequences or cleavage sites, or the sequence upstream and / or downstream of the region between the 5' and 3' target sequences of the exogenous repair template). By way of example, when two guide RNAs are used, the target region may comprise, consist essentially of, or consist of the 5' (i.e., upstream) and 3' (i.e., downstream) sequences flanking the region between the guide RNA recognition sequences or Cas cleavage sites. Examples of flanking sequence lengths are disclosed elsewhere herein.

[0151] In some methods, for example, a foreign repair template may be designed first, and then guide RNAs may be designed within regions flanking the 5' and 3' target sequences of the foreign repair template to maximize sequence identity within and / or adjacent (5', 3', or both) regions of the guide RNA recognition sequence (e.g., if two or more guide RNAs are used, adjacent to the region between the two most distant guide RNA recognition sequences). Alternatively, in some methods, for example, two or more guide RNAs may be designed first, and then the foreign repair template may be designed such that the 5' and 3' target sequences are flanked by two or more guide RNA recognition sequences to maximize sequence identity within and / or adjacent (5', 3', or both) regions of the 5' and 3' target sequences (e.g., adjacent to the region between the 5' and 3' target sequences).

[0152] By way of example, a target region may comprise, consist essentially of, or consist of a guide RNA recognition sequence for one of two or more guide RNAs. Alternatively or additionally, a target region may comprise, consist essentially of, or consist of 5' and / or 3' sequence flanking the guide RNA recognition sequence. The 5' flanking sequence may be, for example, at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 bp of flanking sequence, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150 kb of flanking sequence. Similarly, the 3' flanking sequence can be, for example, at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 bp of flanking sequence, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140 or 150 kb of flanking sequence.

[0153] In another example, a target region may comprise, consist essentially of, or consist of two or more guide RNA recognition sequences. Alternatively or additionally, a target region may comprise, consist essentially of, or consist of 5' and / or 3' sequences adjacent to a guide RNA recognition sequence. In methods using two guide RNAs, for example, a target region may comprise, consist essentially of, or consist of a genomic region adjacent to two guide RNA recognition sequences or cleavage sites, or a genomic region adjacent to and containing two guide RNA recognition sequences or cleavage sites. Alternatively or additionally, a target region may comprise, consist essentially of, or consist of 5' and / or 3' sequences adjacent to the region between two guide RNA recognition sequences or cleavage sites, or 5' and / or 3' sequences adjacent to the region between and containing the two guide RNA recognition sequences or cleavage sites. Instead of the above-described genomic regions flanking two guide RNA recognition sequences or cleavage sites, target regions may be selected similarly to methods using three or more guide RNAs, except that they would be genomic regions flanking the guide RNA recognition sequence of the most distant cleavage site. The 5' flanking sequence may be, for example, at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 bp of flanking sequence, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150 kb of flanking sequence. Similarly, the 3' flanking sequence can be, for example, at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 bp of flanking sequence, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140 or 150 kb of flanking sequence.

[0154] In methods using an exogenous repair template, for example, the target region may comprise, consist essentially of, or consist of the region adjacent to the 5' and 3' target sequences, or the region adjacent to and containing the 5' and 3' target sequences. Alternatively or additionally, the target region may comprise, consist essentially of, or consist of the 5' and / or 3' sequences adjacent to the genomic region between the 5' and 3' target sequences, or the 5' and / or 3' sequences adjacent to the genomic region between the 5' and 3' target sequences. The 5' flanking sequence can be, for example, at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 bp of flanking sequence, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140 or 150 kb of flanking sequence. Similarly, the 3' flanking sequence can be, for example, at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 bp of flanking sequence, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140 or 150 kb of flanking sequence.

[0155] Allelic sequence identity may be maximized in all of the target region or in a portion of the target region. As an example, allelic sequence identity may be maximized in the genomic region corresponding to at least one or each guide RNA recognition sequence, or in the region including at least one or each guide RNA recognition sequence. For example, allelic sequence identity may be maximized in at least one or each guide RNA recognition sequence. Alternatively, allelic sequence identity may be maximized in at least one or each guide RNA recognition sequence and in the 5' and / or 3' sequences adjacent to at least one or each guide RNA recognition sequence. Alternatively, allelic sequence identity may be maximized in the 5' and / or 3' sequences adjacent to at least one or each guide RNA recognition sequence. The 5' flanking sequence can be, for example, at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 bp of flanking sequence, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140 or 150 kb of flanking sequence. Similarly, the 3' flanking sequence can be, for example, at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 bp of flanking sequence, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140 or 150 kb of flanking sequence.

[0156] Alternatively or additionally, allelic sequence identity may be maximized in genomic regions corresponding to the 5' and / or 3' target sequences for the foreign repair template, or in regions including at least one or each of the 5' and 3' target sequences. For example, allelic sequence identity may be maximized in at least one or each of the 5' and 3' target sequences. Alternatively, allelic sequence identity may be maximized in at least one or each of the 5' and 3' target sequences, and in the 5' and / or 3' sequences adjacent to at least one or each of the 5' and 3' target sequences. Alternatively, allelic sequence identity may be maximized in the 5' and / or 3' sequences adjacent to at least one or each of the 5' and 3' target sequences. The 5' flanking sequence can be, for example, at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 bp of flanking sequence, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140 or 150 kb of flanking sequence. Similarly, the 3' flanking sequence can be, for example, at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 bp of flanking sequence, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140 or 150 kb of flanking sequence.

[0157] Alternatively or additionally, allelic sequence identity may be maximized in the sequences adjacent to the region targeted for deletion, substitution, or insertion. For example, in a method using two guide RNAs, allelic sequence identity may be maximized in the 5' and / or 3' sequences adjacent to the region between the two cleavage sites or the two guide RNA recognition sequences. In a method using three or more guide RNAs, allelic sequence identity may be maximized in the 5' and / or 3' sequences adjacent to the region between the two most distant cleavage sites or the two guide RNA recognition sequences. In another example, in a method using an exogenous repair template, allelic sequence identity may be maximized in the 5' and / or 3' sequences adjacent to the region between the 5' and 3' target sequences for the exogenous repair template (i.e., the genomic region targeted for deletion by the exogenous repair template). The 5' flanking sequence can be, for example, at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 bp of flanking sequence, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb or 150 kb of flanking sequence. Similarly, the 3' flanking sequence can be, for example, at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 bp of flanking sequence, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb or 150 kb of flanking sequence.

[0158] Selecting a target region such that sequence identity is maximized across all or a portion of the target region between corresponding first and second chromosomes of a homologous chromosome pair does not necessarily mean examining the target genomic locus on the first and second chromosomes of the homologous chromosome pair and then selecting the region with the highest allelic sequence identity compared to the remainder of the target genomic locus; instead, other factors may be taken into consideration. For example, when a target region comprises, consists essentially of, or consists of one or more guide RNA recognition sequences and / or sequences adjacent to one or more guide RNA recognition sequences, other factors that may be taken into consideration include, for example, which putative guide RNA recognition sequences are located within the region, whether the putative guide RNA recognition sequences are unique, which region the putative guide RNA recognition sequences are located within, how successfully or accurately predicting the putative guide RNA recognition sequences within the region, how successfully or accurately placing the putative guide RNA recognition sequences within the region adjacent to suitable 5' and 3' target sequences for the foreign repair template, how successfully or accurately placing the putative guide RNA recognition sequences within the region adjacent to other putative guide RNA recognition sequences, and how successfully or accurately placing the putative guide RNA recognition sequences within the region adjacent to the mutations targeted for repair. For example, the guide RNA recognition sequences are preferably unique target sites that do not exist elsewhere in the genome. See, e.g., US2014 / 0186843 (the entire contents of which are incorporated herein by reference for all purposes). Similarly, the specificity of the guide RNA may be optimized by varying the GC content and length of the targeting sequence relative to each other, and algorithms are available for designing or evaluating guide RNA targeting sequences that minimize off-target binding or interaction of the guide RNA. See, e.g., WO2016 / 094872 (the entirety of which is incorporated herein by reference for all purposes).In some methods, Cas9 proteins from different species may be considered or used (e.g., S. pyogenes Cas9 and S. aureus Cas9) to increase the number of available PAM sequences and thereby the number of potential guide RNA recognition sequences.

[0159] In one example, the target region may be selected so that all or part of the target region has a high percentage of sequence identity between the corresponding first and second chromosomes of a homologous chromosome pair. For example, the target region may be selected so that all or part of the target region has a minimum percentage of sequence identity between the corresponding first and second chromosomes of a homologous chromosome pair, such as at least 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.55%, 99.6%, 99.65%, 99.7%, 99.75%, 99.8%, 99.85%, 99.9%, 99.95% or 100% sequence identity.

[0160] In another example, a target region may be selected such that all or part of the target region has a low number or low density of single nucleotide polymorphisms between corresponding first and second chromosomes of a homologous chromosome pair. For example, a target region may be selected such that all or a portion of the target region has a maximum density of single nucleotide polymorphisms between corresponding first and second chromosomes of a homologous chromosome pair, such as 5, 4.9, 4.8, 4.7, 4.6, 4.5, 4.4, 4.3, 4.2, 4.1, 4, 3.9, 3.8, 3.7, 3.6, 3.5, 3.4, 3.3, 3.2, 3.1, 3, 2.9, 2.8, 2.7, 2.6, 2.5, 2.4, 2.3, 2.2, 2.1, 2, 1.9, 1.8, 1.7, 1.6, 1.5, 1.4, 1.3, 1.2, 1.1, 1, 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1 or less, or zero, per kb of sequence.

[0161] Optionally, the target region may be identical in corresponding first and second chromosomes of a homologous chromosome pair. Optionally, the target region may be present within the longest possible stretch of contiguous sequence identity within the target genomic locus.

[0162] Alternatively or additionally, a target region within a target genomic locus may be selected such that all or part of the target region has a high percentage of sequence identity or a low number or density of single nucleotide polymorphisms between corresponding first and second chromosomes of a homologous chromosome pair, compared to other regions within the target genomic locus.

[0163] For example, the target region may have a higher percentage of sequence identity or a lower density of single nucleotide polymorphisms compared to all or part of the remainder of the target genomic locus. For example, the target region may have at least 99.9% sequence identity between the corresponding first and second homologous chromosomes, while the remainder of the target genomic locus has 99.8% or less sequence identity between the corresponding first and second chromosomes.

[0164] For example, the target region may comprise, consist essentially of, or consist of one or more target genomic regions corresponding to one or more guide RNA recognition sequences, and the target region may have a high percentage of sequence identity or a low density of single nucleotide polymorphisms compared to other segments of the target genomic locus, such as genomic regions corresponding to one or more other potential guide RNA recognition sequences within the target genomic locus. In one example, the target region may comprise, consist essentially of, or consist of at least one or each of one or more guide RNA recognition sequences, and the target region may have a high percentage of sequence identity or a low density of single nucleotide polymorphisms compared to other segments of the target genomic locus, such as one or more other potential guide RNA recognition sequences within the target genomic locus. In another example, the target region may comprise, consist essentially of, or consist of at least one or each of one or more guide RNA recognition sequences and the 5' and / or 3' sequences adjacent to at least one or each of the one or more guide RNA recognition sequences, and the target region may have a high percentage of sequence identity or a low density of single nucleotide polymorphisms compared to other segments of the target genomic locus, such as one or more other potential guide RNA recognition sequences and their 5' and / or 3' flanking sequences within the target genomic locus. In yet another example, the target region may comprise, consist essentially of, or consist of the 5' and / or 3' sequences adjacent to at least one or each of one or more guide RNA recognition sequences, and the target region may have a high percentage of sequence identity or a low density of single nucleotide polymorphisms compared to other segments of the target genomic locus, such as the 5' and / or 3' flanking sequences of one or more other potential guide RNA recognition sequences within the target genomic locus.The 5' flanking sequence can be, for example, at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 bp of flanking sequence, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140 or 150 kb of flanking sequence. Similarly, the 3' flanking sequence can be, for example, at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 bp of flanking sequence, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140 or 150 kb of flanking sequence.

[0165] In methods using two guide RNAs, the target region may comprise, consist essentially of, or consist of a first target genomic region corresponding to a first guide RNA recognition sequence and / or may be within a second target genomic region corresponding to a second guide RNA recognition sequence, and the target region may have a high percentage of sequence identity or a low density of single nucleotide polymorphisms compared to other segments of the target genomic locus, such as genomic regions corresponding to one or more other potential guide RNA recognition sequences within the target genomic locus. For example, the target region may comprise, consist essentially of, or consist of the first guide RNA recognition sequence and / or the second guide RNA recognition sequence, and the target region may have a high percentage of sequence identity or a low density of single nucleotide polymorphisms compared to other segments of the target genomic locus, such as one or more other potential guide RNA recognition sequences within the target genomic locus. In another example, the target region may comprise, consist essentially of, or consist of a high percentage of first guide RNA recognition sequences and the 5' and / or 3' sequences adjacent to the first guide RNA recognition sequences and / or second guide RNA recognition sequences and the 5' and / or 3' sequences adjacent to the second guide RNA recognition sequences, and the target region may have a high percentage of sequence identity or a low density of single nucleotide polymorphisms relative to other segments of the target genomic locus, such as the genomic regions corresponding to one or more other potential guide RNA recognition sequences and their 5' and / or 3' flanking sequences within the target genomic locus. In yet another example, the target region may comprise, consist essentially of, or consist of 5' and / or 3' sequences adjacent to a first guide RNA recognition sequence and / or 5' and / or 3' sequences adjacent to a second guide RNA recognition sequence, and the target region may have a high percentage of sequence identity or a low density of single nucleotide polymorphisms compared to other segments of the target genomic locus, such as the 5' and / or 3' sequences adjacent to one or more other potential guide RNA recognition sequences within the target genomic locus.The 5' flanking sequence can be, for example, at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 bp of flanking sequence, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140 or 150 kb of flanking sequence. Similarly, the 3' flanking sequence can be, for example, at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 bp of flanking sequence, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140 or 150 kb of flanking sequence.

[0166] Thus, in methods that consider a single guide RNA to select a target region, for example, the selection of the target region can be accomplished by comparing two or more segments of the target genomic locus (each segment containing a distinct guide RNA recognition sequence that is not present elsewhere in the genome and at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, or 900 bp located 5', 3', or both sides of the distinct guide RNA recognition sequence). bp, 900 bp, 1,000 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb of contiguous sequence), selecting as the target region the segment having the highest percentage of sequence identity relative to the other segments. When two or more guide RNAs are used, the method may include selecting as the target region the two or more segments having the highest percentage of sequence identity relative to the other segments. Optionally, the one or more segments may comprise, consist essentially of, or consist of segments that correspond to the respective guide RNA recognition sequences in the target genomic locus but are not present elsewhere in the genome.

[0167] Alternatively or additionally, in methods using two guide RNAs, the target region may comprise, consist essentially of, or consist of the region between the first and second guide RNA recognition sequences or the first and second cleavage sites, and the target region may have a high percentage of sequence identity or a low density of single nucleotide polymorphisms relative to other segments of the target genomic locus, such as the region between one or more other pairs of potential guide RNA recognition sequences or cleavage sites within the target genomic locus. When three or more guide RNAs are used, the relevant region would be the region between the two most distant guide RNA recognition sequences or two cleavage sites.

[0168] Thus, in methods using two guide RNAs, for example, selecting a target region may involve comparing two or more segments of the target genomic locus (each segment comprising, consisting essentially of, or consisting of a region between a different guide RNA recognition sequence pair, where the guide RNA recognition sequence is not present elsewhere in the genome), and selecting as the target region the segment having the highest percentage of sequence identity relative to the other segments. Optionally, one or more segments comprise, consist essentially of, or consist of segments corresponding to each different guide RNA recognition sequence pair within the target genomic locus, where the guide RNA recognition sequence is not present elsewhere in the genome.

[0169] Alternatively or additionally, in methods using two guide RNAs, the target region may comprise, consist essentially of, or consist of the 5' and / or 3' sequences adjacent to the region between the first and second guide RNA recognition sequences or the first and second cleavage sites, and the genomic region between the first and second guide RNA recognition sequences or the first and second cleavage sites, and the target region may have a high percentage of sequence identity or a low density of single nucleotide polymorphisms compared to other segments of the target genomic locus, such as the region between one or more other pairs of potential guide RNA recognition sequences or cleavage sites within the target genomic locus, and the 5' and / or 3' sequences adjacent to the genomic region between one or more other pairs of potential guide RNA recognition sequences or cleavage sites. Preferably, the target region may comprise, consist essentially of, or consist of the 5' and 3' sequences adjacent to the genomic region between the first and second guide RNA recognition sequences or the first and second cleavage sites, and the genomic region between the first and second guide RNA recognition sequences or the first and second cleavage sites, and the target region may have a high percentage of sequence identity or a low density of single nucleotide polymorphisms compared to other segments of the target genomic locus, such as the 5' and 3' sequences adjacent to the genomic region between one or more other pairs of potential guide RNA recognition sequences or cleavage sites within the target genomic locus, and the genomic region between one or more other pairs of potential guide RNA recognition sequences or cleavage sites. When three or more guide RNAs are used, the relevant regions will be the 5' and / or 3' sequences adjacent to the genomic region between the two most distant guide RNA recognition sequences or two cleavage sites.The 5' flanking sequence can be, for example, at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 bp of flanking sequence, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140 or 150 kb of flanking sequence. Similarly, the 3' flanking sequence can be, for example, at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 bp of flanking sequence, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140 or 150 kb of flanking sequence.

[0170] Thus, in methods using two guide RNAs, for example, target regions can be selected by comparing two or more segments of the target genomic locus (each segment being at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1 kb, 2 kb, or 3 kb) between the different guide RNA recognition sequence pairs and on the 5', 3', or both sides of the genomic region between the different guide RNA recognition sequence pairs. Optionally, the one or more segments comprise, consist essentially of, or consist of segments corresponding to each different guide RNA recognition sequence pair within the target genomic locus, and the guide RNA recognition sequence is not present elsewhere in the genome, selecting as the target region the segment having the highest percentage of sequence identity compared to other segments. Optionally, the one or more segments comprise, consist essentially of, or consist of segments corresponding to each different guide RNA recognition sequence pair within the target genomic locus, and the guide RNA recognition sequence is not present elsewhere in the genome.

[0171] Alternatively or additionally, in methods using two guide RNAs, the target region may comprise, consist essentially of, or consist of 5' and / or 3' sequences adjacent to the genomic region between the first and second guide RNA recognition sequences or between the first and second cleavage sites, and the target region may have a high percentage of sequence identity or a low density of single nucleotide polymorphisms compared to other segments of the target genomic locus, such as the 5' and / or 3' sequences adjacent to the genomic region between one or more other pairs of potential guide RNA recognition sequences or cleavage sites within the target genomic locus. Preferably, the target region may comprise, consist essentially of, or consist of the 5' and 3' sequences flanking the genomic region between the first and second guide RNA recognition sequences or the first and second cleavage sites, and the target region may have a high percentage of sequence identity or a low density of single nucleotide polymorphisms compared to other segments of the target genomic locus, such as the 5' and 3' sequences flanking the genomic region between one or more other pairs of potential guide RNA recognition sequences or cleavage sites within the target genomic locus. When three or more guide RNAs are used, the relevant regions will be the 5' and / or 3' sequences flanking the genomic region between the two most distant guide RNA recognition sequences or two cleavage sites. The 5' flanking sequence can be, for example, at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 bp of flanking sequence, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140 or 150 kb of flanking sequence.Similarly, the 3' flanking sequence can be, for example, at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 bp of flanking sequence, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140 or 150 kb of flanking sequence.

[0172] Thus, in methods using two guide RNAs, for example, target regions can be selected by comparing two or more non-contiguous segments of the target genomic locus (each non-contiguous segment being at least 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1 kb, 2 kb, or 3 kb of the region between the different guide RNA recognition sequence pairs and on the 5', 3', or both sides of the genomic region between the different guide RNA recognition sequence pairs). selecting as the target region the non-contiguous segment that has the highest percentage of sequence identity compared to other non-contiguous segments (e.g., 1 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb of contiguous sequence, wherein the guide RNA recognition sequence is not present elsewhere in the genome). Optionally, the one or more non-contiguous segments comprise, consist essentially of, or consist of non-contiguous segments corresponding to each different guide RNA recognition sequence pair within the target genomic locus, wherein the guide RNA recognition sequence is not present elsewhere in the genome.

[0173] In methods using an exogenous repair template, the target region may comprise, consist essentially of, or consist of the region between the 5' target sequence and the 3' target sequence, and the target region may have a high percentage of sequence identity or a low density of single nucleotide polymorphisms compared to other segments of the target genomic locus. Alternatively or additionally, the target region may comprise, consist essentially of, or consist of the 5' target sequence and / or the 3' target sequence, and the target region may have a high percentage of sequence identity or a low density of single nucleotide polymorphisms compared to other segments of the target genomic locus. Preferably, the target region may comprise, consist essentially of, or consist of the 5' target sequence and the 3' target sequence, and the target region may have a high percentage of sequence identity or a low density of single nucleotide polymorphisms compared to other segments of the target genomic locus. For example, a target region may comprise, consist essentially of, or consist of regions adjacent to and containing the 5' target sequence and the 3' target sequence, and the target region may have a high percentage of sequence identity or a low density of single nucleotide polymorphisms compared to other segments of the target genomic locus.

[0174] Similarly, in methods using a foreign repair template, the target region may comprise, consist essentially of, or consist of 5' and / or 3' sequences adjacent to the genomic region between the 5' and 3' target sequences of the foreign repair template, or 5' and / or 3' sequences adjacent to the genomic region between the 5' and 3' target sequences of the foreign repair template and the genomic region including the 5' and 3' target sequences of the foreign repair template, and the target region may have a high percentage of sequence identity or a low density of single nucleotide polymorphisms compared to other segments of the target genomic locus. Preferably, the target region may comprise, consist essentially of, or consist of 5' and 3' sequences adjacent to the genomic region between the 5' and 3' target sequences of the foreign repair template, or 5' and 3' sequences within the genomic region between the 5' and 3' target sequences of the foreign repair template and the genomic region including the 5' and 3' target sequences of the foreign repair template, and the target region may have a high percentage of sequence identity or a low density of single nucleotide polymorphisms compared to other segments of the target genomic locus. Alternatively, the target region may comprise, consist essentially of, or consist of 5' and / or 3' sequences adjacent to the region between the 5' and 3' target sequences of the foreign repair template and the genomic region between the 5' and 3' target sequences, and the target region may have a high percentage of sequence identity or a low density of single nucleotide polymorphisms compared to other segments of the target genomic locus. Preferably, the target region may comprise, consist essentially of, or consist of the region between the 5' and 3' target sequences of the foreign repair template and the 5' and 3' sequences flanking the genomic region between the 5' and 3' target sequences, and the target region may have a high percentage of sequence identity or a low density of single nucleotide polymorphisms compared to other segments of the target genomic locus.The 5' flanking sequence can be, for example, at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 bp of flanking sequence, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140 or 150 kb of flanking sequence. Similarly, the 3' flanking sequence can be, for example, at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 bp of flanking sequence, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140 or 150 kb of flanking sequence.

[0175] The target region modified by the methods disclosed herein can include any segment or region (contiguous or non-contiguous) of DNA within a cell. The target region can be native to the cell, a heterologous or foreign segment of DNA integrated into the genome of the cell, or a combination thereof. Such heterologous or foreign segments of DNA can include transgenes, expression cassettes, polynucleotides encoding selectable markers, or heterologous or foreign regions of genomic DNA.

[0176] B. CRISPR / Cas system The methods disclosed herein utilize Clustered Regularly Interspersed Short Palindromic Repeats (CRISPR) / CRISPR-associated (Cas) systems, or components of such systems, to modify the genome in cells. A CRISPR / Cas system includes transcripts and other elements involved in the expression of Cas genes or inducing Cas gene activity. The CRISPR / Cas system may be a Type I, Type II, or Type III system. Alternatively, the CRISPR / Cas system may be, for example, a Type V system (e.g., subtype VA or subtype VB). The methods and compositions disclosed herein employ a CRISPR / Cas system that utilizes a CRISPR complex (including a guide RNA (gRNA) complexed with a Cas protein) for site-specific cleavage of nucleic acids.

[0177] The CRISPR / Cas systems used in the methods disclosed herein are non-naturally occurring. A "non-natural" system includes any that exhibit human intervention, such as one or more system components that have been altered or mutated from the component's natural state, that are at least substantially free of at least one other component with which the component is actually associated in nature, or that are associated with at least one other component with which the component is not naturally associated. For example, some CRISPR / Cas systems employ non-naturally occurring CRISPR complexes that include a gRNA and a Cas protein, both of which do not occur in nature. Other CRISPR / Cas systems employ a Cas protein that does not occur in nature, and other CRISPR / Cas systems employ a gRNA that does not occur in nature.

[0178] (1) Cas protein Cas proteins typically contain at least one RNA recognition or binding domain capable of interacting with a guide RNA (gRNA, described in more detail below). Cas proteins may also contain a nuclease domain (e.g., a DNase or RNase domain), a DNA-binding domain, a helicase domain, a protein-protein interaction domain, a dimerization domain, and other domains. The nuclease domain possesses catalytic activity for nucleic acid cleavage, which includes cleavage of covalent bonds in nucleic acid molecules. Cleavage may generate blunt or staggered ends and may be single-stranded or double-stranded. For example, wild-type Cas9 proteins typically generate blunt cleavage products. Alternatively, wild-type Cpf1 proteins (e.g., FnCpf1) may cleave 18 base pairs after the PAM sequence on the non-target strand and 23 bases after the target strand, resulting in cleavage products with a 5-nucleotide 5' overhang. The Cas protein may have full cleavage activity to generate a double-stranded break within the target nucleic acid (e.g., a double-stranded break with a blunt end), or may be a nickase that generates a single-stranded break within the target nucleic acid.

[0179] Examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5e (CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9 (Csn1 or Csx12), Cas10, Casl0d, CasF, CasG, CasH, Csy1, Csy2, Csy3, Cse1 (CasA), Cse2 (CasB), and Cse3 (C asE), Cse4 (CasC), Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4 and Cu1966, and homologs or modified forms thereof.

[0180] An exemplary Cas protein is the Cas9 protein or a protein derived from the Cas9 protein derived from the type II CRISPR / Cas system. The Cas9 protein is derived from the type II CRISPR / Cas system and generally shares four key motifs with a conserved structure. Motifs 1, 2, and 4 are RuvC-like motifs, and motif 3 is an HNH motif. Exemplary Cas9 proteins are Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Staphylococcus aureus, Nocardiopsis dassonvillei, Streptomyces pristinaespiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas sp., Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africanus, Acaryochloris marina, Neisseria meningitidis, or Campylobacter jejuni. Additional examples of Cas9 family members are described in WO2014 / 131833, the entire contents of which are incorporated by reference for all purposes. Cas9 from S. pyogenes (SpCas9) (accessed under SwissProt accession number Q99ZW2) is an exemplary Cas9 protein. Cas9 from S. aureus (SaCas9) (accessed under UniProt accession number J7RUA5) is another exemplary Cas9 protein. Cas9 from Campylobacter jejuni (CjCas9) (accessed under Uniprot accession number Q0P897) is another exemplary Cas9 protein. See, e.g., Kim et al. (2017) Nat. Comm. 8:14500 (incorporated herein by reference in its entirety for all purposes). SaCas9 is smaller than SpCas9, and CjCas9 is smaller than both SaCas9 and SpCas9.

[0181] Another example of a Cas protein is the Cpf1 (CRISPR from Prevotella and Francisella 1) protein. Cpf1 is a large protein (approximately 1300 amino acids) that contains an equivalent of the unique arginine-rich cluster of Cas9, in addition to a RuvC-like nuclease domain that is homologous to the corresponding domain in Cas9. However, Cpf1 lacks the HNH nuclease domain present in the Cas9 protein, and in contrast to Cas9, which contains a long insert containing the HNH domain, the RuvC-like domain is contiguous within the Cpf1 sequence. For example, Zetsche et al. See, e.g., J. et al. (2015) Cell 163(3):759-771, the entire contents of which are incorporated by reference herein for all purposes. Exemplary Cpf1 proteins are found in Francisella tularensis 1, Francisella tularensis subsp. novicida, Prevotella albensis, Lachnospiraceae bacterium MC2017, and the like. 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Smithella sp.SCADC, Acidaminococcus sp.BV3L6, Lachnospiraceae bacterium MA2020, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi 237, Leptospira inadai, Lachnospiraceae bacterium ND2006, Porphyromonas crevioricanis 3, Prevotella disiens, and Porphyromonas macacae. Cpf1 from Francisella novicida U112 (FnCpf1; registered under Uniprot accession number A0Q7Q2) is an exemplary Cpf1 protein.

[0182] The Cas protein may be a wild-type protein (i.e., a naturally occurring protein), a modified Cas protein (i.e., a Cas protein mutant), or a fragment of a wild-type or modified Cas protein. The Cas protein may also be an active mutant or fragment of the wild-type or modified Cas protein that is involved in catalytic activity. The active mutant or fragment that is involved in catalytic activity may contain at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the wild-type or modified Cas protein or a portion thereof, and the active mutant retains the ability to cleave at the desired cleavage site and therefore retains nick-inducing or double-strand break-inducing activity. Assays for nick-inducing or double-strand break-inducing activity are well known and typically measure the overall activity and specificity of the Cas protein on a DNA substrate containing the cleavage site.

[0183] One example of a modified Cas protein is the modified SpCas9-HF1 protein, which is a high-fidelity mutant (N497A / R661A / Q695A / Q926A) of Streptococcus pyogenes Cas9 with modifications designed to suppress non-specific DNA contact. See, e.g., Kleinstiver et al. (2016) Nature 529(7587):490-495 (incorporated herein by reference in its entirety for all purposes). Another example of a modified Cas protein is the modified eSpCas9 mutant (K848A / K1003A / R1060A) designed to suppress off-target effects. See, e.g., Slaymaker et al. (2016) Science 351(6268):84-88 (incorporated herein by reference in its entirety for all purposes). Other SpCas9 mutants include K855A and K810A / K1003A / R1060A.

[0184] Cas proteins may be modified to increase or decrease one or more of nucleic acid binding affinity, nucleic acid binding specificity, and enzymatic activity. Cas proteins may also be modified to change any other activity or property of the protein, such as stability. For example, one or more nuclease domains of a Cas protein may be modified, deleted, or inactivated, or the Cas protein may be truncated to remove domains that are not essential for protein function, or to optimize (e.g., enhance or suppress) the activity of a Cas protein.

[0185] A Cas protein may contain at least one nuclease domain, such as a DNase domain. For example, wild-type Cpf1 proteins typically contain a RuvC-like domain (possibly in a dimeric structure) that cleaves both strands of target DNA. A Cas protein may also contain at least two nuclease domains, such as a DNase domain. For example, wild-type Cas9 proteins typically contain a RuvC-like nuclease domain and an HNH-like nuclease domain. The RuvC domain and the HNH domain can each cleave different strands of double-stranded DNA to generate a double-strand break in the DNA. See, e.g., Jinek et al. (2012) Science 337:816-821, the entire contents of which are incorporated herein by reference for all purposes.

[0186] One or both nuclease domains may be deleted or mutated so that they no longer function or their nuclease activity is suppressed. When one of the nuclease domains is deleted or mutated, the resulting Cas protein (e.g., Cas9) may be called a nickase and can generate single-stranded breaks at the guide RNA recognition sequence in double-stranded DNA but cannot generate double-stranded breaks (i.e., it can cleave either the complementary strand or the non-complementary strand, but not both). When both nuclease domains are deleted or mutated, the resulting Cas protein (e.g., Cas9) will have reduced functionality for cleaving both strands of double-stranded DNA (e.g., a nuclease-null Cas protein). One example of a mutation that converts Cas9 into a nickase is the D10A (aspartic acid to alanine at position 10 of Cas9) mutation in the RuvC domain of S. pyogenes-derived Cas9. Similarly, H939A (histidine to alanine at amino acid position 839), H840A (histidine to alanine at amino acid position 840), or N863A (asparagine to alanine at amino acid position N863) within the HNH domain of Cas9 from S. pyogenes can convert Cas9 into a nickase. Other examples of mutations that convert Cas9 into a nickase include the corresponding mutations in Cas9 from S. thermophilus. See, e.g., Sapranauskas et al. See, for example, WO2013 / 141680, WO2013 / 176772, WO2013 / 176772, and WO2013 / 142578, the entire contents of each of which are incorporated herein by reference for all purposes. Such mutations can be generated using methods such as site-directed mutagenesis, PCR-mediated mutagenesis, or total gene synthesis. Other examples of nickase-generating mutations can be found, for example, in WO2013 / 176772 and WO2013 / 142578, the entire contents of each of which are incorporated herein by reference for all purposes. When all of the nuclease domains are deleted or mutated in a Cas protein (e.g., when both nuclease domains are deleted or mutated in a Cas9 protein), the resulting Cas protein (e.g., Cas9) will have reduced functionality for cleaving both strands of double-stranded DNA (e.g., a nuclease-null or nuclease-inactive Cas protein). One particular example is the D10A / H840A S. pyogenes Cas9 double mutation, or a corresponding double mutation in a Cas9 from another species when optimally aligned to S. pyogenes Cas9. Another particular example is the D10A / N863A S. pyogenes Cas9 double mutation, or a corresponding double mutation in a Cas9 from another species when optimally aligned to S. pyogenes Cas9.

[0187] Examples of inactivating mutations within the catalytic domain of the Staphylococcus aureus Cas9 protein are also known. The S. aureus Cas9 enzyme (SaCas9) may include substitutions at positions N580 (e.g., an N580A substitution) and D10 (e.g., a D10A substitution) to generate a nuclease-inactive Cas protein. See, e.g., WO2016 / 106236, the entire contents of which are incorporated herein by reference for all purposes.

[0188] Examples of inactivating mutations within the catalytic domain of the Cpf1 protein are also known. For the Cpf1 protein from Francisella novicida U112 (FnCpf1), the Cpf1 protein from Acidaminococcus sp. BV3L6 (AsCpf1), the Cpf1 protein from Lachnospiraceae bacterium ND2006 (LbCpf1), and the Cpf1 protein from Moraxella bovoculi 237 (MbCpf1 Cpf1), such mutations may include mutations at positions 908, 993, or 1263 of AsCpf1 or corresponding positions in Cpf1 orthologs, or mutations at positions 832, 925, 947, or 1180 of LbCpf1 or corresponding positions in Cpf1 orthologs. Such mutations can include, for example, one or more of the D908A, E993A, and D1263A mutations in AsCpf1 or corresponding mutations in Cpf1 orthologs, or the D832A, E925A, D947A, and D1180A mutations in LbCpf1 or corresponding mutations in Cpf1 orthologs. See, e.g., US2016 / 0208243 (incorporated herein by reference in its entirety for all purposes).

[0189] The Cas protein may also be functionally linked to a heterologous polypeptide as a fusion protein. For example, the Cas protein may be fused to a cleavage domain, an epigenetic modification domain, a transcriptional activation domain, or a transcriptional repressor domain. See WO2014 / 089290 (the entire contents of which are incorporated herein by reference for all purposes). The Cas protein may also be fused to a heterologous polypeptide to improve or decrease stability. The fusion domain or fused heterologous polypeptide may be located at the N-terminus, C-terminus, or internally within the Cas protein.

[0190] An example of a Cas fusion protein is a Cas protein fused to a heterologous polypeptide that confers subcellular localization. Such heterologous polypeptides can include, for example, one or more nuclear localization signals (NLSs), such as an SV40 NLS for targeting the nucleus, a mitochondrial localization signal for targeting the mitochondria, or an ER retention signal. See, e.g., Lange et al. (2007) J. Biol. Chem. 282:5101-5105, the entire contents of which are incorporated herein by reference for all purposes. Other suitable NLSs include the α-importin NLS. Such subcellular localization signals can be located at the N-terminus, C-terminus, or anywhere within the Cas protein. The NLS can include a stretch of basic amino acids and can be a monopartite or bipartite sequence. Optionally, the Cas protein includes two or more NLSs, including an N-terminal NLS (e.g., an α-importin NLS) and / or a C-terminal NLS (e.g., an SV40 NLS).

[0191] The Cas protein may also be functionally linked to a cell-penetrating domain. For example, the cell-penetrating domain can be derived from the TLM cell-penetrating motif from HIV-1 TAT protein, human hepatitis B virus, MPG, Pep-1, VP22, the cell-penetrating peptide from herpes simplex virus, or a polyarginine peptide sequence. See, for example, WO2014 / 089290 (the entire contents of which are incorporated herein by reference for all purposes). The cell-penetrating domain may be located at the N-terminus, C-terminus, or anywhere within the Cas protein.

[0192] The Cas protein may also be operably linked to a heterologous polypeptide, such as a fluorescent protein, a purification tag, or an epitope tag, to facilitate tracking or purification. Examples of fluorescent proteins include green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, eGFP, Emerald, Azami Green, Monomeric Azami). Green, CopGFP, AceGFP, ZsGreenl), yellow fluorescent proteins (e.g., YFP, eYFP, Citrine, Venus, YPet, PhiYFP, ZsYellowl), blue fluorescent proteins (e.g., eBFP, eBFP2, Azurite, mKalamal, GFPuv, Sapphire, T-sapphire), cyan-green fluorescent proteins (e.g., eCFP, Cerulean, CyPet, AmCyanl, Midoriishi-Cyan), red fluorescent proteins (mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-Monomer, HcRed-Tandem, HcRedl, AsRed2, eqFP611, mRaspberry, mStrawberry, Jred), orange-yellow fluorescent proteins (mOrange, mKO, Kusabira-Orange, Monomeric Examples of tags include glutathione-S-transferase (GST), chitin-binding protein (CBP), maltose-binding protein, thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tag, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, hemagglutinin (HA), nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, histidine (His), biotin carboxyl carrier protein (BCCP), and calmodulin.

[0193] The Cas9 protein may also be tethered to an exogenous repair template or labeled nucleic acid. Such tethering (i.e., physical linkage) can be achieved through covalent or non-covalent interactions, and the tethering can be direct (e.g., via direct fusion or chemical conjugates, which can be achieved by modifying cysteine ​​or lysine residues on the protein or by intein modification) or via one or more intervening linker or adapter molecules (such as streptavidin or aptamers). For example, Pierce et al.(2005)Mini Rev.Med.Chem.5(1):41-55;Duckworth et al.(2007)Angew.Chem.Int.Ed.Engl.46(46):8819-8822;Schaeffer and Dixon(2009)Australian J. Chem.62(10):1328-1332; Goodman et al. (2009) Chembiochem.10(9):1551-1557; and Khatwani See, e.g., Wang et al. (2012) Bioorg. Med. Chem. 20(14):4532-4539 (the entire contents of each are incorporated herein by reference for all purposes). Non-covalent strategies for synthesizing protein-nucleic acid conjugates include the biotin-streptavidin method and the nickel-histidine method. Covalent protein-nucleic acid conjugates can be synthesized by appropriately linking functionalized nucleic acids to proteins using a wide variety of chemical reactions. Some of these chemical reactions involve direct binding of oligonucleotides to amino acid residues (e.g., lysine amines or cysteine ​​thiols) on the protein surface, while many other conjugation schemes require post-translational modifications of the protein or the involvement of catalytic or reactive protein domains. Methods for covalently linking proteins to nucleic acids include, for example, chemical cross-linking of oligonucleotides to lysine or cysteine ​​residues of proteins, expressed protein ligation, chemoenzymatic methods, and the use of photoaptamers. The foreign repair template or labeled nucleic acid may be tethered to the C-terminus, N-terminus, or internal region of the Cas9 protein. Preferably, the foreign repair template or labeled nucleic acid is tethered to the C-terminus or N-terminus of the Cas9 protein. Similarly, the Cas9 protein may be tethered to the 5'-terminus, 3'-terminus, or internal region of the foreign repair template or labeled nucleic acid. In other words, the foreign repair template or labeled nucleic acid may be tethered in any direction and polarity. Preferably, the Cas9 protein is tethered to the 5'-terminus or 3'-terminus of the foreign repair template or labeled nucleic acid.

[0194] The Cas protein may be provided in any form. For example, the Cas protein may be provided in the form of a protein, such as a Cas protein complexed with a gRNA. Alternatively, the Cas protein may be provided in the form of a nucleic acid encoding the Cas protein, such as RNA (e.g., messenger RNA (mRNA)) or DNA. Optionally, the nucleic acid encoding the Cas protein may be codon-optimized for efficient translation into a protein in a particular cell or organism. For example, the nucleic acid encoding the Cas protein may be modified to substitute codons more frequently used in bacterial cells, yeast, human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, or any other host cell of interest, compared to the native polynucleotide sequence. When the nucleic acid encoding the Cas protein is introduced into a cell, the Cas protein may be expressed transiently, conditionally, or constitutively within the cell.

[0195] The nucleic acid encoding the Cas protein may be stably integrated into the genome of the cell or operably linked to a promoter active in the cell. Alternatively, the nucleic acid encoding the Cas protein may be operably linked to a promoter within an expression construct. Expression constructs include any nucleic acid construct capable of inducing expression of a gene or other nucleic acid sequence of interest (e.g., a Cas gene) and transferring such a nucleic acid sequence of interest into a target cell. For example, the nucleic acid encoding the Cas protein may be within a targeting vector containing a nucleic acid insert and / or a vector containing DNA encoding a gRNA. Alternatively, the nucleic acid encoding the Cas protein may be within a vector or plasmid separate from the targeting vector containing a nucleic acid insert and / or separate from the vector containing DNA encoding the gRNA. Promoters that can be used in expression constructs include, for example, promoters active in one or more of eukaryotic cells, human cells, non-human cells, mammalian cells, non-human mammalian cells, rodent cells, mouse cells, rat cells, hamster cells, rabbit cells, pluripotent cells, embryonic stem (ES) cells, or zygotes. Such promoters may be, for example, conditional, inducible, constitutive, or tissue-specific promoters. Optionally, the promoter may be a bidirectional promoter, driving both Cas protein expression in one direction and guide RNA expression in the other. Such a bidirectional promoter may consist of (1) a complete, normal, unidirectional Pol III promoter containing three external control elements: a distal sequence element (DSE), a proximal sequence element (PSE), and a TATA box, and (2) a second, basic Pol III promoter containing a PSE and a TATA box fused in reverse orientation to the 5' end of the DSE. For example, in the H1 promoter, the DSE is adjacent to the PSE and TATA box; a promoter can be made bidirectional by adding a PSE and a TATA box from the U6 promoter to create a hybrid promoter in which transcription is controlled in reverse orientation.See, e.g., US2016 / 0074535, the entire contents of which are incorporated by reference for all purposes. The use of bidirectional promoters to simultaneously express genes encoding Cas proteins and guide RNAs allows for the creation of compact expression cassettes that facilitate transport.

[0196] (2) Guide RNA A "guide RNA" or "gRNA" is an RNA molecule that binds to a Cas protein (e.g., a Cas9 protein) and directs the Cas protein to a specific location within a target DNA. A guide RNA may contain two segments: a "DNA-targeting segment" and a "protein-binding segment." A "segment" includes a section or region of a molecule, such as a continuous stretch of nucleotides within an RNA. Some gRNAs, such as the gRNA of Cas9, may contain two independent RNA molecules: an "activating RNA" (e.g., a tracrRNA) and a "targeting RNA" (e.g., a CRISPR RNA or crRNA). Other gRNAs are single RNA molecules (single RNA polynucleotides) and are sometimes referred to as "single-molecule gRNAs," "single guide RNAs," or "sgRNAs." See, e.g., WO2013 / 176772, WO2014 / 065596, WO2014 / 089290, WO2014 / 093622, WO2014 / 099750, WO2013 / 142578, and WO2014 / 131833, the entire contents of each of which are incorporated herein by reference for all purposes. For example, in Cas9, a single guide RNA may comprise a crRNA fused to a tracrRNA (e.g., via a linker). For example, in Cpfl, only the crRNA is required to achieve binding to or cleavage of the target sequence. The terms "guide RNA" and "gRNA" include both dual-molecule gRNAs (i.e., modular gRNAs) and single-molecule gRNAs.

[0197] Exemplary bimolecular gRNAs include a crRNA-like ("CRISPR RNA," or "targeting RNA," or "crRNA" or "crRNA repeat") molecule and a corresponding tracrRNA-like ("trans-activating CRISPR RNA," or "activating RNA," or "tracrRNA") molecule. The crRNA contains both the DNA-targeting segment (single strand) of the gRNA and a stretch of nucleotides that forms one half of the dsRNA duplex of the protein-binding segment of the gRNA.

[0198] The corresponding tracrRNA (activator RNA) contains a stretch of nucleotides that forms the other half of the dsRNA duplex of the protein-binding segment of the gRNA. The stretch of nucleotides in the crRNA is complementary to the stretch of nucleotides in the tracrRNA and hybridizes to the stretch of nucleotides in the tracrRNA to form the dsRNA duplex of the protein-binding domain of the gRNA. Thus, each crRNA is sometimes said to have a corresponding tracrRNA.

[0199] In systems requiring both a crRNA and a tracrRNA, the crRNA and the corresponding tracrRNA hybridize to form a gRNA. In systems requiring only a crRNA, the crRNA may be the gRNA. The crRNA also carries a single-stranded DNA targeting segment that hybridizes to the guide RNA recognition sequence. When used for intracellular modification, the precise sequence of any crRNA or tracrRNA molecule may be designed to be specific to the species in which the RNA molecule is used. See, for example, Mali et al. (2013) Science 339:823-826; Jinek et al. (2012) Science 337:816-821; Hwang et al. (2013) Nat. Biotechnol. 31:227-229; Jiang et al. (2013) Nat. Biotechnol. 31:233-239; and Cong et al. (2013) Science 339:819-823, the entire contents of each of which are incorporated herein by reference for all purposes.

[0200] The DNA-targeting segment (crRNA) of any gRNA contains a nucleotide sequence complementary to a sequence within the target DNA (i.e., the guide RNA recognition sequence). The DNA-targeting segment of the gRNA interacts with the target DNA in a sequence-specific manner through hybridization (i.e., base pairing). Thus, the nucleotide sequence of the DNA-targeting segment can vary and determines the location within the target DNA where the gRNA and target DNA interact. The DNA-targeting segment of the gRNA can be modified to hybridize to any desired sequence within the target DNA. Natural crRNAs vary depending on the CRISPR / Cas system and organism, but often contain a targeting segment of 21-72 nucleotides in length flanked by two direct repeats (DRs) of 21-46 nucleotides in length (see, e.g., WO2014 / 131833, the entire contents of which are incorporated herein by reference for all purposes). In the example of S. pyogenes, the DRs are 36 nucleotides in length and the targeting segment is 30 nucleotides in length. The DR located 3' is complementary to the corresponding tracrRNA, hybridizes to the corresponding tracrRNA, and then binds to the Cas protein.

[0201] A DNA-targeting segment may have a length of at least about 12 nucleotides, at least about 15 nucleotides, at least about 17 nucleotides, at least about 18 nucleotides, at least about 19 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 35 nucleotides, or at least about 40 nucleotides. Such a DNA-targeting segment may have a length of about 12 nucleotides to about 100 nucleotides, about 12 nucleotides to about 80 nucleotides, about 12 nucleotides to about 50 nucleotides, about 12 nucleotides to about 40 nucleotides, about 12 nucleotides to about 30 nucleotides, about 12 nucleotides to about 25 nucleotides, or about 12 nucleotides to about 20 nucleotides. For example, a DNA-targeting segment may be about 15 nucleotides to about 25 nucleotides (e.g., about 17 nucleotides to about 20 nucleotides, or about 17 nucleotides, about 18 nucleotides, about 19 nucleotides, or about 20 nucleotides). See, e.g., US2016 / 0024523 (incorporated herein by reference in its entirety for all purposes). In S. pyogenes-derived Cas9, the standard DNA-targeting segment is 16-20 nucleotides or 17-20 nucleotides in length. In S. aureus-derived Cas9, the standard DNA-targeting segment is 21-23 nucleotides in length. In Cpf1, the standard DNA-targeting segment is at least 16 nucleotides in length or at least 18 nucleotides in length.

[0202] The tracrRNA may be in any form (e.g., full-length tracrRNA or active truncated tracrRNA) and may be of various lengths. The tracrRNA may include the primary transcript or a processed form. For example, the tracrRNA (as part of a single guide RNA or as an independent molecule that is part of a bimolecular gRNA) may comprise or consist of all or a portion of a wild-type tracrRNA sequence (e.g., a wild-type tracrRNA sequence of about or greater than about 20, 26, 32, 45, 48, 54, 63, 67, 85, or more nucleotides). Examples of wild-type tracrRNA sequences from S. pyogenes include 171-nucleotide, 89-nucleotide, 75-nucleotide, and 65-nucleotide versions. See, e.g., Deltcheva et al. (2011) Nature 471:602-607; WO2014 / 093661, the entire contents of each of which are incorporated herein by reference for all purposes. Examples of tracrRNA within a single guide RNA (sgRNA) include tracrRNA segments found within sgRNAs +48, ​​+54, +67, and +85 (where "+n" indicates that up to +n nucleotides of wild-type tracrRNA are included within the sgRNA). See US8,697,359, the entire contents of which are incorporated by reference for all purposes.

[0203] The percent complementarity between the DNA-targeting sequence and the guide RNA recognition sequence within the target DNA may be at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%). The percent complementarity between the DNA-targeting sequence and the guide RNA recognition sequence within the target DNA may be at least 60% over about 20 contiguous nucleotides. As an example, the percent complementarity between the DNA-targeting sequence and the guide RNA recognition sequence within the target DNA is 100% over 14 contiguous nucleotides at the 5'-end of the guide RNA recognition sequence within the complementary strand of the target DNA and only 0% over the remainder. In such a case, the DNA-targeting sequence can be considered to be 14 nucleotides in length. In another example, the percent complementarity between the DNA targeting sequence and the guide RNA recognition sequence in the target DNA is 100% over 7 consecutive nucleotides at the 5' end of the guide RNA recognition sequence in the complementary strand of the target DNA, and only 0% over the remainder. In such cases, the DNA targeting sequence can be considered to be 7 nucleotides long. In some guide RNAs, at least 17 nucleotides in the DNA targeting sequence are complementary to the target DNA. For example, the DNA targeting sequence may be 20 nucleotides long and may contain 1, 2, or 3 mismatches with the target DNA (guide RNA recognition sequence). Preferably, the mismatches are not adjacent to the protospacer adjacent motif (PAM) sequence (e.g., the mismatch is within the 5' end of the DNA targeting sequence, or the mismatch is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 base pairs away from the PAM sequence).

[0204] The protein-binding segment of the gRNA may contain two stretches of nucleotides that are complementary to each other. The complementary nucleotides of the protein-binding segment hybridize to form a double-stranded RNA duplex (dsRNA). The protein-binding segment of the gRNA interacts with a Cas protein, and the gRNA guides the bound Cas protein to a specific nucleotide sequence within the target DNA via the DNA-targeting segment.

[0205] The single guide RNA comprises a DNA targeting segment and a scaffold sequence (i.e., the protein-binding or Cas-binding sequence of the guide RNA). Exemplary scaffold sequences include GTTGGAACCATTCAAAACAGCATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC (SEQ ID NO: 150), GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC (SEQ ID NO: 151), and GTTTAAGAGCTATGCTGGAAACAGCATAGCAAGTTTAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC (SEQ ID NO: 152).

[0206] Guide RNAs may contain modifications or sequences that confer additional desirable characteristics (e.g., altered or modulated stability, intracellular targeting, tracking with fluorescent labels, binding sites for proteins or protein complexes, etc.). Examples of such modifications include, for example, a 5' cap (e.g., a 7-methylguanylate cap (m7G)); a 3' polyadenylation tail (i.e., a 3' poly(A) tail); a riboswitch sequence (e.g., to allow for regulated stability and / or regulated accessibility by proteins and / or protein complexes); a stability control sequence; a sequence that forms a dsRNA duplex (i.e., a hairpin); a modification or sequence that targets the RNA to a subcellular location (e.g., the nucleus, mitochondria, chloroplasts, etc.); a modification or sequence that provides for tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a moiety that facilitates fluorescent detection, a sequence that allows for fluorescent detection, etc.); a modification or sequence that provides binding sites for proteins (e.g., proteins that act on DNA, including transcriptional activators, transcriptional repressors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, etc.); and combinations thereof. Other examples of modifications include genetically engineered stem-loop duplex structures, genetically engineered bulge regions, genetically engineered hairpins 3' of stem-loop duplex structures, or any combination thereof. See, e.g., US2015 / 0376586 (incorporated herein by reference in its entirety for all purposes). The bulge may be an unpaired region of nucleotides within a duplex composed of a crRNA-like region and a minimal tracrRNA-like region. The bulge may contain an unpaired 5'-XXXY-3' (where X is any purine and Y can be a nucleotide that can form a wobble pair with a nucleotide on the opposite strand) on one side of the duplex and an unpaired nucleotide region on the other side of the duplex.

[0207] Guide RNAs may be provided in any form. For example, gRNAs may be provided in the form of RNA, either as two molecules (independent crRNA and tracrRNA) or as a single molecule (sgRNA), and may optionally be provided in the form of a complex with a Cas protein. For example, gRNAs can be prepared by in vitro transcription, for example, using T7 RNA polymerase (see, e.g., WO2014 / 089290 and WO2014 / 065596, the entire contents of each of which are incorporated herein by reference for all purposes). Guide RNAs can also be prepared by chemical synthesis.

[0208] The gRNA may also be provided in the form of DNA encoding the gRNA. The DNA encoding the gRNA may encode a single RNA molecule (sgRNA) or separate RNA molecules (e.g., separate crRNA and tracrRNA). In the latter case, the DNA encoding the gRNA may be provided as a single DNA molecule or as separate DNA molecules encoding the crRNA and tracrRNA, respectively.

[0209] When the gRNA is provided in the form of DNA, the gRNA may be expressed transiently, conditionally, or constitutively in the cell. The DNA encoding the gRNA may be stably integrated into the genome of the cell or operably linked to a promoter active in the cell. Alternatively, the DNA encoding the gRNA may be operably linked to a promoter within an expression construct. For example, the DNA encoding the gRNA may be within a vector containing an exogenous repair template and / or a vector containing a nucleic acid encoding a Cas protein. Alternatively, the DNA encoding the gRNA may be within a vector or plasmid independent of the vector containing an exogenous repair template and / or a vector containing a nucleic acid encoding a Cas protein. Promoters that can be used in such expression constructs include, for example, promoters active in one or more of eukaryotic cells, human cells, non-human cells, mammalian cells, non-human mammalian cells, rodent cells, mouse cells, rat cells, hamster cells, rabbit cells, pluripotent cells, embryonic stem (ES) cells, or zygotes. Such promoters may be, for example, conditional promoters, inducible promoters, constitutive promoters, or tissue-specific promoters. Such promoters may also be, for example, bidirectional promoters. Specific examples of suitable promoters include RNA polymerase III promoters, such as the human U6 promoter, rat U6 polymerase III promoter, or mouse U6 polymerase III promoter.

[0210] (3) Guide RNA recognition sequence The term "guide RNA recognition sequence" includes a nucleic acid sequence present in the target DNA to which the DNA-targeting segment of the gRNA binds when sufficient conditions for binding exist. For example, the guide RNA recognition sequence includes a sequence to which the guide RNA is designed to have complementarity, where hybridization between the guide RNA recognition sequence and the DNA-targeting sequence promotes the formation of a CRISPR complex. Perfect complementarity is not necessarily required, provided there is sufficient complementarity for hybridization to occur and promote the formation of a CRISPR complex. The guide RNA recognition sequence also includes a cleavage site for a Cas protein, which is described in more detail below. The guide RNA recognition sequence may include any polynucleotide that can be located, for example, in the nucleus or cytoplasm of a cell, or in a cell organelle (such as a mitochondria or chloroplast).

[0211] The guide RNA recognition sequence in the target DNA may be targeted by (i.e., bound to, hybridized with, or complementary to) the Cas protein or gRNA. Suitable DNA / RNA binding conditions include physiological conditions typically present within cells. Other suitable DNA / RNA binding conditions (e.g., conditions within a cell-free system) are well known in the art (e.g., Molecular Cloning: A Laboratory Manual, 3rd Ed. (Sambrook et al., 2002)). al., Harbor Laboratory Press 2001), the entire contents of which are incorporated herein by reference for all purposes. The strand of target DNA that is complementary to and hybridizes to a Cas protein or gRNA may be referred to as the "complementary strand," and the strand of target DNA that is complementary to a "complementary strand" (and therefore not complementary to a Cas protein or gRNA) may be referred to as the "non-complementary strand" or "template strand."

[0212] Cas proteins can cleave nucleic acids at sites within or outside the nucleic acid sequence present within the target DNA to which the DNA-targeting segment of the gRNA binds. A "cleavage site" includes the location in the nucleic acid where the Cas protein generates a single-strand or double-strand break. For example, the formation of a CRISPR complex (including a gRNA hybridized to a guide RNA recognition sequence and complexed with a Cas protein) can result in cleavage of one or both strands within or near (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs from) the nucleic acid sequence present within the target DNA to which the DNA-targeting segment of the gRNA binds. If the cleavage site is outside the nucleic acid sequence to which the DNA-targeting segment of the gRNA binds, the cleavage site is still considered to be within the "guide RNA recognition sequence." The cleavage site may be on only one strand of the nucleic acid or on both strands. The cleavage sites can be at the same position on both strands of the nucleic acid (generating a blunt end) or at different sites on each strand (generating a cohesive end (i.e., overhang)). For example, cohesive ends can be generated using two Cas proteins, each of which generates a single-stranded break at a different cleavage site on a different strand, thereby generating a double-stranded break. For example, a first nickase can generate a single-stranded break on a first strand of double-stranded DNA (dsDNA), and a second nickase can generate a single-stranded break on a second strand of dsDNA, such that an overhang sequence is generated. In some examples, the guide RNA recognition sequence of the nickase on the first strand is separated from the guide RNA recognition sequence of the nickase on the second strand by at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 75, 100, 250, 500, or 1,000 base pairs.

[0213] Site-specific binding and cleavage of target DNA by the Cas protein can occur at a location determined by both (i) base pair complementarity between the gRNA and the target DNA and (ii) a short motif within the target DNA called a protospacer adjacent motif (PAM). The PAM may be adjacent to the guide RNA recognition sequence. Optionally, the PAM may be adjacent to the 3' end of the guide RNA recognition sequence. Alternatively, the PAM may be adjacent to the 5' end of the guide RNA recognition sequence. For example, the cleavage site of the Cas protein may be about 1 to about 10 or about 2 to about 5 base pairs (e.g., 3 base pairs) upstream or downstream of the PAM sequence. In some cases (e.g., when using Cas9 from S. pyogenes or a closely related Cas9), the PAM sequence on the non-complementary strand may be 5'-N1GG-3' (where N1 is any DNA nucleotide and is immediately 3' to the guide RNA recognition sequence on the non-complementary strand of the target DNA). Thus, the PAM sequence on the complementary strand will be 5'-CCN2-3' (N2 is any DNA nucleotide immediately 5' to the guide RNA recognition sequence on the complementary strand of the target DNA). In some such examples, N1 and N2 can be complementary, and the N1-N2 base pair can be any base pair (e.g., N1 = C and N2 = G; N1 = G and N2 = C; N1 = A and N2 = T; or N1 = T and N2 = A). In the example of Cas9 from S. aureus, the PAM can be NNGRRT (SEQ ID NO: 146) or NNGRR (SEQ ID NO: 147) (N can be A, G, C, or T, and R can be G or A). In the example of Cas9 from C. jejuni, the PAM can be, for example, NNNNACAC or NNNNRYAC (N can be A, G, C, or T, and R can be G or A). In some cases (eg, FnCpf1), the PAM sequence may be upstream of the 5' end and may have the sequence 5'-TTN-3'.

[0214] Examples of guide RNA recognition sequences include DNA sequences complementary to the DNA targeting segment of the gRNA, or such DNA sequences appended to the PAM sequence. For example, the target motif may be a preceding NGG motif, e.g., GN, that is recognized by the Cas9 protein. 19 NGG (SEQ ID NO: 1) or N 20 The guide RNA recognition sequence may be a 20-nucleotide DNA sequence immediately adjacent to the nucleotide sequence, such as NGG (SEQ ID NO: 2) (see, e.g., WO2014 / 165825, the entire contents of which are incorporated herein by reference for all purposes). The guanine at the 5' end may promote transcription by RNA polymerase in cells. Other examples of guide RNA recognition sequences include those with two guanine nucleotides at the 5' end (e.g., GGN 20 Examples of guide RNA recognition sequences include the nucleotide sequence GG (NGG; SEQ ID NO: 3). See, for example, WO2014 / 065596 (the entire contents of which are incorporated herein by reference for all purposes). Other guide RNA recognition sequences may have a length of 4 to 22 nucleotides of SEQ ID NOs: 1 to 3, including 5'G or GG and 3'GG or NGG. Still other guide RNA recognition sequences may have a length of 14 to 20 nucleotides of SEQ ID NOs: 1 to 3.

[0215] The guide RNA recognition sequence may be any nucleic acid sequence that is endogenous or exogenous to the cell. The guide RNA recognition sequence may be a sequence that encodes a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory sequence), or may include both.

[0216] C. Outpatient repair template The methods and compositions disclosed herein can utilize an exogenous repair template to modify a target genomic locus after cleaving the target genomic locus with a Cas protein. For example, the cell can be a one-cell embryo, and the exogenous repair template can be shorter than 5 kb in length. In cell types other than one-cell embryos, the exogenous repair template (e.g., a targeting vector) can be longer. For example, in cell types other than one-cell embryos, the exogenous repair template can be a large targeting vector (LTVEC) described elsewhere herein (e.g., a targeting vector having a length of at least 10 kb, or a targeting vector having 5' and 3' homology arms with a combined length of at least 10 kb). The use of an exogenous repair template in combination with a Cas protein can result in more precise modification of the target genomic locus by promoting homologous recombination repair.

[0217] In such methods, a Cas protein cleaves a target genomic locus to generate a single-strand break (nick) or a double-strand break, and an exogenous repair template recombines with the target nucleic acid via non-homologous end joining (NHEJ)-mediated ligation or a homology-directed repair event. Optionally, repair by the exogenous repair template removes or destroys the guide RNA recognition sequence or the Cas cleavage site so that the targeted allele cannot be retargeted by the Cas protein.

[0218] The exogenous repair template may comprise deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), may be single-stranded or double-stranded, and may be linear or circular. For example, the exogenous repair template may be a single-stranded oligodeoxynucleotide (ssODN). See, for example, Yoshimi et al. (2016) Nat. Commun. 7:10431 (the entire contents of which are incorporated by reference herein for all purposes). Exemplary exogenous repair templates are about 50 nucleotides to about 5 kb in length, about 50 nucleotides to about 3 kb in length, or about 50 to about 1,000 nucleotides in length. Other exemplary exogenous repair templates are about 40 to about 200 nucleotides in length. For example, the exogenous repair template may be about 50 to about 60, about 60 to about 70, about 70 to about 80, about 80 to about 90, about 90 to about 100, about 100 to about 110, about 110 to about 120, about 120 to about 130, about 130 to about 140, about 140 to about 150, about 150 to about 160, about 160 to about 170, about 170 to about 180, about 180 to about 190, or about 190 to about 200 nucleotides in length. Alternatively, the exogenous repair template may be about 50 to about 100, about 100 to about 200, about 200 to about 300, about 300 to about 400, about 400 to about 500, about 500 to about 600, about 600 to about 700, about 700 to about 800, about 800 to about 900, or about 900 to about 1,000 nucleotides in length. Alternatively, the exogenous repair template may be about 1 kb to about 1.5 kb, about 1.5 kb to about 2 kb, about 2 kb to about 2.5 kb, about 2.5 kb to about 3 kb, about 3 kb to about 3.5 kb, about 3.5 kb to about 4 kb, about 4 kb to about 4.5 kb, or about 4.5 kb to about 5 kb in length. Alternatively, the exogenous repair template may be, for example, 5 kb, 4.5 kb, 4 kb, 3.5 kb, 3 kb, 2.5 kb, 2 kb, 1.5 kb, 1 kb, 900 nucleotides, 800 nucleotides, 700 nucleotides, 600 nucleotides, 500 nucleotides, 400 nucleotides, 300 nucleotides, 200 nucleotides, 100 nucleotides, or 50 nucleotides or less in length. In cell types other than one-cell embryos, the exogenous repair template (e.g., targeting vector) may be longer.For example, in cell types other than one-cell stage embryos, the foreign repair template may be a large targeting vector (LTVEC) as described elsewhere herein.

[0219] In one example, the exogenous repair template is an ssODN about 80 nucleotides to about 200 nucleotides in length. In another example, the exogenous repair template is an ssODN about 80 nucleotides to about 3 kb in length. Such an ssODN may have, for example, homology arms each about 40 nucleotides to about 60 nucleotides in length. Such an ssODN may also have, for example, homology arms each about 30 nucleotides to 100 nucleotides in length. The homology arms may be symmetric (e.g., each 40 nucleotides in length or each 60 nucleotides in length) or asymmetric (e.g., one homology arm 36 nucleotides in length and one homology arm 91 nucleotides in length).

[0220] The foreign repair template may contain modifications or sequences that confer additional desirable characteristics (e.g., altered or regulated stability, tracking or detection by fluorescent labeling, binding sites for proteins or protein complexes, etc.). The foreign repair template may contain one or more fluorescent labels, purification tags, epitope tags, or combinations thereof. For example, the foreign repair template may contain one or more fluorescent labels (e.g., fluorescent proteins or other fluorophores or dyes), e.g., at least one, at least two, at least three, at least four, or at least five fluorescent labels. Exemplary fluorescent labels include fluorophores such as fluorescein (e.g., 6-carboxyfluorescein (6-FAM)), Texas Red, HEX, Cy3, Cy5, Cy5.5, Pacific Blue, 5-(and -6)-carboxytetramethylrhodamine (TAMRA), and Cy7. A wide range of fluorescent dyes for labeling oligonucleotides are commercially available (e.g., from Integrated DNA Technologies). For example, such fluorescent labels (e.g., internal fluorescent labels) can be used to detect exogenous repair templates directly incorporated into cleaved target nucleic acids with overhangs that match the ends of the exogenous repair templates. The labels or tags can be at the 5' end, 3' end, or internal within the exogenous repair templates. For example, the exogenous repair templates can be conjugated at the 5' end with IR700 fluorophore (5'IRDYE® 700) manufactured by Integrated DNA Technologies.

[0221] The foreign repair template may also include a nucleic acid insert comprising a segment of DNA to be integrated into the target genome locus. Introduction of the nucleic acid insert into the target genome locus can result in the addition of a nucleic acid sequence of interest to the target genome locus, the deletion of a nucleic acid sequence of interest at the target genome locus, or the replacement (i.e., deletion and insertion) of a nucleic acid sequence of interest at the target genome locus. Some foreign repair templates are designed to insert a nucleic acid insert into the target genome locus without any corresponding deletion at the target genome locus. Other foreign repair templates are designed to delete a nucleic acid sequence of interest at the target genome locus without any corresponding insertion of a nucleic acid insert. Still other foreign repair templates are designed to delete a nucleic acid sequence of interest at the target genome locus and replace it with a nucleic acid insert.

[0222] The nucleic acid insert, or the corresponding nucleic acid at the target genomic locus to be deleted and / or replaced, can be of various lengths. Exemplary nucleic acid inserts, or the corresponding nucleic acid at the target genomic locus to be deleted and / or replaced, are from about 1 nucleotide to about 5 kb in length, or from about 1 nucleotide to about 1,000 nucleotides in length. For example, the nucleic acid insert, or the corresponding nucleic acid at the target genomic locus to be deleted and / or substituted, may be about 1 to about 10, about 10 to about 20, about 20 to about 30, about 30 to about 40, about 40 to about 50, about 50 to about 60, about 60 to about 70, about 70 to about 80, about 80 to about 90, about 90 to about 100, about 100 to about 110, about 110 to about 120, about 120 to about 130, about 130 to about 140, about 140 to about 150, about 150 to about 160, about 160 to about 170, about 170 to about 180, about 180 to about 190, or about 190 to about 200 nucleotides in length. Similarly, the nucleic acid insert, or the corresponding nucleic acid in the target genomic locus to be deleted and / or replaced, may be about 1 to about 100, about 100 to about 200, about 200 to about 300, about 300 to about 400, about 400 to about 500, about 500 to about 600, about 600 to about 700, about 700 to about 800, about 800 to about 900, or about 900 to about 1,000 nucleotides in length. Similarly, the nucleic acid insert, or the corresponding nucleic acid in the target genomic locus to be deleted and / or replaced, may be from about 1 kb to about 1.5 kb, from about 1.5 kb to about 2 kb, from about 2 kb to about 2.5 kb, from about 2.5 kb to about 3 kb, from about 3 kb to about 3.5 kb, from about 3.5 kb to about 4 kb, from about 4 kb to about 4.5 kb, or from about 4.5 kb to about 5 kb in length.The nucleic acid to be deleted from the target genomic locus may also be about 1 kb to about 5 kb, about 5 kb to about 10 kb, about 10 kb to about 20 kb, about 20 kb to about 30 kb, about 30 kb to about 40 kb, about 40 kb to about 50 kb, about 50 kb to about 60 kb, about 60 kb to about 70 kb, about 70 kb to about 80 kb, about 80 kb to about 90 kb, about 90 kb to about 10 The length may be 0 kb, about 100 kb to about 200 kb, about 200 kb to about 300 kb, about 300 kb to about 400 kb, about 400 kb to about 500 kb, about 500 kb to about 600 kb, about 600 kb to about 700 kb, about 700 kb to about 800 kb, about 800 kb to about 900 kb, or about 900 kb to about 1 Mb, or it may be longer. Alternatively, the nucleic acid to be deleted from the target genomic locus may be about 1 Mb to about 1.5 Mb, about 1.5 Mb to about 2 Mb, about 2 Mb to about 2.5 Mb, about 2.5 Mb to about 3 Mb, about 3 Mb to about 4 Mb, about 4 Mb to about 5 Mb, about 5 Mb to about 10 Mb, about 10 Mb to about 20 Mb, about 20 Mb to about 30 Mb, about 30 Mb to about 40 Mb, about 40 Mb to about 50 Mb, about 50 Mb to about 60 Mb, about 60 Mb to about 70 Mb, about 70 Mb to about 80 Mb, about 80 Mb to about 90 Mb, or about 90 Mb to about 100 Mb.

[0223] The nucleic acid insert may comprise genomic DNA or any other type of DNA. For example, the nucleic acid insert may be derived from a prokaryote, a eukaryote, a yeast, a bird (e.g., a chicken), a non-human mammal, a rodent, a human, a rat, a mouse, a hamster, a rabbit, a pig, a cow, a deer, a sheep, a goat, a cat, a dog, a ferret, a primate (e.g., a marmoset, a rhesus monkey), a domestic mammal, an agricultural mammal, a turtle, or any other organism of interest.

[0224] The nucleic acid insert may contain a sequence homologous or orthologous to all or a portion of a gene encoding an autoantigen (e.g., a portion of a gene encoding a particular motif or region of an autoantigen). The homologous sequence may be from a different species or from the same species. For example, the nucleic acid insert may contain a sequence containing one or more point mutations (e.g., 1, 2, 3, 4, 5, or more) compared to the sequence targeted for substitution at the target genomic locus. Optionally, such point mutations may result in conservative amino acid substitutions (e.g., substitution of aspartic acid [Asp, D] for glutamic acid [Glu, E]) within the encoded polypeptide.

[0225] The nucleic acid insert, or the corresponding nucleic acid in the target genomic locus to be deleted and / or replaced, may be a coding region, such as an exon; a non-coding region, such as an intron, an untranslated region, or a regulatory region (e.g., a promoter, enhancer, or transcriptional repressor binding element); or any combination thereof.

[0226] The nucleic acid insert may also contain a conditional allele. The conditional allele may be a multifunctional allele, as described in US2011 / 0104799 (the entirety of which is incorporated herein by reference for all purposes). For example, the conditional allele may contain (a) an operating sequence in a sense orientation relative to transcription of the target gene, (b) a drug selection cassette (DSC) in a sense or antisense orientation, (c) a nucleotide sequence of interest (NSI) in an antisense orientation, and (d) a conditional by inversion module (utilizing COIN, exon split intron, and inverted gene trap-like modules) in an inverted orientation. See, e.g., US2011 / 0104799. The conditional allele may further contain a recombinase unit that recombines upon exposure to a first recombinase to form a conditional allele that (i) lacks the operating sequence and DSC and (ii) contains the NSI in a sense orientation and a COIN in an antisense orientation. See, for example, US2011 / 0104799.

[0227] The nucleic acid insert may also include a polynucleotide encoding a selection marker. Alternatively, the nucleic acid insert may lack a polynucleotide encoding a selection marker. The selection marker may be included in a selection cassette. Optionally, the selection cassette may be a self-deletion cassette. See, e.g., US8,697,851 and US2013 / 0312129 (the entire contents of each are incorporated by reference herein for all purposes). As an example, the self-deletion cassette may include a Crei gene (comprising two exons encoding Cre recombinase, separated by an intron) operably linked to a mouse Prm1 promoter and a neomycin resistance gene operably linked to a human ubiquitin promoter. By using the Prm1 promoter, the self-deletion cassette can be deleted, particularly in male germ cells of F0 animals. Exemplary selection markers include neomycin phosphotransferase (neomycin phosphotransferase) and the like. r ), hygromycin B phosphotransferase (hygr ), puromycin N-acetyltransferase (puro r ), blasticidin S deaminase (bsr r ), xanthine / guanine phosphoribosyltransferase (gpt), or herpes simplex virus thymidine kinase (HSV-k), or a combination thereof. The polynucleotide encoding the selectable marker may be operably linked to a promoter that is active in the target cells. Examples of promoters are described elsewhere herein.

[0228] The nucleic acid insert may also contain a reporter gene. Exemplary reporter genes include those encoding luciferase, β-galactosidase, green fluorescent protein (GFP), enhanced green fluorescent protein (eGFP), cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), enhanced yellow fluorescent protein (eYFP), blue fluorescent protein (BFP), enhanced blue fluorescent protein (eBFP), DsRed, ZsGreen, MmGFP, mPlum, mCherry, tdTomato, mStrawberry, J-Red, mOrange, mKO, mCitrine, Venus, YPet, Emerald, CyPet, Cerulean, T-Sapphire, and alkaline phosphatase. Such reporter genes may be operably linked to a promoter that is active in the target cells. Examples of promoters are described elsewhere herein.

[0229] The nucleic acid insert may also contain one or more expression or deletion cassettes. Any cassette may contain one or more of a nucleotide sequence of interest, a polynucleotide encoding a selectable marker, and a reporter gene, along with various regulatory components that affect expression. Examples of selectable markers and reporter genes that may be included are described in detail elsewhere herein.

[0230] The nucleic acid insert may contain nucleic acids flanked by site-specific recombination target sequences. Alternatively, the nucleic acid insert may contain one or more site-specific recombination target sequences. The entire nucleic acid insert may be flanked by such site-specific recombination target sequences, but any region or individual polynucleotide of interest within the nucleic acid insert may also be flanked by such sites. Site-specific recombination target sequences that may flank the nucleic acid insert or any polynucleotide of interest within the nucleic acid insert may include, for example, loxP, lox511, lox2272, lox66, lox71, loxM2, lox5171, FRT, FRT11, FRT71, attp, att, FRT, rox, or combinations thereof. In one example, the site-specific recombination sites flank a polynucleotide encoding a selectable marker and / or a reporter gene contained within the nucleic acid insert. After the nucleic acid insert is introduced into the target locus, the sequence between the site-specific recombination sites may be removed. Optionally, two exogenous repair templates, each with a nucleic acid insert containing a site-specific recombination site, may be used. The exogenous repair template may target th...

Claims

[Claim 1] The invention as described in the drawings.