Mammalian cells containing an integrated Cas9 gene to generate a stable integration site, and mammalian cells containing a stable integration site and other sites

JP2024539076A5Pending Publication Date: 2025-10-27REGENERON PHARMACEUTICALS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024523257
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-10-18
Filing Date
2022-10-18
Publication Date
2025-10-27

AI Technical Summary

Technical Problem

Mammalian cell lines used for producing therapeutic proteins face issues of reduced production due to genetic and epigenetic instability, particularly when integrating polynucleotides into genomic safe harbors like AAVS1, which can lead to instability and inefficiencies in DNA integration.

Method used

The development of mammalian cells with multiple stable integration sites, including one in a genomic safe harbor (such as AAVS1) and another outside it, utilizing integrated Cas9 genes to enhance integration efficiency through recombinase-mediated cassette exchange (RMCE), allowing for the stable incorporation of DNA cassettes and polynucleotides of interest.

Benefits of technology

This approach significantly enhances the stability and efficiency of DNA integration, leading to improved production of proteins of interest, with up to 10 times higher homologous directed repair (HDR) efficiency compared to conventional methods, and supports the production of multiple proteins or genes within the same cell line.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000099_0000
    Figure 00000099_0000
  • Figure 00000099_0001
    Figure 00000099_0001
  • Figure 00000100_0000
    Figure 00000100_0000
Patent Text Reader

Abstract

The present invention provides mammalian cells comprising multiple stable integration sites. The present invention provides sites genomically introduced into a genomic safe harbor and sites genomically introduced outside that particular genomic safe harbor (including but not limited to another genomic safe harbor). A polynucleotide of interest encoding a polypeptide or RNA of interest can be inserted into a stable integration site provided according to the present invention. The cells and methods of the present invention can be used for high yield production of any protein, including viral proteins. Additionally, the cells and methods of the present invention are useful for the production of viral vectors such as AAV, antibodies, and other proteins.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority to U.S. Application No. 63 / 256,675, filed October 18, 2021, which is incorporated herein by reference in its entirety. The present invention provides mammalian cells (including cell lines), including human and rodent cells (including cell lines), that contain multiple stable integration sites (SISs) that can be produced using an integrated Cas9 gene. The present invention provides stable integration sites (1) genomically introduced into genomic safe harbors (GSHs), such as AAVS1 (adeno-associated virus integration site 1) and AAVS1-like sites, and (2) genomically introduced outside of that particular genomic safe harbor, such as a different genomic safe harbor or another region that is not a genomic safe harbor. A polydeoxyribonucleotide of interest encoding a polypeptide or RNA of interest can be inserted into the stable integration site provided in accordance with the present invention.

[0002] Reference to Electronic Sequence Listing This application contains a Sequence Listing that has been submitted electronically in .XML format and is incorporated herein by reference in its entirety. The .XML copy, created on October 7, 2022, is named "135975-97402.xml" and is 709,205 bytes in size. The Sequence Listing contained in this .XML file is made a part of the present specification and is incorporated herein by reference in its entirety. [Background technology]

[0003] Mammalian cell lines are a preferred approach for producing commercial quantities of therapeutic proteins, such as antibodies. However, engineered mammalian cells have been reported to often exhibit reduced production due to genetic and epigenetic instability. (Hilliard and Lee, Biotech. Bioeng. 118:659-75 (2021))

[0004] Polynucleotide integration is a preferred approach for generating and maintaining transformed cells. Integration of specific sequences into human AAVS1 is discussed in Liu et al., BMC Research Note, 7:626 (2014) and Ramachandra et al., Nucl. Acids Res. 39:e107 (2011). Human AAVS1 is known as a genomic safe harbor. Papapetrou et al., Molecular Therapy 24:678-84 (2016). Gaidukov et al., Nucl. Acids Res. 46:4072-86 (2018) discloses a site for DNA integration into the landing pad.

[0005] Chinese hamster ovary (CHO) cells and baby hamster kidney (BHK) cells are used to produce therapeutic proteins, and the hamster genome has been extensively tested. Hamaker and Lee reported CHO chromosomal loci as potential sites for stable integration, calling them "genomic hotspots." Curr. Op. Chem. Eng. 22:152-60 (2018) p. 153. In Table 1, Hamaker and Lee identified 30 hotspot loci, 17 of which were identified by genes and 13 were unannotated. Curr. Op. Chem. Eng. 22:152-60 (2018) p. 154. This work was continued by Hilliard and Lee, who attempted to identify safe harbor regions in CHO using epigenomic analysis. Hilliard and Lee, Biotech. Bioeng. 118:659-75 (2021). The authors determined that 10.9% of the CHO genome contained chromatin structures with enhanced genetic and epigenetic stability. They further determined that five of the 30 hotspots identified by Hamaker and Lee in Table 1 overlapped with stable regions determined by high-throughput chromosome conformation capture (Hi-C). The genes closest to the regions were ALDH5A1, SMAD6, and CLCN3; the other two regions were unannotated. Table 3 (S3) in Hilliard and Lee, Biotech. Bioeng. 118:659-75 (2021) also identifies genetic loci for integration in CHO cells. Lee et al., Scientific Reps. 5:8572 (2015) identifies the COSMC locus.

[0006] The present invention advantageously uses an integrated Cas9 gene to efficiently generate mammalian cell intermediates that are further modified to provide mammalian cells with multiple stable integration sites for stable integration of multiple DNA cassettes and other polydeoxyribonucleotides of interest. According to the present invention, the stable integration sites can be located in genomic safe harbors or other regions that include newly identified genomic safe harbors. Summary of the Invention

[0007] The present invention provides mammalian cells, any of which can comprise a first stable integration site located in a genomic safe harbor and a second stable integration site not located in a genomic safe harbor, wherein the first stable integration site comprises a first reporter gene encoding a first reporter protein, and the second stable integration site comprises a second reporter gene encoding a second reporter protein, wherein the first and second reporter proteins are different. The first and second stable integration sites can comprise recombinase recognition sites (RRS). The first and second reporter genes can be under the control of an SV40 promoter. The first and second reporter genes can be fluorescent proteins. The cells can further comprise a polynucleotide encoding a repressor protein under the control of a CMV promoter. The cells can be human amniotic epithelial cells, HEK293 cells, CHO cells, or BHK cells. A polynucleotide encoding a protein of interest can be inserted into the first stable integration site or the second stable integration site. The second stable integration site can be located in a second genomic safe harbor that is different from the first genomic safe harbor or in a region that is not a genomic safe harbor.

[0008] The present invention also provides mammalian cells, any of which can include a first stable integration site located in a genomic safe harbor and a second stable integration site not located in the first genomic safe harbor, where the first stable integration site includes a first polynucleotide encoding a first protein and the second stable integration site includes a second polynucleotide encoding a second protein. The first and second proteins can be viral proteins, such as adenovirus-associated virus proteins or adenovirus proteins. For example, the mammalian cell can include a polynucleotide encoding an adeno-associated virus protein and a polynucleotide encoding an adenovirus protein. Other polynucleotides encoding proteins include, but are not limited to, antibody genes. The cell can have a second stable integration site located in a second genomic safe harbor different from the first genomic safe harbor in which the first stable integration site is located, or in a region that is not a genomic safe harbor.

[0009] The present invention further provides mammalian cells, any of which can comprise a first stable integration site located in a genomic safe harbor and a second stable integration site not located in a genomic safe harbor, wherein the first stable integration site comprises a polynucleotide encoding a first reporter gene encoding a first reporter protein, and the second stable integration site comprises a polynucleotide encoding Cas9 and a polynucleotide encoding a second reporter gene encoding a second reporter protein, wherein the first reporter protein and the second reporter protein are different. The second stable integration site can further comprise a selectable marker gene and an internal ribosome entry site (IRES). The first and second stable integration sites can comprise recombinase recognition sites. The first and second reporter genes can be under the control of an SV40 promoter. The first and second reporter genes can be fluorescent proteins. The cell can further comprise a polynucleotide encoding a repressor (e.g., TetR) under the control of a promoter (e.g., CMV). The cells can be human amniotic epithelial cells, HEK293 cells, CHO cells, or BHK cells. A polynucleotide encoding a protein of interest can be inserted into the first stable integration site or the second stable integration site. The selectable marker protein can confer drug resistance. The second reporter gene, selectable marker gene, IRES, and SV40 promoter can be arranged in the DNA cassette. The cells can further comprise a polynucleotide encoding a repressor protein under the control of a promoter (e.g., CMV). The second stable integration site can be located in a second genomic safe harbor different from the first genomic safe harbor in which the first stable integration site is located, or in a region that is not a genomic safe harbor. The first reporter gene can be flanked by a 5' genomic safe harbor homology arm and a 3' genomic safe harbor homology arm.The 5' genomic safe harbor homology arm can comprise a CRISPR sgRNA target site, and the 3' genomic safe harbor homology arm can comprise a CRISPR sgRNA target site.

[0010] The present invention further provides methods for producing at least one protein of interest, any of which may include: (a) providing a mammalian cell containing a first stable integration site located in a genomic safe harbor and a second stable integration site not located in the first genomic safe harbor, wherein the first stable integration site contains a first reporter gene encoding a first reporter protein, and the second stable integration site contains a second reporter gene encoding a second reporter protein, the first and second reporter proteins being distinct, and the first and second stable integration sites containing recombinase recognition sites; (b) introducing a polynucleotide encoding the protein of interest into the stable integration site by recombinase-mediated cassette exchange; and (c) culturing the mammalian cell under conditions allowing expression of the polynucleotide encoding the polynucleotide of interest. The first and second reporter genes may be under the control of an SV40 promoter. The first and second reporter genes may be fluorescent proteins. The cell may further contain a polynucleotide encoding a repressor protein under the control of a CMV promoter. The cells can be human amniotic epithelial cells, HEK293 cells, CHO cells, or BHK cells. The polynucleotide encoding the protein of interest can be inserted into a first stable integration site or a second stable integration site. The second stable integration site can be located in a second genomic safe harbor different from the first genomic safe harbor in which the first stable integration site is located, or in a region that is not a genomic safe harbor. The first stable integration site contains a first polynucleotide encoding a first protein, and the second stable integration site contains a second polynucleotide encoding a second protein. The first and second proteins can be viral proteins, such as adenovirus-associated virus proteins or adenovirus proteins.For example, the mammalian cell can contain a polynucleotide encoding an adeno-associated virus protein and a polynucleotide encoding an adenovirus protein. Other polynucleotides encoding proteins include, but are not limited to, antibody genes. The second stable integration site can also be located in a region that is not a genomic safe harbor.

[0011] The present invention further provides methods for generating mammalian cells with multiple stable integration sites, any of the methods including: (A) providing a mammalian cell containing a first DNA cassette comprising, in 5' to 3' order, a first lox site, a promoter, a selectable marker gene encoding a selectable marker protein, an IRES, a first reporter gene encoding a first reporter protein, a promoter operably linked to an operator, a Cas9 gene, and a polynucleotide encoding a second lox site; and (B) providing a mammalian cell containing, in 5' to 3' order, a first genomic safe harbor homology arm comprising a CRISPR sgRNA target site, a third lox site, a second reporter gene encoding a second reporter protein, a fourth lox site, and a CRISPR sgRNA target site. (C) replacing the first DNA cassette with a third DNA cassette comprising, in 5' to 3' order, the first lox site, a promoter, a third reporter gene encoding a third reporter protein, and a polynucleotide encoding a second lox site, thereby providing a mammalian cell with multiple stable integration sites. The mammalian cells can be human amniotic epithelial cells, HEK293 cells, CHO cells, or BHK cells. The reporter gene used can be a fluorescent protein. The cells of step (A) can further comprise a polynucleotide encoding a repressor (e.g., TetR) under the control of a promoter (e.g., CMV).The cells of step (B) can further comprise a polynucleotide encoding a repressor (e.g., TetR) under the control of a promoter (e.g., CMV). The cells of step (C) can further comprise a polynucleotide encoding a repressor (e.g., TetR) under the control of a promoter (e.g., CMV). The selectable marker protein can confer drug resistance. Lox sites are the most commonly used type of RRS, although different RRSs can be used as well.

[0012] The present invention also provides methods for producing a mammalian cell having multiple recombinase-mediated cassette exchange sites, any of the methods comprising: (A) randomly integrating into the cellular genome a polynucleotide encoding a promoter and a repressor, wherein the repressor is capable of binding to a ligand; (B) randomly integrating into the cellular genome a first DNA cassette comprising, in 5' to 3' order, a first lox site, a promoter and optionally an operator, a first reporter gene encoding a first reporter protein, an IRES, a first selectable marker gene encoding a first selectable marker protein, and a polynucleotide encoding a second lox site, wherein the first lox site and the second lox site are different; and (C) exchanging the first DNA cassette with a second DNA cassette, wherein the second DNA cassette comprises, in 5' to 3' order, a first lox site, a promoter, a second selectable marker gene encoding a second selectable marker protein, an IRES, and a second reporter gene encoding a second reporter protein. (D) a polynucleotide encoding a gene, a promoter and optionally an operator, a Cas9 gene, and a second lox site, wherein the first selection marker protein and the second selection marker protein are different, and the first reporter protein and the second reporter protein are different, and (E) a polynucleotide encoding, in order from 5' to 3', a first genomic safe harbor (GSH) homology arm including an sgRNA (single guide RNA) target site, a third lox site, a third reporter gene encoding a third reporter protein, a fourth lox site, a fourth reporter gene encoding a fourth reporter protein, a fourth reporter gene encoding a fourth reporter protein, a fifth ... incorporating a third DNA cassette comprising a polynucleotide comprising a GSH homology arm comprising a first lox site, a second lox site, a third lox site, and a fourth lox site, wherein the first lox site, the second lox site, the third lox site, and the fourth lox site are different, the first and second guide arms can comprise at least one region with an alteration (optionally to avoid re-creation of a targetable site), and the third reporter protein is different from the second reporter protein and can be the same as or different from the first reporter protein;(E) Replacing the second DNA cassette with a fourth DNA cassette, the fourth DNA cassette comprising, in 5' to 3' order, a first lox site, a promoter, a fourth reporter gene encoding a fourth reporter protein, and a polynucleotide encoding a second lox site, where the fourth reporter protein is different from the third reporter protein and the second reporter protein, and preferably different from the first reporter protein, thereby providing cells with multiple stable integration sites. Lox sites are the most commonly used type of RRS, although different RRSs can be used as well.

[0013] The present invention further provides mammalian cells comprising a modified genome, wherein the given genome is modified by insertion of at least three DNA cassettes within different regions of the genome, and the modified genome has at least 50% to 99%, 75% to 99%, 85% to 99%, 90% to 99%, 95% to 98%, 98% to 99%, 50%, 55%, 60%, 65% of a sequence selected from the group consisting of SEQ ID NOs: 1 and 2 prior to modification. , 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to at least one of SEQ ID NOs: 5 to 10, and (2) a first deoxyribonucleic acid sequence that is at least 50% to 99%, 75% to 99%, 85% to 99%, 90% to 99%, 95% to 98%, 98% to 99%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to at least one of SEQ ID NOs: 5 to 10, before modification. , 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to at least one of SEQ ID NO: 11 and SEQ ID NO: 12, before modification; and (3) a second deoxyribonucleic acid sequence that is at least 50% to 99%, 75% to 99%, 85% to 99%, 90% to 99%, 95% to 98%, 98% to 99%, 50%, 55%, 60%, 65%, or identical to at least one of SEQ ID NO: 11 and SEQ ID NO: 12, before modification. and a third deoxyribonucleic acid that is 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the first deoxyribonucleic acid sequence, wherein the first deoxyribonucleic acid sequence is modified by the insertion of a first DNA cassette, the second deoxyribonucleic acid sequence is modified by the insertion of a second DNA cassette, and the third deoxyribonucleic acid sequence is modified by the insertion of a third DNA cassette. The mammalian cells can each have (a) a first DNA cassette comprising a promoter and at least one selected from the group consisting of a selectable marker gene and a reporter gene, (b) a second DNA cassette comprising a promoter and at least one selected from the group consisting of a selectable marker gene and a reporter gene, and (c) a third DNA cassette comprising a promoter and at least one selected from the group consisting of a selectable marker gene and a reporter gene.Furthermore, the mammalian cells can each have (a) a first DNA cassette containing a promoter, a selectable marker gene, and a reporter gene, (b) a second DNA cassette containing a promoter, a selectable marker gene, and a reporter gene, and (c) a third DNA cassette containing a promoter, a selectable marker gene, and a reporter gene. The first deoxyribonucleic acid sequence contains a stable integration site and a gene of interest inserted therein. The gene of interest can encode a polypeptide of interest selected from the group consisting of an antibody, an antibody chain, a receptor, an Fc-containing protein, a trap protein, an enzyme, a factor, an inhibitor, an activator, a ligand, a reporter protein, a selection protein, a protein hormone, a protein toxin, a structural protein, a storage protein, a transport protein, a neurotransmitter, and a contractile protein. The mammalian cell can be a human cell, and the first deoxyribonucleic acid sequence is at least 50% to 99%, 75% to 99%, 85% to 99%, 90% to 99%, 95% to 98%, 98% to 99%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:1. Alternatively, the mammalian cell can be a CHO cell, and the first deoxyribonucleic acid sequence is at least 50% to 99%, 75% to 99%, 85% to 99%, 90% to 99%, 95% to 98%, 98% to 99%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 2. The first deoxyribonucleic acid sequence can comprise a stable integration site produced using a guide sequence selected from the group consisting of SEQ ID NOs: 13-419.Additionally, the first deoxyribonucleic acid sequence binds to and / or is complementary to the target sequence of SEQ ID NO: 2, and the stable integration site produced by using the guide sequence is (a) 1 to 2000, (b) 2001 to 4000, (c) 4001 to 6000, (d) 6001 to 8000, (e) 8001 to 10,000, (f) 10,001 to 12,000, (g) 12,001 to 14,000, (h) 14,001 to 16,000, (i) 16,001 to 18,000, (j) 18,001 to 20,000, (k) 20 (r) 34,001 to 36,000, (s) 36,001 to 38,000, (t) 38,001 to 40,000, (u) 40,001 to 42,000, and (v) 42,001 to 44,232.

[0014] Additionally provided is a mammalian cell comprising a modified genome, the modified genome comprising a deoxyribonucleic acid sequence comprising an AAVS1-like region modified by insertion of at least one DNA cassette, wherein the guide sequence is selected from the group consisting of SEQ ID NOs: 13-419, which binds to and / or is complementary to the sense or antisense strand of the AAVS1-like region. The mammalian cell is further characterized by the presence of a second deoxyribonucleic acid sequence that is at least 50% to 99%, 75% to 99%, 85% to 99%, 90% to 99%, 95% to 98%, 98% to 99%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to at least one selected from the group consisting of SEQ ID NOs: 5 to 10 before modification, and a second deoxyribonucleic acid sequence that is at least 50% to 99%, 75% to 99%, 85% to 99%, 90% to 99%, 95% to 98%, 98% to 99%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to at least one selected from the group consisting of SEQ ID NOs: 11 and 12 before modification. and a third deoxyribonucleic acid that is 75% to 99%, 85% to 99%, 90% to 99%, 95% to 98%, 98% to 99%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the first deoxyribonucleic acid sequence, wherein the first deoxyribonucleic acid sequence is modified by insertion of a first DNA cassette, the second deoxyribonucleic acid sequence is modified by insertion of a second DNA cassette, and the third deoxyribonucleic acid sequence is modified by insertion of a third DNA cassette.The second deoxyribonucleic acid sequence, before modification, has a similarity to at least one selected from the group consisting of at least one selected from the group consisting of SEQ ID NOs: 5 to 10, of at least 50% to 99%, 75% to 99%, 85% to 99%, 90% to 99%, 95% to 98%, 98% to 99%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of at least one selected from the group consisting of SEQ ID NOs: 5 to 10. 9% identical to at least one selected from the group consisting of SEQ ID NO:11 and SEQ ID NO:12, and the third deoxyribonucleic acid is at least 50% to 99%, 75% to 99%, 85% to 99%, 90% to 99%, 95% to 98%, 98% to 99%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to at least one selected from the group consisting of SEQ ID NO:11 and SEQ ID NO:12, before modification.The first deoxyribonucleic acid sequence has at least 50% to 99%, 75% to 99%, 85% to 99%, 90% to 99%, 95% to 98%, 98% to 99%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO:2. Stable integration sites, generated by using guide sequences that bind to and / or are complementary to at least one target sequence, are located at nucleotide positions: (a) 1 to 2000, or (b) 2001 to 4000, or (c) 4001 to 6000, or (d) 6001 to 8000, or (e) 8001 to 10,000, or (f) 10,001 to 12 ,000, or (g) 12,001 to 14,000, or (h) 14,001 to 16,000, or (i) 16,001 to 18,000, or (j) 18,001 to 20,000, or (k) 20,001 to 22,000, or (l) 22,001 to 24,000, or (m) 24,001 to 26,000, or (n) 26,001 to 28,000 , or (o) 28,001 to 30,000, or (p) 30,001 to 32,000, or (q) 32,001 to 34,000, or (r) 34,001 to 36,000, or (s) 36,001 to 38,000, or (t) 38,001 to 40,000, or (u) 40,001 to 42,000, or (v) 42,001 to 44,232, inclusive.

[0015] Also provided is a mammalian cell comprising a modified genome, the modified genome comprising a stable integration site in an AAVS1-like region, the stable integration site being located between nucleotide positions: (a) 1 and 2000, or (b) 2001 and 4000, or (c) 4001 and 6000, or (d) 6001 and 8000, or (e) 8001 and 10,000, by using a guide sequence that binds to and / or is complementary to at least one target sequence having at least 50% to 99%, 75% to 99%, 85% to 99%, 90% to 99%, 95% to 98%, 98% to 99%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO:2. , or (f) 10,001 to 12,000, or (g) 12,001 to 14,000, or (h) 14,001 to 16,000, or (i) 16,001 to 18,000, or (j) 18,001 to 20,000, or (k) 20,001 to 22,000, or (l) 22,001 to 24,000, or (m) 24,001 to 26,000, or (n) 26,001 to 28,000, or (o) 28,001 to 30,000, or (p) 30,001 to 32,000, or (q) 32,001 to 34,000, or (r) 34,001 to 36,000, or (s) 36,001 to 38,000, or (t) 38,001 to 40,000, or (u) 40,001 to 42,000, or (v) 42,001 to 44,232.

[0016] Further provided is a mammalian cell according to the preceding paragraph, comprising a second deoxyribonucleic acid sequence that, prior to modification, is at least 50% to 99%, 75% to 99%, 85% to 99%, 90% to 99%, 95% to 98%, 98% to 99%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to at least one selected from the group consisting of SEQ ID NOs: 5-10; and a second deoxyribonucleic acid sequence that, prior to modification, is at least 50% to 99%, 75% to 99%, 85% to 99%, 90% to 99%, 95% to 98%, 98% to 99%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to at least one selected from the group consisting of SEQ ID NOs: 11 and 12. and a third deoxyribonucleic acid that is 0% to 99%, 75% to 99%, 85% to 99%, 90% to 99%, 95% to 98%, 98% to 99%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the first deoxyribonucleic acid sequence, wherein the first deoxyribonucleic acid sequence is modified by insertion of a first DNA cassette, the second deoxyribonucleic acid sequence is modified by insertion of a second DNA cassette, and the third deoxyribonucleic acid sequence is modified by insertion of a third DNA cassette. The mammalian cell, prior to modification, contains a second deoxyribonucleic acid sequence that is at least 50% to 99%, 75% to 99%, 85% to 99%, 90% to 99%, 95% to 98%, 98% to 99%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to at least one selected from the group consisting of at least one selected from the group consisting of SEQ ID NOs: 5 to 10. and a third deoxyribonucleic acid that, prior to modification, is at least 50% to 99%, 75% to 99%, 85% to 99%, 90% to 99%, 95% to 98%, 98% to 99%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to at least one selected from the group consisting of SEQ ID NO:11 and SEQ ID NO:12.

[0017] Additionally, methods of producing a protein of interest are provided, comprising: (1) culturing the mammalian cells described above; and (2) harvesting the protein of interest. Also provided are cells produced according to any of the above methods, as well as methods of using the disclosed cells.

[0018] The following diagrams illustrate exemplary progression and generation of intermediate cells useful for generating cells with stable integration sites in different regions of the genome, and subsequently generating cells with stable integration sites in different regions of the genome. These diagrams illustrate embodiments of the present invention and do not limit the invention in any way. [Brief explanation of the drawings]

[0019] [Figure 1] 1 shows a schematic representation of the engineering of a cell with a polynucleotide encoding a repressor protein under the transcriptional control of a promoter and a polyadenylation signal, where the polynucleotide is randomly inserted into the cell genome.

[0020] [Figure 2] Figure 1 shows a schematic representation of a cell modified after a DNA cassette (1) has been inserted randomly or site-specifically into the cell genome. The DNA cassette (1) contains flanking lox sites (1 and 2), a promoter, a reporter gene (1), an IRES, a selectable marker gene (1), and a polyadenylation signal. Instead of the lox sites, other RRSs can be used as well.

[0021] [Figure 3]2 shows a schematic representation of the modification of the cell in which DNA cassette (1) is replaced with DNA cassette (2) by recombinase-mediated cassette exchange. DNA cassette (2) contains flanking lox sites (1 and 2), a promoter, a selectable marker gene (2), an IRES, and a reporter gene (2), and a polyadenylation signal, as well as a Cas9 gene with a second polyadenylation signal under the control of a second promoter (an operator is optional). Other RRSs can be used in place of the lox sites as well.

[0022] [Figure 4] Figure 3 shows a schematic representation of the modification of the cell with a DNA cassette (3) containing a reporter gene (3) under the control of a promoter inserted into the genomic safe harbor (GSH) with flanking genomic safe harbor (GSH) homology arms, lox sites (3 and 4), and a polyadenylation signal. The insertion is site-specific, creating a stable integration site between Lox3 and Lox4. Instead of lox sites, other RRSs can be used as well.

[0023] [Figure 5] Schematic representation of the cell modification in Figure 4, in which DNA cassette (2) is replaced with DNA cassette (4) by recombinase-mediated cassette exchange. DNA cassette (4) contains flanking lox sites (1 and 2), a reporter gene (4), and a polyadenylation signal under the control of a promoter. This exchange removes the Cas9 gene. Other RRSs can be used in place of the lox sites as well.

[0024] [Figure 6] Schematic representation of the sgRNA plasmids used in Example 6.

[0025] [Figure 7]1 shows the plot from Example 6 showing the green fluorescent protein positive population (Q1) for no HDR template (control), 104-mer HDR template, 401-mer HDR template, and 1030-mer HDR template. GFP positivity is on the vertical axis and CFP positivity is on the horizontal axis.

[0026] [Figure 8] A mammalian cell (e.g., HEK293) with a stably integrated Cas9 gene flanked by Lox sites 3 and 4 is shown schematically. The Cas9 gene is under the control of at least one promoter (not shown). AAVS1 is also shown schematically. Other RRSs can be used in place of the lox sites as well. A promoter is present 5' of the gene but is not shown.

[0027] [Figure 9A] The targeting plasmid is shown schematically, including an sgRNA target site, a left homology arm (here, a GSH homology arm) for insertion into a region such as a genomic safe harbor (here, AAVS1), a Lox1 site, a reporter gene (color 1), a Lox2 site, and a right homology arm (here, a GSH homology arm) for insertion into a region such as a genomic safe harbor (here, AAVS1). At the 3' end, A schematically indicates a reporter gene (color 2). Promoters and optionally other portions (such as operators) are represented by arrows pointing in the 5' to 3' direction. Both plasmids insert color 1 into regions such as a genomic safe harbor. Other RRSs can be used in place of the lox site. [Figure 9B]Figure 1 shows a schematic representation of a targeting plasmid containing an sgRNA target site, a left homology arm (here, a GSH homology arm) for insertion into a region such as a genomic safe harbor (here, AAVS1), a Lox1 site, a reporter gene (color 1), a Lox2 site, and a right homology arm (here, a GSH homology arm) for insertion into a region such as a genomic safe harbor (here, AAVS1). At the 3' end, B shows a schematic representation of a negative selection gene (negative selection 1) at the 3' end. Promoters and, optionally, other portions (such as operators) are represented by arrows pointing in the 5' to 3' direction. Both plasmids insert color 1 into regions such as a genomic safe harbor. Instead of lox sites, other RRSs can be used as well.

[0028] [Figure 10] Figure 9B shows a schematic representation of the results after Cas9-mediated integration into a genome safe harbor (AAVS1) of mammalian cells (e.g., HEK293). Color 1 is flanked by Lox1 and Lox2. A gene of interest can replace Color 1 via RMCE. When a targeting plasmid according to Figure 9A is properly integrated, the cells will be Color 1 positive and Color 2 negative. When a targeting plasmid according to Figure 9B is properly integrated, the cells will be Color 1 positive and will be able to grow due to the removal of the negative selection gene. Instead of lox sites, other RRSs can be used as well. The promoter and optionally other moieties (such as operators) are represented by arrows pointing in the 5' to 3' direction.

[0029] [Figure 11]The insertion in Figure 10 is shown in more detail. The cellular genome containing AAVS1 is flanked by the insert and the 5' and 3' ends. Color 1 is flanked by Lox1 and Lox2. Figure 11, left, shows the locations of the 5' genomic primer and 3' insertion primer used in 5' junction PCR. Figure 11, right, shows the locations of the 5' insertion primer and 3' genomic primer used in 3' junction PCR. Instead of lox sites, other RRSs can be used as well. The promoter 5' of the Color 1 gene is shown as a 5' arrow.

[0030] [Figure 12] It is shown that a fragment of the correct size is amplified in HEK293 cells by junction PCR, as shown schematically in Figure 11. Stable Cas9-targeted HEK293 cells and 5' and 3' junctions are obtained and detected, establishing the correct insertion.

[0031] [Figure 13] A fragment of the correct size is shown to be amplified in CHO cells by junction PCR, as shown schematically in Figure 11. Stable Cas9-targeted CHO cells and 5' and 3' junctions are obtained and detected, confirming the correct insertion. Instead of lox sites, other RRSs can be used as well.

[0032] [Figure 14]An exemplary cell is shown schematically containing three cassettes integrated into regions of the genome with adjacent RRSs (here, lox1 and lox2). Depending on the cell type, each of the three cassettes can be integrated into a different stable integration site (e.g., AAVS1-like), shown schematically at position A, as well as other available sites (e.g., stable site 1 and stable site 2), shown schematically at positions B and C. The reporter genes can be the same or different. The negative selection genes can be the same or different, but are preferably the same. The cell can include additional stable integration sites and integrated cassettes in accordance with the teachings contained herein. A promoter is present 5' of the gene but is not shown.

[0033] [Figure 15] Modifications of the cell of Figure 14 are shown schematically at positions A, B, and C. Three cassettes each contain flanking RRSs (here, lox1 and lox2), a gene of interest, a positive selection marker gene, and a reporter* gene. The positive selection marker genes can be the same or different, but are preferably the same. The reporter* genes can be the same or different, but each must be different from any of the reporter genes in the cell of Figure 14. The genes of interest can be the same or different. The cassette of Figure 14 is replaced with the cassette of Figure 15 by RMCE. The cell can include additional stable integration sites and integrated cassettes according to the teachings contained herein. A promoter is present 5' of the gene but is not shown.

[0034] [Figure 16] 1 is a bar graph comparing proteins produced by 3-site CHO-K1 cells (A, B, and C) compared to 2-site CHO-K1 cells (B and C).

[0035] [Figure 17]An exemplary cell is shown schematically containing four cassettes integrated into regions of the genome with adjacent RRSs (here, lox1 and lox2, or lox3 and lox4). Depending on the cell type, each of the four cassettes can integrate into a different stable integration site (and other available sites, such as stable site 1 and stable site 2), shown schematically as positions A and B (SIS) and C and D (stable sites 1 and 2). The reporter genes can be the same or different. The negative selection genes can be the same or different, but are preferably the same. The cell can include additional stable integration sites and integrated cassettes according to the teachings contained herein. A promoter is present 5' of the gene but is not shown.

[0036] [Figure 18] Modifications of the cell in Figure 17 are shown schematically with positions A, B, C, and D indicated. Each of the four cassettes contains flanking RRSs (here, lox1 and lox2, or lox3 and lox4), a gene of interest, a positive selection marker gene, and a reporter* gene. The positive selection marker genes can be the same or different, but are preferably the same. The reporter* genes can be the same or different, but each must be different from any of the reporter genes in the cell in Figure 17. The genes of interest can be the same or different. In this illustration, there are two copies of gene of interest 1 and two copies of gene of interest 2. The cassette in Figure 17 is replaced by the cassette in Figure 18 by RMCE. The cell can contain additional stable integration sites and integrated cassettes according to the teachings contained herein. A promoter is present 5' of the gene but is not shown. DETAILED DESCRIPTION OF THE INVENTION

[0037] definition Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0038] The term "about" in the context of numerical values ​​and ranges refers to a value or range that is approximately or very close to the recited value or range, such that the invention can be performed to have the desired rate, amount, degree, increase, decrease, or degree of expression, concentration, or time, as is clear from the teachings contained herein. Thus, the term encompasses values ​​beyond those that result solely from systematic error. For example, "about" can indicate a value that is either above or below the recited value by a range of approximately + / - 10% or more or less, depending on the ability to perform.

[0039] "AAVS1" can be a genomic safe harbor and refers to adeno-associated virus integration site 1, which is reported to be located on native human chromosome 19 and encompasses approximately 4.7 kilobases. The AAVS1 locus can be used in accordance with the present invention.

[0040] "AAVS1-like" refers to the AAVS1 homolog found in CHO cells and is disclosed herein. AAVS1-like regions containing AAVS1-like genomic safe harbors (GSH) can be used in accordance with the present invention. SEQ ID NO: 2 is an example of an AAVS1-like region.

[0041] A "DNA cassette" or "cassette" is a type of nucleic acid segment that includes at least a promoter, at least one open reading frame, and, optionally, a polyadenylation signal, e.g., an SV40 polyadenylation signal. Other nucleic acid segments, such as an operator, are also optional. Thus, a DNA cassette is a polynucleotide that includes two or more shorter polynucleotides. A cassette can include one or more genes, as well as promoters, enhancers, operators, repressors, transcription termination signals, ribosome entry sites, introns, and polyadenylation signals.

[0042] "COSMC" has been reportedly found in hamster cells. Homologs of partial or complete COSMC loci are candidates for use according to the present invention.

[0043] "CCR5" refers to the CC chemokine receptor type 5 gene, reportedly found in human, mouse, and rat cells. Homologs of partial or entire CCR5 loci are candidates for use according to the present invention.

[0044] "Genomic safe harbor" or "GSH" refers to a site within a cell's genome that can accommodate the insertion of a polynucleotide, such as a DNA cassette, allowing the inserted polynucleotide to function without imposing undue burden on the transformed cell. Thus, a genomic safe harbor is an ideal location for creating a stable integration site for the insertion of a DNA cassette through the practice of the present invention. Genomic safe harbors that can be utilized herein include, but are not limited to, AAVS1 and AAVS1-like. Reported candidate loci include, but are not limited to, CCR5, COSMC, and Rosa26.

[0045] A "genomic safe harbor homology arm" or "GSH homology arm" is derived from and has homology to a genomic safe harbor. Preferably, a genomic safe harbor homology arm contains about 100 to 2000 bases, more preferably about 300 to 1800 bases, more preferably about 400 to 1600 bases, more preferably about 500 to 1500 bases, more preferably about 500 to 1300 bases, more preferably about 500 to 1100 bases, more preferably about 500 to 1000 bases, more preferably about 600 to 1000 bases, more preferably about 700 to 1000 bases, more preferably about 800 to 1000 bases, and even more preferably about 900 to 1000 bases. Typically, a polynucleotide to be inserted into a genomic safe harbor is flanked by a 5' GSH homology arm and a 3' GSH homology arm. See, for example, Figures 4 and 5, which show a lox site-flanked DNA cassette further flanked by GSH homology arms.

[0046] "hRosa26" refers to the human homolog of the mouse Rosa26 locus ("reverse splice acceptor"). "Rosa26" refers to a partial or entire Rosa26 locus, which is reportedly found in hamster cells in addition to mouse and human cells. Homologs of the partial or entire Rosa26 locus are candidates for use according to the present invention.

[0047] An "intron" is a segment of DNA located between exons. Introns are removed to form mature messenger RNA. Preferred introns are those that can influence the start point of translation; examples are the hCMV-IE intron (human cytomegalovirus immediate early protein) and the FMDV intron (foot-and-mouth disease virus).

[0048] A "nucleic acid segment" includes any arrangement of single- or double-stranded nucleotide sequences, which may include, but are not limited to, polynucleotides, promoters, enhancers, operators, repressors, transcription termination signals, ribosome entry sites, and polyadenylation signals.

[0049] "Operably linked" refers to one or more nucleotide sequences that are in a functional relationship with one or more other nucleotide sequences. Such a functional relationship can directly or indirectly control, generate, regulate, enhance, facilitate, permit, attenuate, suppress, or block an action or activity, depending on the selected design. Typical examples include single- or double-stranded nucleic acid segments, which may contain two or more nucleotide sequences arranged within a given segment such that one or more sequences can exert at least one functional effect on the other(s). For example, a promoter operably linked to the coding region of a DNA polynucleotide sequence can promote transcription of the coding region. Other elements, such as enhancers, operators, repressors, transcription termination signals, ribosome entry sites, and polyadenylation signals, can also be operably linked to a polynucleotide of interest to control its expression. The positioning and spacing required to achieve operably linked sequences can be confirmed by approaches available to those skilled in the art, such as screening using Western blots and RT-PCR.

[0050] "Operator" refers to a DNA sequence introduced within or near a polynucleotide sequence such that the polynucleotide sequence can be regulated by the interaction of a molecule capable of binding to the operator, thereby preventing or enabling transcription of the polynucleotide sequence, as the case may be. Those skilled in the art will recognize that an operator must be placed sufficiently proximal to a promoter to control or influence transcription by the promoter and can be considered a type of operable linkage. Operators can be placed either downstream or upstream of a promoter. These include, but are not limited to, the LexA peptide and the operator region of the LexA gene of E. coli, which binds to lactose, and the 45 tryptophan operator that binds to the repressor proteins encoded by the Lad and trpR genes of E. coli. Bacteriophage operators from lambda Pi and phage P22 Mnt, as well as Arc. Preferred operators are the Tet (tetracycline) operator (TetO or TO) and the Arc operator (ArcO or AO). Operators can have native or mutant sequences. For example, mutant sequences of the Tet operator are disclosed in Wissmann et al., Nucleic Acids Res. 14:4253-4266 (1986).

[0051] The Tet operator is preferred and can be used to control transcription using a repressor such as the tetracycline repressor (TetR). Suitable ligands for the repressor are tetracycline (tet), doxycycline (dox), and their derivatives. When the ligand binds to TetR, the affinity of the Tet repressor for the Tet operator is reduced, causing the Tet repressor to dissociate from the operator, thereby allowing the operator to permit transcription. Other repressors can be paired with their own respective operators.

[0052] The phrases "percent identity" or "% identical," in their various grammatical forms, when describing a sequence, are meant to include homologous sequences that exhibit the recited identities along the contiguous region of homology, but the presence of gaps, deletions, or insertions in the compared sequences that do not have a homolog are not taken into account when calculating percent identity. As used herein, the determination of "percent identity" or "% identical" between homologs does not include comparisons of sequences where the homolog does not have a homologous sequence to which it is compared in an alignment. Thus, "percent identity" and "% identical" do not include penalties for gaps, deletions, and insertions.

[0053] "Homologous sequence," in its various grammatical forms, in the context of nucleic acid sequences, refers to a sequence that is substantially homologous to a reference nucleic acid sequence. In some embodiments, two sequences are considered to be substantially homologous if at least 50%-99%, 75%-99%, 85%-99%, 90%-99%, 95%-98%, 98%-99%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more of their corresponding residues are identical over the relevant stretch of residues. In some embodiments, the relevant stretch is a complete (i.e., entire) sequence.

[0054] A "polynucleotide" includes a sequence of covalently linked nucleotides and includes RNA and DNA. An oligonucleotide is considered a shorter polynucleotide. A gene is a DNA polynucleotide (polydeoxyribonucleic acid) that ultimately encodes a polypeptide, typically transcribed from DNA and translated from RNA (polyribonucleic acid). A DNA polynucleotide can also encode an RNA polynucleotide that is not translated, but rather functions as an RNA "product." The type of polynucleotide (i.e., DNA or RNA) is clear from the context of the use of the term. A polynucleotide referred to or specified by the polypeptide it encodes refers to and encompasses all suitable sequences, according to codon degeneracy. Polynucleotides, including those disclosed herein, include percent identity and homologous sequences, where indicated.

[0055] "Polypeptide" and "peptide" refer to a sequence or sequences of covalently linked amino acids. Polypeptides include natural, semi-synthetic, and synthetic proteins, as well as protein fragments. "Polypeptide" and "protein" can be used interchangeably. An oligopeptide is considered a shorter polypeptide.

[0056] A "promoter" refers to a DNA sequence that causes transcription of a DNA sequence to which it is operably linked, i.e., linked in such a way as to allow transcription of the nucleotide sequence of interest when the appropriate signals are present and repressors are absent. Expression of a polynucleotide of interest can be placed under the control of any promoter or enhancer element known in the art. A eukaryotic promoter can be operably linked to a TATA box, which is typically located upstream of the transcription start site.

[0057] Useful promoters that can be used include, but are not limited to, the SV40 early promoter region, the SV40E / L (early-late) promoter, the promoter contained in the 3' long terminal repeat of Rous sarcoma virus, the regulatory sequence of the metallothionein gene, the mouse or human cytomegalovirus major-immediate-early (CMV-MIE) promoter, and other CMV promoters, including the CMVmin promoter. Plant expression vectors containing the nopaline synthase promoter region, the cauliflower mosaic virus 35S RNA promoter, and the promoter of the photosynthetic enzyme ribulose bisphosphate carboxylase; promoter elements from yeast or other fungi, such as the Gal4 promoter, the ADC (alcohol dehydrogenase) promoter, the PGK (phosphoglycerol kinase) promoter, the alkaline phosphatase promoter, and animal transcriptional control regions that exhibit tissue specificity and have been utilized in transgenic animals: elastase I, insulin, immunoglobulins, mouse mammary tumor virus, albumin, C-fetoprotein, C.1-antitrypsin, C.3-globin, and myosin light chain-2. Various forms of the CMV promoter can be used in accordance with the present invention.

[0058] Minimal promoters, such as the CMVmin promoter, can be truncated promoters or core promoters and are preferred for use in regulated expression systems. Minimal promoters and development approaches are widely known and are disclosed, for example, in Saxena et al., Methods Molec. Biol. 1651:263-73 (2017); Ede et al., ACS Synth Biol. 5:395-404 (2016); Brown et al., Biotech Bioeng. 111:1638-47 (2014); Morita et al., Biotechniques 0:1-5 (2012); Lagrange et al., Genes Dev. 12:34-44 (1998). Many CMVmin promoters have been reported in the field.

[0059] A "protein of interest" or "polypeptide of interest" can have any amino acid sequence and includes any protein, polypeptide, or peptide, as well as derivatives, components, domains, chains, and fragments thereof. These include, but are not limited to, viral proteins, bacterial proteins, fungal proteins, plant proteins, and animal (including human) proteins. Protein types can include, but are not limited to, antibodies, bispecific antibodies, multispecific antibodies, antibody chains (including heavy and light chains), antibody fragments, Fv fragments, Fc fragments, Fc-containing proteins, Fc fusion proteins, receptor-Fc fusion proteins, receptors, receptor domains, trap and mini-trap proteins, enzymes, factors, inhibitors, activators, ligands, reporter proteins, selection proteins, protein hormones, protein toxins, structural proteins, storage proteins, transport proteins, neurotransmitters, and contractile proteins. Derivatives, components, chains, and fragments of the above are also included. Sequences can be natural, semi-synthetic, or synthetic. Proteins and polypeptides of interest are encoded by a "gene of interest," which may also be referred to as a "polynucleotide of interest." Where multiple genes (the same or different) are integrated, they may be referred to as "first," "second," "third," "fourth," "fifth," "sixth," "seventh," "eighth," "ninth," "tenth," etc., as is clear from the context of use.

[0060] "Recombinase recognition sites" (RRS), also known as "heterospecific recombination sites," are used in recombinase-mediated cassette exchange (RMCE). For example, Cre / Lox, Dre / Rox, Vre / Vlox, SCre / Slox, and Flp / Frt are suitable RRS systems. RRSs suitable for use with the present invention include LoxP, Lox66, Lox71, Lox511, Lox2272, Lox2372, Lox5171, LoxM2, LoxM3, loxM7, and LoxM11. These sites may be generally referred to as first (1), second (2), third (3), fourth (4), fifth (5), sixth (6), seventh (7), eighth (8), ninth (9), tenth (10), etc., as is clear from the context of use. Cre / Lox is the most commonly used RRS, although other RRSs can be used in place of Cre / Lox in accordance with the present invention.

[0061] As used herein, "reporter protein" refers to any protein capable of directly or indirectly generating a detectable signal. Reporter proteins typically fluoresce or catalyze a colorimetric or fluorescent reaction and are often referred to as "fluorescent proteins" or "color proteins." However, reporter proteins can also be non-enzymatic and non-fluorescent, so long as they can be detected by another protein or moiety, such as a cell surface protein detected with a fluorescent ligand. Reporter proteins can also be inert proteins that become functional through interaction with another protein, either fluorescent or catalyzing a reaction. Therefore, as will be understood by those skilled in the art, any suitable reporter protein can be used. In some embodiments, the reporter protein can be selected from fluorescent proteins, luciferase, alkaline phosphatase, β-galactosidase, β-lactamase, dihydrofolate reductase, ubiquitin, and variants thereof. Fluorescent proteins are sometimes useful for identifying gene cassettes that have or have not been successfully inserted and / or exchanged. Fluid cytometry and fluorescence-activated cell sorting are suitable for detection. Examples of fluorescent proteins are well known in the art and include Discosoma Reporter proteins include, but are not limited to, a first (1), a second (2), a third (3), a fourth (4), a fifth (5), a sixth (6), a seventh (7), an eighth (8), a ninth (9), a tenth (10), etc., as is clear from the context of the usage.A reporter can be considered a type of marker. "Color" or "fluorescence" in their various grammatical forms can also be used to refer more specifically to a reporter protein or gene.

[0062] A "repressor protein," also referred to as a "repressor," is a protein capable of binding to DNA to repress transcription and is encoded by a polynucleotide also referred to herein as a "repressor gene" or "repressor protein gene." Repressors are of eukaryotic and prokaryotic origin. Prokaryotic repressors are preferred. Examples of repressor families include the TetR, LysR, LacI, ArsR, IcIR, MerR, AsnC, MarR, DeoR, GntR, and Crp families. Repressor proteins in the TetR family include ArcR, ActII, AmeR, AmrR, ArpR, BpeR, EnvR, EthR, HemR, HydR, IfeR, LanK, LfrR, LmrA, MtrR, Pip, PqrA, QacR, RifQ, RmrR, SimReg2, SmeT, SrpR, TcmR, TetR, TtgR, TrgW, UrdK, and VarR. YdeS, ArpA, BarA, Aur1B, CalR, CprB, FarA, JadR*, JadR2, MphB, NonG, PhlF, TylQ, VanT, Ta rA, TylP, BM1P1, Bm3R1, ButR, CampR, CamR, DhaR, KstR, LexA-like, AcnR, PaaRR, PsbI, Th1R, Ui dR, YDH1, BetI, McbR, MphR, PhaD, Q9ZF45, TtK, Yhgd, YixD, CasR, IcaR, LitR, LuxR, LuxT, O Examples include paR, Orf2, SmcR, HapR, Ef0113, HlyIIR, BarB, ScbR, MmfR, AmtR, PsrA, and YjdC proteins. See Ramos et al., Microbiol. Mol. Biol. Rev., 69:326-56 (2005). Still other repressors include PurR, LacR, MetJ, and PadR.

[0063] "Selectable" or "selectable" marker proteins include proteins that confer a particular trait, including, but not limited to, drug resistance or other selective advantage. Selectable markers can result in cells receiving the selectable marker gene being resistant to a particular toxin, drug, antibiotic, or other compound, allowing the cell to produce the protein and grow in the presence of the toxin, drug, antibiotic, or other compound, and are often referred to as "positive selectable markers." Suitable examples of antibiotic resistance markers include, but are not limited to, proteins that confer resistance to various antibiotics, such as kanamycin, spectinomycin, neomycin, gentamicin (G418), ampicillin, tetracycline, chloramphenicol, puromycin, hygromycin, zeocin, and / or blasticidin. Other selectable markers, often referred to as "negative selectable markers," cause cells to stop growing, stop protein production, and / or are lethal to the cell in the presence of the negative selectable marker protein. Thymidine kinase and certain fusion proteins can function as negative selectable markers, including, but not limited to, GyrB-PKR. See White et al., Biotechniques, 50:303-309 (May 2011). Selectable marker proteins and corresponding genes (selectable marker genes) can be generally referred to as first (1), second (2), third (3), fourth (4), fifth (5), sixth (6), seventh (7), eighth (8), ninth (9), tenth (10), etc., as is clear from the context of use. In the figures, a selectable marker is a positive selectable marker unless otherwise specified as a negative (neg.) marker.

[0064] A "single guide RNA" or "sgRNA" is used to target Cas9 to a site and is typically 17-24 nucleotides in length.

[0065] A "stable integration site" or "SIS" is a region for site-specific integration of a DNA polynucleotide of interest, and includes a cassette containing a gene and / or other open reading frame, a promoter, and optionally other elements. Stable integration sites include exogenously supplied DNA cassettes and can be generated, preferably in GSH, according to the methods of the invention described and depicted herein. Constructs can be inserted into the SIS by a variety of approaches. Multiple stable integration sites can be generated and located on different chromosomes, in different regions of the same chromosome, or at different locations in the same region of a chromosome.

[0066] "Tetracycline response elements" or "TREs" contain seven copies of the 19-nucleotide TetO, separated by spacers containing 17-18 nucleotides, and are commercially available. The TetO sequence can vary, and nucleotide substitutions are known. For example, altered sequences based on the Tet operator are disclosed in Wissmann et al., Nucleic Acids Res. 14:4253-66 (1986). The spacer is not sequence-specific. The spacers can be similar, but need not all be identical. A TRE is considered a type of operator as used herein.

[0067] All numerical limits and ranges set forth herein include all numbers or values ​​surrounding or between the numbers in the range or limit. The ranges and limits set forth herein expressly represent and denote all integer, decimal, and fractional values ​​defined and encompassed by the range or limit. The ranges and limits set forth herein expressly represent and denote all integer, decimal, and fractional values ​​defined and encompassed by the range or limit. Accordingly, the recitation of ranges of values ​​herein is merely intended to serve as a shorthand method of individually referring to each separate value falling within that range, unless otherwise indicated herein, and each separate value is incorporated into the specification as if it were individually set forth herein.

[0068] Detailed Description The present invention provides mammalian cells with multiple stable integration sites, suitable for producing proteins of interest, including viral proteins, and for producing viral vectors, including adeno-associated viral vectors (AAV). One or more stable integration sites can be within a genomic safe harbor, and one or more stable integration sites can be outside a particular genomic safe harbor. Multiple stable integration sites can be created and located on different chromosomes, different regions of the same chromosome, or different locations in the same region of a chromosome.

[0069] Genomic safe harbors are discussed in Pellenz et al., Hum. Gene Therapy 30:814-28 (2019) and Papapetrou et al., Molecular Therapy 24:678-84 (2016).

[0070] Preferably, the stable integration site contains a recognition site that allows recombinase-mediated cassette exchange (RMCE). Modification of the cell genome can be achieved using known approaches using heterospecific recombination sites (also known as RRSs), such as Cre / Lox, Flp / Frt, transcription activator-like effector nucleases (TALENs), TAL effector domain fusion proteins, zinc finger nucleases (ZFNs), ZFN dimers, or RNA-guided DNA endonuclease systems such as CRISPR / Cas9. See U.S. Patent No. 9,816,110 at cols. 17-18; Sajgo et al., PLoS ONE 9:e91435 (2014); Suzuki et al., Nucl. Acids. Res. 39:e49 (2011). Integration using Bxb1 integrase can also be achieved in human, mouse, and rat cells. Russell et al., Biotechniques 40:460-64 (2006).

[0071] Recombinase recognition sites, also known as heterospecific integration sites, may be generally referred to as first (1), second (2), third (3), fourth (4), fifth (5), sixth (6), seventh (7), eighth (8), ninth (9), tenth (10), etc., as will be apparent from the context of use. Lox sites suitable for use in accordance with the present invention include, but are not limited to, LoxP, Lox66, Lox71, Lox511, Lox2272, Lox2372, Lox5171, LoxM2, LoxM3, loxM7, and LoxM11. Other RRSs may be used as well. Lox sites are the most commonly used type of RRS, but different RRSs may be used as well.

[0072] The homology arms preferably begin approximately 10-20 bases, more preferably 10-15 bases, within the cleavage site. Larger distances can be used as well, but with lower efficiency. To ensure that the DNA cassette inserted into the genomic safe harbor maintains stability in the event that homologous repair could potentially recreate the targetable site determined by one of skill in the art, the guide arm region of the DNA cassette can be engineered to contain alterations (e.g., base mismatches) that disrupt the function of the CRISPR target site. There are two approaches that can be used independently or together. The first approach is to insert a base substitution into the 20-base target site of the CRISPR or into the protospacer adjacent motif (PAM), which is usually 2-6 bases, to create a base mismatch. The second approach is to create a donor plasmid whose insertion splits the CRISPR target site or splits the CRISPR target site from the PAM.

[0073] Human cell lines include amniotic cells (such as human amniotic epithelial cells), Hela cells, Per.C6 cells, and HEK293 cells. Examples of HEK293 cells include, but are not limited to, HEK293, HEK293A, HEK293E, HEK293F, HEK293FT, HEK293FTM, HEK293H, HEK293MSR, HEK293S, HEK293SG, HEK293SGGD, and HEK293T, as well as mutants and variants thereof. Rodent cell lines, such as Sp2 / 0 cells, BHK cells, and CHO cells, as well as mutants and variants thereof, can also be used in accordance with the present invention. CHO cells include, but are not limited to, CHO-ori, CHO-K1, CHO-s, CHO-DHB11, CHO-DXB11, and CHO-K1SV, as well as mutants and variants thereof.

[0074] The mammalian cells of the present invention are produced by advantageously producing and utilizing a cellular intermediate having a cassette containing a Cas9 endonuclease gene flanked by recombinase recognition sites and integrated into the genome via RCME. Without being bound by any theory, the present use of an integrated Cas9 gene appears to increase the efficiency of integration of homology arms into the genome safe harbor by increasing the occurrence of cleavage in genomic DNA caused by the Cas9 endonuclease when expressed. The use of a stably integrated Cas9 gene of the present invention is 10, 10 times more efficient than HDR without a stably integrated Cas9 gene. 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , or 10 10 This provides high HDR efficiency. Finally, this intermediate cell can be further subjected to RMCE to remove the cassette containing the Cas9 gene.

[0075] As a starting point for engineering a cell, a polynucleotide sequence of interest, and an operably linked promoter and optional operator, can be introduced into a cell by transfection of a plasmid containing the polynucleotide sequence and elements. Thus, the invention includes the generation of cells as described.

[0076] Suitable plasmid constructs can be prepared by those skilled in the art. Useful regulatory elements already reported or known in the art can also be included in the plasmid construct used to transfect cells. Some non-limiting examples of useful regulatory elements include, but are not limited to, promoters, enhancers, sequences encoding suitable mRNA ribosomal binding sites, and sequences controlling the termination of transcription and translation. Suitable plasmid constructs may also include non-transcribed elements such as origins of replication, other 5' or 3' flanking non-transcribed sequences, and 5' or 3' non-translated sequences such as splice donor and acceptor sites. One or more selectable marker genes may also be incorporated. Selectable marker proteins and reporter proteins useful for use in the present invention are known and can be easily identified by those skilled in the art.

[0077] Plasmid constructs encoding the gene of interest can be delivered to cells using viral vectors or via non-viral transfer methods.

[0078] Non-viral methods of nucleic acid transfer include naked nucleic acid, liposomes, and protein / nucleic acid conjugates. The plasmid construct introduced into the cell can be linear or circular, single-stranded or double-stranded, and can be DNA, RNA, or any modification or combination thereof.

[0079] Plasmid constructs can be introduced into cells by transfection. Those skilled in the art are aware of many different transfection protocols and can select an appropriate system for use in transfecting cells. Generally, transfection methods include, but are not limited to, viral transduction, cationic transfection, liposome transfection, dendrimer transfection, electroporation, heat shock, nucleofection transfection, magnetofection, nanoparticles, biolistic delivery (gene gun), and proprietary transfection reagents such as Lipofectamine, Dojindo Hilymax, Fugene, jetPEI, Effectene, or DreamFect. [Example]

[0080] The present invention is further described by the following examples which illustrate various embodiments and aspects of the invention, but which are not intended to limit the invention in any way. In the examples, selectable markers are positive selectable markers unless otherwise specified as negative (neg.) markers.

[0081] Example 1 This example relates to the generation of mammalian cells containing a repressor, such as TetR, under the control of a promoter, such as the CMV promoter. See Figure 1. Cells are transfected with a polynucleotide containing the promoter and repressor gene. The polynucleotide is randomly inserted into the cell genome. Western blots and TaqMan can be used on cell pools to identify transformants and determine the average copy number. The incorporation of a repressor, such as TetR, allows for the regulation of transcription of the polynucleotide under the control of the promoter and operator.

[0082] Example 2 This example relates to further manipulation of the cells of Example 1. DNA cassette 1 is shown schematically in FIG. 2 and contains flanking lox sites (1 and 2), and further contains, in 5' to 3' order, a promoter, a reporter gene (1) encoding a reporter protein (1), an IRES, and a selectable marker gene (1) encoding a selectable marker protein (1), and a polyadenylation signal. DNA cassette (1) can optionally contain an operator operably linked to the promoter. DNA cassette (1) is inserted randomly or site-specifically into the cell genome. The first and second lox sites in DNA cassette (1) are different.

[0083] When the tet operator is used in the DNA cassette (1), multiple rounds of -ligand / +ligand sorting and single-cell sorting identify Lox site-stable cells for dox-regulated expression. Thus, in the presence of a ligand such as doxycycline or tetracycline, TetR does not bind to the operator, thus creating permissive conditions for transcription of the reporter gene (1) and selectable marker polynucleotide (1).

[0084] Example 3 In this example, RMCE is performed to replace DNA cassette (1) with DNA cassette (2) in the cell of Example 2. As shown schematically in Figure 3, DNA cassette (2) contains flanking lox sites (1 and 2) and further contains, in 5' to 3' order, a promoter, a selectable marker gene (2) encoding a selectable marker protein (2), an IRES, and a reporter gene (2) encoding a reporter protein (2) and a Cas9 gene under the control of a second promoter (optionally operably linked to an operator).

[0085] In one embodiment, the CMV promoter is operably linked to a tet operator to control transcription of the Cas9 gene. When the cell is in the presence of doxycycline or tetracycline, TetR can no longer bind to the tet operator, thus allowing transcription of the Cas9 gene to occur. Reporter protein (1) is different from reporter protein (2), and selectable marker protein (1) is different from selectable marker protein (2).

[0086] Example 4 This example relates to the integration of a DNA cassette (3) into a genomic safe harbor. See Figure 4. The DNA cassette (3) comprises a polynucleotide comprising, in 5' to 3' order, a first genomic safe harbor homology arm comprising an sgRNA target site, a lox site (3), a promoter operably linked to a reporter gene (3) encoding a reporter protein (3), a polyadenylation signal, a lox site (4), and a second genomic safe half harbor homology arm comprising an sgRNA target site. The first and second guide arm target sites can each comprise a region with alterations if required to avoid recreating a targetable site. The lox sites (1), (2), (3), and (4) are different from one another. The reporter protein (3) is different from the reporter protein (2). The reporter protein (3) and the reporter protein (1) can be the same or different. A homology arm of approximately 1,000 bases is used in this example.

[0087] When the Cas9 endonuclease is expressed, the efficiency of integration of DNA cassette 3 increases. Without being bound by any theory, the use of the integrated Cas9 gene of the present invention appears to increase the efficiency of integration by increasing the occurrence of cleavage in genomic DNA caused by the Cas9 endonuclease. The use of the stably integrated Cas9 gene of the present invention is 10, 10 times more efficient than HDR without a stably integrated Cas9 gene. 2, 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , or 10 10 Provides high HDR efficiency.

[0088] Optionally, alterations in the first and second genomic safe harbor homology arms allow DNA cassette (3) to remain integrated by avoiding re-creation of a targetable site, and the smaller cassette therein, i.e., the region between lox site (3) and lox site (4), is available for RMCE and is referred to as the stable integration site.

[0089] Example 5 This example, which concerns the final form of the cell line, is shown diagrammatically in Figure 5. To ensure stability of the cell line over time, it is preferable to remove the Cas9 gene. Therefore, DNA cassette (2) is replaced by DNA cassette (4) by RMCE, which removes the Cas9 gene. DNA cassette (4) contains flanking lox sites (1 and 2) and a reporter gene (4) encoding a reporter protein (4) under the control of a promoter. Reporter protein (4) is different from reporter protein (2) and reporter protein (3), and preferably different from reporter protein (1).

[0090] The resulting cell has two integration sites within the genome, one integration site within the genomic safe harbor (e.g., a stable integration site), and one integration site outside that particular genomic safe harbor. Additional integration sites can still be created by applying the approaches described above, including the use of an integrated Cas9 gene and the use of additional, different GSH homology arms.

[0091] Example 6 This example compares the efficiency of using Cas9 with the HDR disclosed herein compared to conventional homology-directed repair (HDR). As reported in the literature, HDR is accurate, but the desired recombination event occurs infrequently: 6 ~10 9 One in every 100 cells (0.0001% to 0.0000001%). Hsu et al., Cell 157:1262-78(2014).

[0092] To evaluate the benefits of a stably integrated Cas9 gene, CHO cells were engineered with the sites disclosed in U.S. Patent No. 7,771,997 ("Stable Site 1") and U.S. Patent No. 9,816,110 ("Stable Site 2"). Regeneron offers a line of products and services called EESYR®. CHO cells with sequences integrated into Stable Site 1 and Stable Site 2 are disclosed in US 2019 / 0233544 A1, each referred to herein as an "enhanced expression locus." The sequences in these patents and in Examples 11 and 12 can be used in accordance with the inventions described and illustrated herein.

[0093] CHO cells were engineered to contain a cyanofluorescent protein reporter gene under the control of a promoter at stable site 1, and a selectable marker gene and a yellow fluorescent protein reporter gene under the control of the same promoter at stable site 2. Additionally, a Cas9 gene under the control of a second promoter with an operator was also inserted into stable site 2. The Cas9 gene can be eventually removed according to the teachings contained herein.

[0094] Cyanofluorescent proteins can be made fluorescent green by changing the tyrosine residue at position 66 to tryptophan. The sgRNA delivery plasmid contains a selectable marker (ampicillin resistance), a POL III promoter (RNA polymerase III promoter), a target sequence, and a gRNA scaffold, a POL III terminator, and digestion sites 1 and 2. The Pol III promoter contains H1 and U6.

[0095] As shown in Figure 6, sgRNA delivery plasmids were constructed containing an HDR template: a 104-mer insert (with a 57-bp arm and a 45-bp arm), a 401-mer insert (with a 198-bp arm and a 201-bp arm), or a 1030-mer (with a 524-bp arm and a 504-bp arm) containing homology arms, and a sequence conferring a cyano-to-green color change, which in this example consisted of two nucleotides ("repair nucleotides"). The HDR template was inserted into the digestion site (e.g., NotI and / or other appropriate site) of the sgRNA delivery plasmid to form the sgRNA target plasmid. An sgRNA delivery plasmid without an insert (no HDR template) was used as a control.

[0096] Figure 7 shows that the control did not express green positivity in Q1. Cells with HDR templates showed green positivity in Q1, and the green-positive population in Q1 consistently increased with increasing HDR template size (from left to right). Cells with the 1030-mer HDR template showed the highest repair efficiency, approximately 6.5%.

[0097] The cells in this example have stable site 1 and stable site 2 engineered into GSH according to the present invention, as well as SIS, and therefore have three sites for stable integration of a gene of interest.

[0098] Example 7 - Generation of intermediate human cells containing stable integration sites in the genomic safe harbor (AAVS1) In this example, the starting point is HEK293 cells with a stably integrated Cas9 gene flanked by Lox sites 3 and 4. The Cas9 gene is under the control of at least one promoter (not shown). AAVS1 is also shown schematically. See Figure 8. The cells can be generated according to Examples 1-4 and Figures 1-4.

[0099] A targeting plasmid comprising an sgRNA target site, a left homology arm (here, a GSH homology arm) for insertion into a region such as a genomic safe harbor (here, AAVS1), a Lox1 site, a reporter gene (color 1), a Lox2 site, and a right homology arm (here, a GSH homology arm) for insertion into a region such as a genomic safe harbor (here, AAVS1). For alternative targeting plasmids, see Figures 9A and 9B. At the 3' end, one targeting plasmid has a reporter gene (color 2), see Figure 9A. The other targeting plasmid has a negative selection gene (negative selection 1) at the 3' end, see Figure 9B. The promoter and optionally other portions (such as operators) are represented by arrows pointing in the 5' to 3' direction in Figures 9A and 9B. Both plasmids insert color 1 into a region such as a genomic safe harbor (here, AAVS1).

[0100] Cas9-mediated integration of a targeting plasmid (e.g., Figure 9A or Figure 9B) into the genomic safe harbor (AAVS1) of HEK293 cells is shown schematically in Figure 10. Color 1 is flanked by Lox1 and Lox2. A gene of interest can replace Color 1 via RMCE.

[0101] When the targeting plasmid according to Figure 9A is properly integrated, the cells will be color 1 positive and color 2 negative. When the targeting plasmid according to Figure 9B is properly integrated, the cells will be color 1 positive and will be able to grow because the negative selection gene has been removed. This cell is considered an intermediate. Finally, as shown in Figure 8, the cells can be further subjected to RMCE at lox sites 3 and 4 to remove the cassette containing the Cas9 gene. See, e.g., Example 5.

[0102] The accuracy of the methodology of the present invention is demonstrated in Figures 10 and 11. Figure 11 shows the insert of Figure 10 in more detail. The cell genome containing AAVS1 is flanked by the insert and the 5' and 3' ends. Color 1 is flanked by Lox1 and Lox2. Figure 11, left, identifies the locations of the 5' genomic primer and 3' insert primer used in the 5' junction PCR. Figure 11, right, identifies the locations of the 5' insert primer and 3' genomic primer used in the 3' junction PCR.

[0103] Junction PCR demonstrates that the correct size fragment is amplified and labeled as "stable Cas9-targeted cells." See Examples 12 and 13. Stable Cas9-targeted cells and 5' and 3' junctions are obtained and detected, confirming the correct insertion. Positive and negative controls are in the right column of each gel.

[0104] Example 8 - CHO regions and sequences In CHO cells, the sequences shown in U.S. Patent Nos. 7,771,997 (stable site 1) and 9,816,110 (stable site 2) can be utilized. Sequences within the percent identity values ​​of U.S. Patent Nos. 7,771,997 and 9,816,110 and homologous sequences are incorporated herein by reference. The AAVS1-like regions disclosed herein can be used to generate stable integration sites according to the present invention.

[0105] Candidate loci for use in the present invention have been reported in the literature. Hamaker and Lee, Curr. Op. Chem. Eng. 22:152-60 (2018) identify 30 hotspot loci. Hilliard and Lee, Biotech. Bioeng. 118:659-75 (2021) attempted to identify safe harbor regions in CHO using epigenomic analysis of Hi-C stable regions and found overlap with five of the 30 regions identified by Hamaker and Lee. See Supplementary Table 3 in Hilliard and Lee. Gaidukov et al., Nucl. Acids Res. 46:4072-86 (2018) also identify loci for integration in CHO cells, including the putative Rosa26 locus. Lee et al., Scientific Reps. 5:8572 (2015) reported the COSMC locus in hamster cells. In summary, these papers identify several unannotated and genic regions in CHO, the genic regions are shown below. [Table 1]

[0106] Example 9 - CHO cells with more than two insertion sites CHO cells containing multiple insertions are exemplified using the cells disclosed in US2019 / 0233544A1. Stable site 1 and stable site 2 can be used according to the teachings contained herein, first utilizing an integrated Cas9 gene. One or more stable integration sites are created in a genomic safe harbor, such as an AAVS1-like region (see, e.g., SEQ ID NO: 2) and corresponding guide sequence (see, e.g., SEQ ID NOs: 13-419). The guide sequence is assigned to the target sequence of SEQ ID NO: 2 at the following positions: (a) 1 to 2000, (b) 2001 to 4000, (c) 4001 to 6000, (d) 6001 to 8000, (e) 8001 to 10,000, (f) 10,001 to 12,000, (g) 12,001 to 14,000, (h) 14,001 to 16,000, (i) 16,001 to 18,000, (j) 18,001 to 20,000, (k) 20,001 to 22,000, (l) 22,001 to 24,000, ( (m) 24,001 to 26,000, (n) 26,001 to 28,000, (o) 28,001 to 30,000, (p) 30,001 to 32,000, (q) 32,001 to 34,000, (r) 34,001 to 36,000, (s) 36,001 to 38,000, (t) 38,001 to 40,000, (u) 40,001 to 42,000, and (v) 42,001 to 44,232.

[0107] Stable Site 1 and Stable Site 2 of U.S. Patent Nos. 7,771,997 and 9,816,110 can be used for expression of a gene of interest encoding a protein of interest. Cells with SISs can ultimately have 3, 4, 5, 6, 7, 8, 9, 10, or more sites for expressing a gene of interest.

[0108] Preferably, CHO cells containing stable sites 1 and 2 are modified to create a third site, i.e., a stable integration site, in the genomic safe harbor. A preferred genomic safe harbor for the creation of such CHO cells is in the AAVS1-like region. Other CHO cell types can be used to create multiple sites according to the teachings contained herein.

[0109] Figure 14 shows a schematic representation of an exemplary cell containing three cassettes integrated into regions of the genome with adjacent RRSs (here, lox1 and lox2). Depending on the cell type, each of the three cassettes can be integrated into a different stable integration site, as well as other available sites (e.g., stable site 1 and stable site 2), shown schematically as positions A, B, and C. The reporter genes can be the same or different. The negative selection genes can be the same or different, but are preferably the same. The cell can include additional stable integration sites and integrated cassettes according to the teachings contained herein.

[0110] Figure 15 shows a schematic representation of the modification of the cell of Figure 14, with positions A, B, and C shown. Each of the three cassettes contains flanking RRSs (here, lox1 and lox2), a gene of interest, a positive selection marker gene, and a reporter* gene. The positive selection marker genes can be the same or different, but are preferably the same. The reporter* genes can be the same or different, but each must be different from any of the reporter genes in the cell of Figure 14. The genes of interest can be the same or different. The cassette of Figure 14 is replaced by the cassette of Figure 15 by RMCE. The cell can contain additional stable integration sites and integrated cassettes according to the teachings contained herein.

[0111] A combination of negative and positive selection ensures the isolation of cells that have undergone recombination at all sites. If the gene of interest is the same in each of the three cassettes, the cells can produce high-yield protein expression, for example, 7, 8, 9, 10, or even more grams of protein per liter (g / L).

[0112] Figure 16 shows the results of five different human IgG antibodies stably integrated using Cre-lox recombination into CHO K1-derived hosts engineered with either two integration sites (stable sites 1 and 2) or three integration sites (stable site 1, stable site 2, and AAVS1-like (see SEQ ID NO: 2)). Isogenic cell lines (ICLs) were isolated using flow cytometry. Fed-batch production of ICLs was inoculated into chemically defined production medium, and production cultures were run for 13 days. Antibody titers in the conditioned medium were determined using a Protein A HPLC-based method; each triple-site cell expressing a given antibody (1, 2, 3, 4, or 5) expressed greater amounts of protein than the comparable double-site cells. Three-site cells can provide an increase of 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150% or more over two-site cells.

[0113] Alternatively, different genes of interest can be used in the cassettes, for example, antibody heavy and light chain sequences can be genes of interest.

[0114] Turning to four-site cells, preferably, CHO cells containing stable sites 1 and 2 are modified to create a third and fourth site, i.e., a stable integration site, in a genomic safe harbor. A preferred genomic safe harbor for the creation of such CHO cells is in the AAVS1-like region and can be the third site. The fourth site can be created in other loci, including, but not limited to: [Table 2]

[0115] Other CHO cell types can be used to generate multiple sites according to the teachings contained herein.

[0116] Figure 17 shows a schematic representation of an exemplary cell containing four cassettes integrated into regions of the genome with adjacent RRSs (here, lox1 and lox2, or lox3 and lox4). Depending on the cell type, each of the four cassettes can be integrated into a different stable integration site, as well as other available sites (e.g., stable site 1 and stable site 2), shown schematically as positions A, B, C, and D. The reporter genes can be the same or different. The negative selection genes can be the same or different, but are preferably the same. The cell can contain additional stable integration sites and integrated cassettes according to the teachings contained herein.

[0117] FIG. 18 shows the modification of the cell of FIG. 17, shown at positions A, B, C, and D. 17. The four cassettes each contain flanking RRSs (here, lox1 and lox2, or lox3 and lox4), a gene of interest, a positive selection marker gene, and a reporter* gene. The positive selection marker genes can be the same or different, but are preferably the same. The reporter* genes can be the same or different, but each must be different from any of the reporter genes in the cell of FIG. 17. The genes of interest can be the same or different. In this illustration, there are two copies of gene of interest 1 and two copies of gene of interest 2. The cassette of FIG. 17 is replaced by the cassette of FIG. 18 by RMCE. The cell can contain additional stable integration sites and integrated cassettes according to the teachings contained herein.

[0118] A combination of negative and positive selection ensures the isolation of cells that have undergone recombination at all sites. Quadruple-site cells are useful for generating bispecific antibodies, where two different heavy / light chain plasmids can be targeted to different sites.

[0119] Example 10 - Genomic Safe Harbor Sequences Genomic safe harbor sequences and the like are described herein, and many are available in the literature and are publicly available. Exemplary sequences are shown below. Human AAVS1 sequence JPEG2024539076000003.jpg14169 (sequence number 1) CHO AAVS1-like region sequence (Guidelines for insertion are further provided below in Example 13) (SEQ ID NO: 2) CCR5 sequence JPEG2024539076000005.jpg8169 (sequence number 3) hRosa26 sequence (SEQ ID NO: 4) JPEG2024539076000007.. ATTTGGGGTAGAATCTAAAGGGCATATTTTTAAAAAAACTTTTAGTTCTAAAGACAAAAGAGTTTAACCTAAAACAGAACAAAGAGAAGGGCCTTTGAAGCAGTATGATTGATTATAT Example 11 - CHO and Mouse Stable Site 1 Sequence - U.S. Patent No. 7,771,997 211>6473 <212>DNA <213>Cricetulus griseus <400>1 (SEQ ID NO: 5) tctagaaaca aaaccaaaaa tattaagtca ggcttggctt caggtgctgg ggtggagtgc 60 tgacaaaaat acacaaattc ctggctttct aaggcttttt cggggattca ggtattgggt 120 gatggtagaa taaaaatctg aaacataggt gatgtatctg ccatactgca tgggtgtgta 180 tgtgtgtgta tgtgtgtctg tgtgtgtgcc cagacagaaa taccatgaag gaaaaaaaca 240 cttcaaagac aggagagaag agtgacctgg gaaggactcc ccaatgagat gagaactgag 300 cacatgccag aggaggtgag gactgaacca ttcaacacaa gtggtgaata gtcctgcaga 360 cacagagagg gccagaagca ctcagaactc cagggggtca ggagtggttc tctggaggct 420 tctgcccttg gaggttcctg aggaggaggc ttccatattg aaaatgtagt tagtggccgt 480 ttccattagt acagtgacta gagagagctg agggaccact ggactgaggc ctagatgctc 540 agtcagatgg ccatgaaagc ctagacaagc acttccgggt ggaaaggaaa cagcaggtgt 600 gaggggtcag gggcaagtta gtgggagagg tcttccagat gaagtagcag gaacggagac 660 gcactggatg gccccacttg tcaaccagca aaagcttgga tcttgttcta agaggccagg 720 gacatgacaa gggtgatctc ggtttttaaa aggctttgtg ttacctaatc acttctatta 780 gtcagatact ttgtaacaca aatgagtact tggcctgtat tttagaaact tctgggatcc 840 tgaaaaaaca caatgacatt ctggctgcaa cacctggaga ctcccagcca ggccctggac 900 ccgggtccat tcatgcaaat actcagggac agattcttca ctaggtactg atgagctgtc 960 ttggatgcaa atgtggcctc ttcattttac tacaagtcac catgagtcag gaggtgctgt 1020 ttgcacagtg tgactaagtg atggagtgtt gactgcagcc attcccggcc ccagcttgtg 1080 agagagatcc ttttaaattg aaagtaagct caaagttacc acgaagccac acatgtataa 1140 actgtgtgaa taatctgtgc acatacacaa accatgtgaa taatctgtgt acatgtataa 1200 actgtgtgaa taatctgtgt gcagcctttc cttacctact accttccagt gatcaggttt 1260 ggactgcctg tgtgctactg gaccctgaat gtccccaccg ctgtcccctg tcttttacga 1320 ttctgacatt tttaataaat tcagcggctt cccctctgct ctgtgcctag ctataccttg 1380 gtactctgca ttttggtttc tgtgacattt ctctgtgact ctgctacatt ctcagatgac 1440 atgtgacaca gaaggtgttc cctctggaga catgtgatgt ccctgtcatt agtggaatca 1500 gatgccccca aactgttgtc cagtgtttgg gaaagtgaca cgtgaaggag gatcaggaaa 1560 agaggggtgg aaatcaagat gtgtctgagt atctcatgtc cctgagtggt ccaggctgct 1620 gacttcactc ccccaagtga gggaggccat ggtgagtaca cacacctcac acatactata 1680 tccaacacac acacacacac acacacacac acgcacgcac gcacgcacgc acgcacacat 1740 gcacacacac gaactacatt tcacaaacca catacgcata ttacacccca aacgtatcac 1800 ctatacatac cacacataca cacccctcca cacatcacac acataccaca cccacacaca 1860 gcacacacat acataggcac acattcacac accacacata tacatttgtg tatgcataca 1920 tgcatacaca cacaggcaca cagacaccac acacatgcat tgtgtacaca cacatgcata 1980 cacacacata ggcacacatt gagcacacac atacatttgt gtacgcacac tacatagaca 2040 tatatgcatt tgtatatgca cacatgcatg cacacataca taggcacaca tagagcacac 2100 acatacattt gtgtatgcac acatgcacac accaatcaca tgggaagact caggttcttc 2160 actaaggttc acatgaactt agcagttcct ggttatctcg tgaaacttgg aagattgctg 2220 tggagaagag gaagcgttgg cttgagccct ggcagcaatt aaccccgccc agaagaagta 2280 ggtttaaaaa tgagagggtc tcaatgtgga acccgcaggg cgccagttca gagaagagac 2340 ctacccaagc caactgagag caaaggcaga gggatgaacc tgggatgtag tttgaacctc 2400 tgtaccagct gggcttcatg ctattttgtt atactcttat taaatattct tttagtttta 2460 tgtgcgtgaa taccttgctt gcataaatgt atgggcactg tatgtgttct tggtgccggt 2520 ggaggccagg agagggcatg gatcctccgg agctggcgtt tgagacagtt gtgacccaca 2580 gtgtggggtc tgggaactgg gtcttagtgt tccgcaagtg cagctggggc tcttaacctc 2640 tgagccatcc ctccagcttc aagaaactta ttttcttagg acatggggga agggatccag 2700 ggctttaggc ttgtttgttc agcaaatact cttttcgtgt attttgaatt ttattttatt 2760 ttactttttt gggatagaat cacattctgc agctcaggct gggcctgaac tcatcaaaat 2820 cctcctgtct cagtctacca ggtgataaga ttactgatgt gagcctggct ttgacaagca 2880 ctttagagtc cccagccctt ctggacactt gttccaagta taatatatat atatatatat 2940 atatatatat atatatatat atatattgtg tgtgtgtgtt tgtgtgtgta tgagacactt 3000 gctctaaggg tatcatatat atccttgatt tgcttttaat ttatttttta attaaaaatg 3060 attagctaca tgtcacctgt atgcgtctgt atcatctata tatccttcct tccttctctc 3120 tctttctctc ttcttcttct cacccccaag catctatttt caaatccttg tgccgaggag 3180 atgccaagag tctcgttggg ggagatggtg agggggcgat acaggggaag agcaggagga 3240 aagggggaca gactggtgtg ggtctttgga gagctcagga gaatagcagc gatcttccct 3300 gtccctggtg tcacctctta cagccaacac cattttgtgg cctggcagaa gagttgtcaa 3360 gctggtcgca ggtctgccac acaaccccaa tctggcccca agaaaaggca cctgtgtgtg 3420 actctggggt taaaggcgct gcctggtcgt ctccagctgg acttgaaact cccgtttaat 3480 aaagagttct gcaaaataat acccgcagag tcacagtgcc aggttcccgt gctttcctga 3540 agcgccaggc acgggttccc taggaaatgg ggccttgctt gccaagctcc cacggcttgc 3600 cctgcaaacg gcctgaatga tctggcactc tgcgttgcca ctgggatgaa atggaaaaaa 3660 gaaaaagaag aagtgtctct ggaagcgggc gcgctcacac aaacccgcaa cgattgtgta 3720 aacactctcc attgagaatc tggagtgcgg ttgccctcta ctggggagct gaagacagct 3780 agtgggggcg gggggaggac cgtgctagca tccttccacg gtgctcgctg gctgtggtgc 3840 atgccgggaa ccgaaacgcg gaactaaagt caagtcttgc tttggtggaa ctgacaatca 3900 acgaaatcac ttcgattgtt ttcctctttt tactggaatt cttggatttg atagatgggg 3960 gaggatcaga gggggagggg aggggcgggg agacggaggg aggaggggag gaggggagga 4020 ggggaggagg ggaggagggg aagggatgga ggaaaatact aacttttcta attcaacatg 4080 acaaagattc ggagaaagtg caccgctagt gaccgggagg aggaatgccc tattgggcat 4140 tatattccct gtcgtctaat ggaatcaaac tcttggttcc agcaccaagg attctgagcc 4200 tatcctattc aagacagtaa ctacagccca cacggaagag gctatacaac tgaagaaata 4260 aaattttcac tttatttcat ttctgtgact gcatgttcac atgtagagag ccacctgtgt 4320 ctaggggctg atgtgctggg cagtagagtt ctgagcccgt taactggaac aacccagaac 4380 tcccaccaca gttagagctt gctgagagag ggaggccctt ggtgagattt ctttgtgtat 4440 ttatttagag acagggtctc atactgtagt ccaagctagc ctccagctca cagaaattct 4500 cctgttccgg tttccaaagt actggagtta tgagtgtgtg ttaattgaac gctaagaatt 4560 tgctgattga agaaaacctc aagtgggttt ggctaatccc cacgacccca gaggctgagg 4620 caggaggaat gagagaattc aaggtttgcc agagccacag ggtgagctca atgtggagac 4680 tgtgagggtg agctcaatgt ggagactgtg agggtgagct caatgtggag actgtgaggg 4740 tgagctcaat gtggagactg tgagggtgag ctcaatgtgg agactgtgag ggtgagctca 4800 atgtggagac ctgtatcaag ataatatag tagtagtac atgcaggcg agggtgtggt 4860 tgagtgtag agcagttagt tgattgaca tgcttgaggt ctcccggtcc atctgtggcc 4920 ctgcaacagg agggaggga gggagggggg gagagagagagagagag agahagaagc 4980 taggatagggg atgagag gagggaa acgggaa attcagactc cttcctgagt 5040 tccgccaacg cctagtgaca tcctgtgcac accctaggt ggccttgtg tgcactgc 5100 ttgggtggtc gggaaaggca tttcagctt gttgcagaac tgccacagta gcatgctggg 5160 tccgtgaaag ttctcccg ttaacaagaa gtcttacta cttgtgacct caccagtgaa 5220 aatttcttta attgtctcct gttgttctgg gttttgcatt ttgtttcta aggatacatt 5280 cctgggtgat gtcatgaagt ccccaagac acagtggggc tgtgttgat tgggaagat 5340 gatttatctg gggtgtcaa aggaaagaa gggaacagg cacttgggaa atgtcctcc 5400 cgcccacccg aattttggct tggcaccgt ggtggaggag cagaacac gtggacgttt 5460 gaggaggcat ggggtcctag gaggacagga agcagaagga gagagctggg ctgacagcct 5520 gcaggcattg cacagttca gaggagatt acagcatgac tgagtttta gggatccaac 5580 agggacctgg gtagagattc tgtgggctct gaggcaactt gacctcagcc agatgtatt 5640 tgaataacct gctcttagag ggaaaacaga catagcaac agagccacgt ttagtgag 5700 aactctcact ttgcctgagt catgtgcggc catgcccagg gtcaggctg acacctcact 5760 caaaaacaag tgagaattg agacaatcc gtggtggcag ctactggaag ggccaccaca 5820 tccccagaaa gagtggagct gctaaaaagc cattgtgat aggcacagtt atcttgaatg 5880 catggagcag agattacgga aaatcgaga atgttaatga ggcaacattc gagttgagtc 5940 attcagtgtg ggaaacccag acgctccat ccctaaag gaacatctg ctctcagtca 6000 aaatggaat aaaattggg gcttgaattt ggcaatgat tcagactct gtgtaggtat 6060 tttcacacgc acagtggata atttcatgt tggatttat tgtgctaaa aggcagaaaa 6120 gggtaaaag cacatcttaa gagttatgag gttctacgaa taaaataat gttacttaca 6180 gctattcctt aattagtacc cccttccacc tgtgtatt tcctgagata gtcagtgggg 6240 aaaagatctc tccttctctt ctttctcccc ctcccctcct ctccctccct ccctccctcc 6300 ctccctcctc tccctccctc cccctttcct tctttctttg ctccttctcc tctgcctcct 6360 tctccctttc ttcttcattt attctaagta gcttttaaca gcacaccaat tacctgtgta 6420 taacgggaaa acacaggctc aagcagctta gagaagattg atctgtgttc act 6473 <211>7045 <212>DNA <213>Cricetulus griseus '<400>2 (SEQ ID NO: 6) actagcgtgc aattcagagg tgggtgaaga taaaaggcaa acatttgagg ccatttcctt 60<{ atttggcacg gcacttagga agtggaacat gcctaatcta ctggtttgta ccacctttcc 120 ctataatgga ctgtttggga agctcctggg caaccgattc tggcatctca ttggtcagag 180 gcctgttaaa tggtactctt atttgcaaag aaggctgtaa cttgtagctt taaaagcctc 240 tcctcaagaa agaagggaga aaggatatgg ctagacatat ctaatagact taaccactgt 3'00 gaaaagcctt agtatgaatc agatagaacc tatttttaac tcagttttga aaaaaataat 360 ctttatattt atttgtgtgt gtgtgtgtgt gtgtgtgtgt gtgtgtgtgt gtgtgtgtgt 420 gaaccacatg tagcaggtgc tggaggaggc cagaagaggg caccagatct cctggaactg 480 acaccacaca tggttatgag ctgcctgatg tgggtgctgg gaactgaact ctcgtgttct 540 gcaagagcag caactgttct cttaactgat gagccatctc tccagccccc cccataattt 600 taattgttca ttttagtaaa ttttattcat aatcaattat cacagtataa aacaatgatt 660 ttatatatat catatacata tcaaggatga cagtgagggg gatatgtgtg tgtgtgtgtg 720 tgtgtgtgtg tgtgtgtgtg tgtgttattt gtgtgtgtgc tttttaagaa ggtgccatag 780 tcactgcatt tctctgaagg atttcaaagg aatgagacat gtctgtctgc caggaaccct 840 atcttcctct ttgggaatct gacccaaatg aggtattctg aggaactgaa tgaagagctc 900 aagtagcagt gtcttaaacc caaatgtgct gtctagagaa agtcaacgtc atcagtgagc 960 tgaggagaga tttactgagc ggaagacaag cgctctttga tttaagtggc tcgaacagtc 1020 acggctgtgg agtggagcct gtgctcaggt ctgaggcagt ctttgctagc cagctgtgat 1080 gagcagtgaa gaaagggtgg agatggaggc agggtgggag cagggctatg gttcagacta 1140 ggtatcgtga gcacaccagc tggttgactt gtggtctgtg ggtcaggcgt tgtaaacgcc 1200 ctcagggtca ggcagtcaca ttgcttgaag ctgaatgggt gaggcaacac agagagtgca 1260 aagaaggcaa agtaccacct cttccccgac ccaggtcact tctggggtat agctgagact 1320 ccggacagca tgcaaccagc tggttagagc ttcagggaaa acttgatgtc tgcatgttgc 1380 tatgaaatgt gattcggtac atctggagaa aatttataat gctggctcag tcaagcactg 1440 aaaaggta ccttggcttt gggagctaca tgacattgac ttgtaggcag actttttttt 1500 ttctgcccgc caattcccag ataaccaata tggaggctca atattaatta taaatgctcg 1560 gctgatagct caggcttgtt actagctaac tcttccaact taaatgaacc catttctatt 1620 atctacattc tgccacgtga cttaccttg tacttcctgt ttcctctcct tgtctgactc 1680 tgcccttctg cttcccagag tccttagtct ggttctcctg cctaacctta tcctgcccag 1740 ctgctgacca agcatttata attaatatta agtctcccag tgagactctc atccagggag 1800 gacttgggtg ctcccccctc ctcattgcca tccgtgtctt cctcttccct cgcttccccc 1860 tcctcttcct gctcttcctc ctccacccct cctttcatag tattgatggc aagggtgttc 1920 tagaatggag gagtgcccat aggcatgcaa agaaaccagt taggatgctc tgtgaggggt 1980 tgtaatcata agcgatggac acaattcaag ccacagagtg aagacggaag gatgcactgt 2040 gctctagagc aacttctggg gcagaatcac agggtgagtt tctgacttga gggcgaagag 2100 gccacgagga agggagtgag tttgtctgag ctagaagcta cggcccacct cttggtagca 2160 gacctgccca caagcatgct ttgttaatca tgtgggatct gattttcctc taaatctatg 2220 ttcaactctt aagaaaatgt gaattctcac attaaaattt agatatacgt cttttggtgg 2280 ggggggtgta aaaaatcctc aagaatatgg atttctgggg gccggagaga tggctcagag 2340 gttaagagaa ctggttgctc ttctagacat tctgagttca attcccagca accacatggt 2400 ggctcacaac catctgtaat gcgacctggt gccatcttct gacatgcatg gatacatgca 2460 ggcagaaagc tgtatacata gtaaattgat aaatcttttt ttaaaaagag tatggattct 2520 gccgggtgtt ggtggcgcac gcctttaatc ccagcactct ggaggcagag gcaggtggat 2580 ctctgtgagt tcgagaccag cctggtctat aagagctagt tccaggacag cctccaaagc 2640 cacagaaa ccctgtctcg aaaaaccaaa aaaaaaaaaa aaaaaaaaaa aaaaaaaga 2700 gtatggattc taagaaagcc gtaacagctg gagctgtgta cggagttcag cgtggtacta 2760 gaacaaga cattcatgat gaacacccc aggattttta cttagtatct agtttccatt 2820 gttgttttga gaccggctct tatgctctcc aggctggcct caaactgctg atcttcccgc 2880 ctctacctct caagtcctgg gactacttgg ctcataaaac agttttgtc gggctccctg 2940 aagttatggt tgtacaaacc gtgggggtca atatactcac ttgggcagag agaagaggtc 3000 tgaatcccag acaatgactg catctcagga cagttgggaa gaggacaatg ccagaaggac 3060 ttagaaaaga tagactggag ggtggaaaag cagcaggaac agaaacaa aacaggaagc 3120 ttgctatcca gggccactct ggagtcctgt ggcaagatgg aagcgggcta ggggaataca 3180 tttgtgctac tgtgtgtgtg tgtgtgtgtg tgtgtgtgtg tgtgtgtgat caatgcctat 3240 caatgttgaa ggggaaatat gtataccaca ttgattctgg gagcaattct cagtatctgg 3300 cctagagaaa ggaatggcccc ctgcagaata gacagagtga atggtgccct ttatcatttg 3360 ctaaagtgaa ggagaaataa acatccttcc atagagtttc aggtaaatga accccacagt 3420 tcatctgtgc cgtggtggag gcctggccaa cagttaaaaa gattagacac ggacaaagtc 3480 tgaaggaaac acctcgaata ggaagaggag agccacctca ttctgtaact ttcctcaagg 3540 ggaagatgtt ccaagagtgg gaataaatgg tcaaaggggg gatttttaat taggaaaacg 3600 atttcctgta tcacttgtga aactggaggt tgatttgggg cataggacaa tagatttgat 3660 gctttgcaaa aagctgtttc aaagcagaga aatggaatag agacaattat gtagcgagga 3720 gggagggtgg ggcgaagatg gagacagaga agtggaagct gactttaggg aagaggaaca 3780 tagaccacag gggcggggcg gggggcaggg gcggggggcg gggctcaaag gaggcagtgg 3840 gaacgttgct agtgtcgca gcgtaagcgt gaatgtgcaa gcgtctttgt ggtgtgtgac 3900 caggagtagc gtggctggct tgtgtgctgc ttgtaatccc agtctttgag gtttccacac 3960 tgttccacag tgggtgtgat tttccctcgg agagcatgag ggctctgctt tccccacatc 4020 ctccccagcg ttcgttggta tttgtttcca agatgttagt gggtgagaca aagcctctct 4080 gttgatttgc ctttaacagg tgacaaaaaa agctcaacca ggagacattt ttgccttctt 4140 ggaaggtaat gctcccatgt agagcaatgg gacccatctc taaggtgagg ctactcttgc 4200 agtttgcacc cagctcttct gatgcaggaa ggaagttggt gggcaagcaa gactgtttgc 4260 ttcttgcgat ggacacattc tgcacacaaa ggctcaggag gggagaaggc tgtttgatgt 4320 ttagcactca ggaaggcccc tgatgcatct gtgattagct gtctccatct gtggagcaga 4380 cacggactaa ctaaaaacca gtgtttttaa attgtcaagc ctttaaggtg aggaaattga 4440 cttattgtgc tgggccatac gtagagcaag tgctctgcat tgggccaacc cccggctctg 4500 gtttctaggc accagaatgg cctagaacta actcacaatc ctcccattcc aggtctcagg 4560 tgctagaatg aaccactata ccagcctgcc tgcctgccta cctgccttcc taaattttaa 4620 atcatgggga gtaggggaga atacacttat cttagttagg gtttctattg ctgtgaagag 4680 acaccatgag catggcaact cttataaagg aaacattta gttgggtggc agtttcagag 4740 gttttagtac attgtcatca tggctgggaa catgatggca tgcagacaga catggtgctg 4800 gagaaaggga tgagagtcct acatcttgca ggcaacagga cctcagctga gacactggct 4860 ggtaccctga gcataggaaa cctcacagcc caccctcaca gtgacatatt tccttcaaca 4920 aagccatacc tcctaatagt gccactccct atgagatgac agggccaatt acattcaaac 4980 tgctataaca ctttaaagta ttttattttt attattgtaa attatgtatg tagctgggtg 5040 gtggcagccg aggtgcacgc cttaatccc agcacttggg aggcagaggc agatggatct 5100 ctgtgagttc aagaccagcc tggtctataa gagctagttg caaggaagga tatacaaaga 5160 acagttctag gatagccttc aaagccacag agaagtgctg tcttgaaaac caaaaattgt 5220 gctgggacct gtctctgctt tggttgcttc ccactccccc agagctggac tcttggtcaa 5280 cactgaatca gctgcaaaat aaactcctgg attcctctct tgtaacagga gcccgaagtc 5340 aggcgcccac ttgtcttctc gcaggattgc catagacttt ttctgtgtgc ccaccattcc 5400 agactgaagt agagatggca gtggcagaga ctgggaaggc tgcaacgaaa acaggaagtt 5460 attgcaccct gggaatagtc tggaaatgaa gcttcaaaac ttgcttcatg ttcagttgta 5520 cacagactca ctcccaggtt gactcacacg tgtaaatatt cctgactatg tctgcactgc 5580 ttttatctga tgcttccttc ccaaaatgcc aagtgtacaa ggtgagggaa tcacccttgg 5640 attcagagcc cagggtcgtc ctccttaacc tggacttgtc tttctccggc agcctctgac 5700 acccctcccc ccattttctc tatcagaagg tctgagcaga gttggggcac gctcatgtcc 5760 tgatacactc cttgtcttcc tgaagatcta acttctgacc cagaaagatg gctaaggtgg 5820 tgaagtgttt gacatgaaga cttggtctta agaactggag caggggaaaa aagtcggatg 5880 tggcagcatg tacccgaaat cccagaactg gggaggtaga gacggatgag tgcccggggc 5940 tagctggctg ctcagccagc ctagctgaat tgccaaattc caactcctat tgaaaaacct 6000 ttaccaaaca aacaaacaaa caaataataa caacaacaac aacaacaaac taccccatac 6060 aaggtgggcg gctcttggct cttgaggaat gactcaccca aacccaaagc ttgccacagc 6120 tgttctctgg cctaaatggg gtgggggtgg ggcagagaca gagacagaga gagacatgac 6180 ttcctgggct gggctgtgtg ctctaggcca ccaggaactt tcctgtcttg ctctctgtct 6240 ggcacagcca gagcaccagc acccagcagg tgcacacacc tccctccgtg cttcttgagc 6300 aaacacaggt gccttggtct gtctattgaa ccggagtaag ttcttgcaga tgtatgcatg 6360 gaaacaacat tgtcctggtt ttatttctac tgttgtgata aaaaccgggg aactccagga 6420 agcagctgag gcagaggcaa atgcaaggaa tgctgcctcc tagcttgctc cccatggctt 6480 gccgggcctg ctttctgcaa gcccttctct ccccattggc atgcctgaca tgaacagcgt 6540 ttgaaatgct ctcaaatgtc actttcaaag aaggcttctc tgatcttgct aactaaatca 6600 gaccatgttt caccgtgcat tatctttctg ctgtctgtct gtctgtctgt ctgtctatct 6660 gtctatcatc tatcaatcat ctatctatct atcttctatt tatctaccta tcattcaatc 6720 atctatcttc taactagtta tcatttattt atttgtttac ttactttttt tatttgagac 6780 agtatttctc tgagtgacag ccttggctgt cctggaaccc attctgtaac caggctgtcc 6840 tcaaactcac agagatccaa ctgcctctgc ctctctggtg ctggggttaa agacgtgcac 6900 caccaacgcc ccgctctatc atctatttat gtacttatta ttcagtcatt atctatcctc 6960 taactatcca tcatctgtct atccatcatc tatctatcta tctatctatc tatctatcta 7020 tctatcatcc atctataatc aattg 7045 <211>6473 <212>DNA <213>Cricetulus griseus <400>3 (SEQ ID NO: 7) agtgaacaca gatcaatctt ctctaagctg cttgagcctg tgttttcccg ttatacacag 60 gtaattggtg tgctgttaaa agctacttag aataaatgaa gaagaaaggg agaaggaggc 120 agaggagaag gagcaaagaa agaaggaaag ggggagggag ggagaggagg gagggaggga 180 gggagggagg gagaggaggg gagggggaga aagaagagaa ggagagatct tttccccact 240 gactatctca ggaaattacc acaggtggaa gggggtacta attaaggaat agctgtaagt 300 aacattattt ttattcgtag aacctcataa ctcttaagat gtgcttttta cccttttctg 360 ccttttagca caaataaact ccaacatgaa aattatccac tgtgcgtgtg aaaataccta 420 cacagagttc tgaatcattt gccaaattca agccccaatt tttatttcca ttttgactga 480 gagcaagatg ttccttttag gggatggaag cgtctgggtt tcccacactg aatgactcaa 540 ctcgaatgtt gcctcattaa cattctcgat ttttccgtaa tctctgctcc atgcattcaa 600 gataactgtg cctatcacaa atggcttttt agcagctcca ctctttctgg ggatgtggtg 660 gcccttccag tagctgccac cacggattgt cttcaatttc tcacttgttt ttgagttgag 720 tgtcagcctg acccctgggc atggccgcac atgactcagg caaagtgaga gtttcatcac 780 taaacgtggc tctgtttgct atgtctgttt tccctctaag agcaggttat tcaaatacca 840 tctggctgag gtcaagttgc ctcagagccc acagaatctc tacccaggtc cctgttggat 900 ccctaaaaac tcagtcatgc tgtaatctcc ttctgaaact gtgcaatgcc tgcaggctgt 960 cagcccagct ctctccttct gcttcctgtc ctcctaggac cccatgcctc ctcaaacgtc 1020 cacgtgtttc ttgctcctcc accacggttg ccaagccaaa attcgggtgg gcgggaggac 1080 attttcccaa gtgcctgttt cccttctttt ccttttgaca ccccagataa atcatctttc 1140 ccaatccaac acagccccac tgtgtctttg gggacttcat ccaatcacc aggaatgtat 1200 ccttagaaac aaaaatgca aacccagaac accaggac attaaagaa atttcactg 1260 gtgaggtcac aagtagtaga gactcttgt taacgggcag aactttcac ggacccagca 1320 tgctactgtg gcagttctgc aacaagctga aaatgccttt cccgaccacc caagccagtg 1380 ccacacaag gccaccttag ggtgtgcaca ggatgtcact aggcttggc ggaactcagg 1440 aaggagtctg aatttctcc cgttctctcc ttccctctc atccctc ttagctctg 1500 tctctcttc ctctctctcg tccccccct tcctccctcc cttcctgtg cagggccaca 1560 gatggaccgg gagacctca gcatgtcaa tcacaac gctctaccac tcaccac 1620 cctcgcctgc attgttacta ctacttat tatcttgata caggtctcca cattgagctc 1680 accctcacag tctccacatt gagctcaccc tcacagtc cacattgagc tcacccac 1740 agtctccaca ttgagctcac cctcacagtc tccattga gctcaccctc acagtctcca cattgagctc accctgtggc tctgcaac cttgaatttct ctcattccctc ctgcctcagc 1860 ctctggggtc gtggggatta gccaaaccca cttcaatca gcaatttctt 1920 agcgttcaat tcacacac tcatactcc agtactttgg aaccggaac aggagaattt 1980 ctgtgagctg gaggctagct tggactacag tatgagaccc tgtctctaaa taaatacaca 2040 aagaaatctc accaagggcc tccctctc agcaagctct aactgtgtg ggagttctgg 2100 gttgttccag ttaacggct cagaactcta ctgcccagca catcagcccc tagacacagg 2160 tggctctcta catgtgaaca tgcagtcaca gaaatgaat aaagtgaaa ttttattct 2220 tcagttgtat agcccttcc gtgtgggctg tagttactt cttgaatagg atggctcag 2280 aatccttggt gctggaacca agagttgat tccattagac gagggaat ataatgccca 2340 ataggcatt cctcctcccg gtcactagcg gtgcactttc tccgaatct tgtcatgttg 2400 aattagaaaa gttagtattt tcctccatcc cttcccctcc tccctccccctccccc 2460 ctcctcccct cctccctccg tctccccgcc cctcccctcc ccctgatc ctccccctc 2520 tatcaatcc aagaattcca gtaaaaagg gaaaacaatc gaagtgattt cgttgattgt 2580 cagttccacc aaagcaagac ttgactttag ttccgcgttt cggttcccgg catgcaccac 2640 agccagcgag caccgtggaa ggatgctagc acggtcctcc ccccgccccc actagctgtc 2700 ttcagctccc cagtagaggg caaccgcact ccagattctc aatggagagt gtttacacaa 2760 tcgttgcggg tttgtgtgag cgcgcccgct tccagagaca cttcttcttt ttcttttttc 2820 catttcatcc cagtggcaac gcagagtgcc agatcattca ggccgtttgc agggcaagcc 2880 gtgggagctt ggcaagcaag gccccatttc ctagggaacc cgtgcctggc gcttcaggaa 2940 agcacgggaa cctggcactg tgactctgcg ggtattattt tgcagaactc tttattaaac 3000 gggagtttca agtccagctg gagacgacca ggcagcgcct ttaaccccag agtcacacac 3060 aggtgccttt tcttggggcc agattggggt tgtgtggcag acctgcgacc agcttgacaa 3120 ctcttctgcc aggccacaaa atggtgttgg ctgtaagagg tgacaccagg gacagggaag 3180 atcgctgcta ttctcctgag ctctccaaag acccacacca gtctgtcccc ctttcctcct 3240 gctcttcccc tgtatcgccc cctcaccatc tcccccaacg agactcttgg catctcctcg 3300 gcacaaggat ttgaaatag atgcttgggg gtgagagaa gagagagaa aggagagaa 3360 ggaaggaagg atatatagat gatacagacg catacaggtg acatgtagct atcattttt 3420 aattaaaaa taaatttaaa gcaatcaag gatatatg ataccttag agcaagtgtc 3480 tcatacacac acacacac acacacata father father father father 3540 tatatatata ttatactgg aacaagtgtc cagaagggct ggggactcta aagtgcttgt 3600 caaagccagg ctcacatcag taatcttac acctggtaga ctgagacagg aggattttga 3660 tgagttcagg cccagcctga gctgcagaat gtgatctat cccaaaaaag taaaataaaa 3720 taaaattcaa atacacgaa aagagtattt gctgaacaa caagcctaaa gccctggatc 3780 ccttccccca tgtcctaaga aaataagttt cttgaagctg gagggatggc tcagaggtta 3840 agagccccag ctgcacttgc ggaacactaa gacccagttc ccagacccca cactgtggtt 3900 cacaactgtc tcaacgcca gctccggagg atccatgccc tctcctggcc tccaccggc 3960 catacaacac atacagtgcc catacattta tgcaagcaag gtattcacgc actaaaact 4020 aaaagaatat ttaataaaga tataacaaaa tagcatgaag cccagctggt acagaggttc 4080 aaactacatc ccaggttcat ccctctgcct ttgctctcag ttggcttggg taggtctctt 4140 ctctgaactg gcgccctgcg ggttccacat tgagaccctc tcatttttaa acctacttct 4200 tctgggcggg gttaattgct gccagggctc aagccaacgc ttcctcttct ccacagcaat 4260 cttccaagtt tcacgagata accaggaact gctaagttca tgtgaacctt agtgaagaac 4320 ctgagtcttc ccatgtgatt ggtgtgtgca tgtgtgcata cacaaatgta tgtgtgtgct 4380 ctatgtgtgc ctatgtatgt gtgcatgcat gtgtgcatat acaaatgcat atatgtctat 4440 gtagtgtgcg tacacaaatg tatgtgtgtg ctcaatgtgt gcctatgtgt gtgtatgcat 4500 gtgtgcgtac acaatgcatg tgtgtggtgt ctgtgtgcct gtgtgtgtat gcatgtatgc 4560 atacacaaat gtatatgtgt ggtgtgtgaa tgtgtgccta tgtatgtgtg tgctgtgtgt 4620 gggtgtggta tgtgtgtgat gtgtggaggg gtgtgtatgt gtggtatgta taggtgatac 4680 gtttggggtg taatatgcgt atgtggtttg tgaaatgtag ttcgtgtgtg tgcatgtgtg 4740 cgtgcgtgcg tgcgtgcgtg cgtgtgtgtg tgtgtgtgtg tgtgtgtgtt 4800 tgtgtgaggt gtgtgtactc accatggcct ccctcacttg ggggagtgaa gtcagcagcc 4860 tggaccactc agggacatga gatactcaga cacatcttga tttccacccc tcttttcctg 4920 atcctccttc acgtgtcact ttcccaaaca ctggacaaca gtttggggc atctgattcc 4980 actaatgaca gggacatcac atgtctccag agggaacacc ttctgtgtca catgtcatct 5040 gagaatgtag cagagtcaca gagaaatgtc acagaaacca aatgcagag taccaaggta 5100 tagctaggca cagagcagag gggaagccgc tgaatttatt aaaaatgtca gaatcgtaaaa 5160 agacaggggga cagcggtggg gacattcagg gtccagtagc acacaggcag tccaaacctg 5220 atcactggaa ggtagtaggt aaggaaaggc tgcacacaga ttattcacac agtttataca 5280 tgtacacaga ttatcacat ggtttgtgta tgtgcacaga ttattcacac agtttataca 5340 tgtgtggctt cgtggtaact ttgagcttac tttcaattta aaaggatctc tctcacaagc 5400 tggggccggg aatggctgca gtcaacactc catcacttag tcacactgtg caaacagcac 5460 ctcctgactc atggtgactt gtagtaaaat gaagaggcca catttgcatc caagacagct 5520 catcagtacc tagtgagaa tctgtccctg agtatttgca tgaatggacc cgggtccagg 5580 gcctggctgg gagtctccag gtgttgcagc cagaatgtca ttgtgttttt tcaggatccc 5640 5700 agtgattagg taacaaag ccttttaaaa accgagatca cccttgtcat gtccctggcc 5760 tcttagaaca agatccaagc ttttgctggt tgacaagtgg ggccatccag tgcgtctccg 5820 ttcctgctac ttcatctgga agacctctcc cactaacttg cccctgaccc ctcacaccctg 5880 ctgtttcctt tccacccgga agtgcttgtc taggctttca tggccatctg actgagcatc 5940 taggcctcag tccagtggtc cctcagctct ctctagtcac tgtactaatg gaacggcca 6000 ctaactacat tttcaatatg gaagcctcct cctcaggaac ctccaagggc agaagcctcc 6060 agagaaccac tcctgacccc ctggagttct gagtgcttct ggccctctct gtgtctgcag 6120 gactattcac cacttgtgtt gaatggttca gtcctcacct cctctggcat gtgctcagtt 6180 ctcatctcat tggggagtcc ttcccaggtc actctctct cctgtctttg aagtgttttt 6240 ttccttcatg gtatttctgt ctgggcacac acatacacac acatacacac 6300 ccatgcagta tggcagatac atcacctag tttcagattt ttattctacc atcaccaat 6360 acctgaatcc ccgaaaaagc cttagaaagc caggaatttg tgtatttg tcagcactcc 6420 acccagcac ctgaagccaa gcctgactta atatttgg tttgttttct aga 6473 <211> 7045 <212> DNA <213> Gray Rhinoceros <400> 4 (8) caattgatta tagatag agatagatagatagatagatagatagatagatagatagata 60 tggatagaca gatgatgat agttagaga tagataatga ctgaataata agtacataaa 120 tagatgatag agcggggcgt tggtggtgca cgtctttaac cccagcacca gagaggcaga 180 ggcagttgga tctctgtgag tttgaggaca gctgttac agaatggtt ccaggacagc 240 aaggctgtc actcagagaa atactgtctc aaaaaaaa agtaagtaaa caataata 300 aatgatact aggataga taggattg atgatagt agataatag agatagata 360 tgagatt tgagatga tgahagagagagagagagagagagagagagagagagagagagagaga 420 agataatgca cggtgaaca tggtctgatt tagttagca gatcagagaa gccttctttg 480 aaagtgacat ttgagagcat ttcaacgct gttcatgtca ggcatgccaa tggggagaga 540 agggcttgca gaaagcaggc ccggcaagcc atgggagca agctaggagg cagcattcct 600 tgcatttgcc tctgcctcag ctgcttcctg gagttccccg gttttatca caacagtaga 660 aaaaacca ggacaatgtt gttccatgc atacatctgc aagaacttac tccggttca 720 tagacagacc aaggcacctg tgttgctca agaagcacgg agggaggtgtgcacctgc 780 tgggtgctgg tgctctggct gtgccagaca gagagcaga caggaagtt cctggtggcc 840 tagcacac agcccagccc aggagtcat gtctctctct gtctctgtct ctgccccacc 900 cccacccat ttaggccaga gaacagctgt ggcaagcttt gggtttggtt gagtcattcc 960 tcaagagcca agagccgccc accttgtatg gggtagtttg tgttgttgt tgttgttatt 1020 atttgtttgt ttgtttgttt ggtaaaggtt tttcaatagg agttggaatt tggcaattca gctaggctgg ctgagcagcc agctagcccc gggcactcat ccgtctctac ctccccagtt 1140 ctgggatttc gggtacatgc tgccacatcc gacttttttc ccctgctcca gttcttaaga ccaagtcttc atgtcaaaca cttcaccacc ttagccatct ttctgggtca gaagttagat cttcaggaag acaaggagtg tatcaggaca tgagcgtgcc ccaactctgc tcagaccttc tgatagagaa aatgggggga ggggtgtcag aggctgccgg agaaagacaa gtccaggtta aggaggacga ccctgggctc tgaatccaag ggtgattccc tcaccttgta cacttggcat tttgggaagg aagcatcaga taaaagcagt gcagacatag tcaggaatat ttacacgtgt gagtcaacct gggagtgagt ctgtgtacaa ctgaacatga agcaagtttt gaagcttcat ttccagacta ttcccagggt gcaataactt cctgttttcg ttgcagcctt cccagtctct 1620 gccactgcca tctctacttc agtctggat ggtgggcaca cagaaaaagt ctatggcaat cctgcgagaa gacaagtggg cgcctgactt cgggctcctg ttacaagaga ggaatccagg agtttatttt gcagctgatt cagtgttgac caagagtcca gctctggggg agtgggaagc 1800 aaccaaagca gagacaggtc ccagcacaat ttttggtttt caagacagca cttctctgtg 1860 gctttgaagg ctatcctaga actgttcttt gtatatcctt ccttgcaact agctcttata 1920 gaccaggctg gtcttgaact cacagagatc catctgcctc tgcctcccaa gtgctgggat 1980 taaaggcgtg cacctcggct gccaccaccc agctacatac ataatttaca ataataaaa 2040 taaaatactt taaagtgtta tagcagtttg aatgtaattg gccctgtcat ctcataggga 2100 gtggcactat taggaggtat ggctttgttg aaggaaatat gtcactgtga gggtgggctg 2160 tgaggtttcc tatgctcagg gtaccagcca gtgtctcagc tgaggtcctg ttgcctgcaa 2220 gatgtaggac tctcatccct ttctccagca ccatgtctgt ctgcatgcca tcatgttccc 2280 agccatgatg acaatgtact aaaacctctg aaactgccac ccaactaaat gttttccttt 2340 ataagagttg ccatgctcat ggtgtctctt cacagcaata gaaaccctaa ctaagataag 2400 tgtattctcc cctactcccc atgatttaaa atttaggaag gcaggtaggc aggcaggcag 2460 gctggtatag tggttcattc tagcacctga gacctggaat gggaggattg tgagttagtt 2520 ctaggccatt ctggtgccta gaaaccagag ccgggggttg gcccaatgca gagcacttgc 2580 tctacgtatg gcccagcaca ataagtcaat ttcctcacct taaaggcttg acaatttaaa 2640 aacactggtt tttagttagt ccgtgtctgc tccacagatg gagacagcta atcacagatg 2700 catcaggggc cttcctgagt gctaaacatc aaacagcctt ctcccctcct gagcctttgt 2760 gtgcagaatg tgtccatcgc aagaagcaaa cagtcttgct tgcccaccaa cttccttcct 2820 gcatcagaag agctgggtgc aaactgcaag agtagcctca ccttagagat gggtcccatt 2880 gctctacatg ggagcattac cttccaagaa ggcaaaaatg tctcctggtt gagctttttt 2940 tgtcacctgt taaaggcaaa tcaacagaga ggctttgtct cacccactaa catcttggaa 3000 acaaatacca acgaacgctg gggaggatgt ggggaaagca gagccctcat gctctccgag 3060 ggaaaatcac acccactgtg gaacagtgtg gaaacctcaa agactgggat tacaagcagc 3120 acacaagcca gccacgctac tcctggtcac acaccacaaa gacgcttgca cattcacgct 3180 tacgctgcga acactagcaa cgttcccact gcctcctttg agccccgccc cccgcccctg 3240 ccccccgccc cgcccctgtg gtctatgttc ctcttcccta aagtcagctt ccacttctct 3300 gtctccatct tcgccccacc ctccctcctc gctacataat tgtctctatt ccatttctct 3360 gctttgaaac agctttttgc aaagcatcaa atctattgtc ctatgcccca aatcaacctc 3420 cagtttcaca agtgatacag gaaatcgttt tcctaattaa aaatcccccc tttgaccatt 3480 tattcccact cttggaacat cttccccttg aggaaagtta cagaatgagg tggctctcct 3540 cttcctattc gaggtgtttc cttcagactt tgtccgtgtc taatcttttt aactgttggc 3600 caggcctcca ccacggcaca gatgaactgt ggggttcatt tacctgaaac tctatggaag 3660 gatgtttatt tctccttcac tttagcaaat gataaagggc accattcact ctgtctattc 3720 tgcaggggcc attcctttct ctaggccaga tactgagaat tgctcccaga atcaatgtgg 3780 tatacatatt tccccttcaa cattgatagg cattgatcac acacacacac acacacacac 3840 acacacacac acacagtagc acaaatgtat tcccctagcc cgcttccatc ttgccacagg 3900 actccagagt ggccctggat agcaagcttc ctgttttgtt tctctgttcc tgctgctttt 3960 ccaccctcca gtctatcttt tctaagtcct tctgccattg tcctcttccc aactgtcctg 4020 agatgcagtc attgtctggg attcagacct tctctctctg cccaagtgag tatattgacc 4080 cccacggttt gtacaaccat aacttcaggg agcccgacaa aaactgtttt atgagccaag 4140 tagtcccagg acttgagagg tagaggcggg aagatcagca gtttgaggcc agcctggaga 4200 gcataagagc cggtctcaaa acaacaatgg aaactagata ctaagtaaaa atcctggggt 4260 gtttcatcat gaatgtctgt tcttctagta ccacgctgaa ctccgtacac agctccagct 4320 gttacggctt tcttagaatc catactcttt tttttttttt tttttttttt ttttttttgg 4380 tttttcgaga cagggtttct ctgtggcttt ggaggctgtc ctggaactag ctcttataga 4440 ccaggctggt ctcgaactca cagagatcca cctgcctctg cctccagagt gctgggatta 4500 aaggcgtgcg ccaccaacac ccggcagaat ccatactctt tttaaaaaaa gatttatcaa 4560 tttactatgt atacagcttt ctgcctgcat gtatccatgc atgtcagaag atggcaccag 4620 gtcgcattac agatggttgt gagccaccat gtggttgctg ggaattgaac tcagaatgtc 4680 tagaagagca accagttctc ttaacctctg agccatctct ccggccccca gaaatccata 4740 ttcttgagga ttttttacac cccccccacc aaaagacgta tatctaaatt ttaatgtgag 4800 aattcacatt ttcttaagag ttgaacatag atttagagga aaatcagatc ccacatgatt 4860 aacaaagcat gcttgtgggc aggtctgcta ccaagaggtg ggccgtagct tctagctcag 4920 acaaactcac tcccttcctc gtggcctctt cgccctcaag tcagaaactc accctgtgat 4980 tctgccccag aagttgctct agagcacagt gcatccttcc gtcttcactc tgtggcttga 5040 attgtgtcca tcgcttatga ttacaacccc tcacagagca tcctaactgg tttcttttgca 5100 tgcctatggg cactcctcca ttctagaaca cccttgccat caatactatg aaaggagggg 5160 tggaggagga agagcaggaa gaggaggggg aagcgaggga agaggaagac acggatggca 5220 atgaggaggg gggagcaccc aagtcctccc tggatgagag tctcactggg agacttaata 5280 ttaattataa atgcttggtc agcagctggg caggataagg ttaggcagga gaaccagact 5340 aggactctg ggaagcagaa gggcagagtc agacaaggag aggaaacagg aagtacaagg 5400 taaagtcacg tggcagaatg tagaatag aaatgggttc atttaagttg gaagagttag 5460 ctagtaacaa gcctgagcta tcagccgagc atttataatt aatattgagc ctccatattg 5520 gttatctggg aattggcggg cagaaaaaaa aaagtctgcc tacaagtcaa tgtcatgtag 5580 ctcccaaagc caaggtacct ttgttcagtg cttgactgag ccagcattat aaattttctc 5640 cagatgtacc gaatcacatt tcatagcaac atgcagacat caagttttcc ctgaagctct 5700 aaccagctgg ttgcatgctg tccggagtct cagctataac ccagaagtga cctgggtcgg 5760 ggaagaggtg gtactttgcc ttctttgcac tctctgtgtt gcctcaccca ttcagcttca 5820 agcaatgtga ctgcctgacc ctgagggcgt ttacaacgcc tgacccacag accacaagtc 5880 aaccagctgg tgtgctcacg atacctagtc tgaaccatag ccctgctccc accctgcctc 5940 catctccacc ctttcttcac tgctcatcac agctggctag caaagactgc ctcagacctg 6000 agcacaggct ccactccaca gccgtgactg ttcgagccac ttaaatcaaa gagcgcttgt 6060 cttccgctca gtaaatctct cctcagctca ctgatgacgt tgactctc tagacagcac 6120 atttgggttt agacactgc tacttgagct cttcattcag tcctcagaa tacctcattt 6180 gggtcagatt cccaaagagg aagatagggt tcctggcaga cagacatgtc tcattccttt 6240 gaaatccttc agagaaatgc agtgactatg gcaccttctt aaaagcaca cacacaata 6300 acacacac acacacac acacacac acacacac atatccccct cactgtcatc 6360 cttgatatgt atatgatata tataaaatca ttgttttata ctgtgataat tgattatgaa 6420 taaaatttac taaaatgac aattaaaatt atgggggggg ctggagagat ggctcatcag 6480 ttaagagaac agttgctgct ctgcagaac acgagagttc agttcccagc acccacatca 6540 ggcagctcat aaccatgtgt gtgtcagtt ccaggagatc tggtgccctc ttctggccctc 6600 ctccagcacc tgctacatgt ggttcacacacacacacacacacacacacacacacacacacacacacacacacacacacacacacacacacacaca 6660 cacacacaca caaataa window ttttcaaa actgagttaa aaataggttc 6720 tatctgattc atactaggc tttcacagt ggttaagtct attagatg tctagccata 6780 tcctttctcc cttctttctt gaggagaggc ttttaaagct acagttaca gccttctttg 6840 caataagag taccatta caggcctctg accaatgaga tgccagaatc ggttgcccag 6900 gagcttccca aacagtccat tatagggaaa ggtggtacaa accagtagat taggcatgtt 6960 ccacttccta agtgccgtgc CAAAgga aatggcctca atgtttgcc ttttacttc 7020 acccaccctct gattgcacg ctagt 7045 <211> 13515 <212> DNA <213> Gray Rhinoceros <400> 5 (9) tctagaaaca aaaccaaaaa tattaagtca ggcttggctt caggtgctgg ggtggagtgc 60 tgacaaaaat acacaaatttc ctggctttct aaggctttt cggggattca ggtattggtt 120 gatggtagaa taaaaatctg aaacataggt gatgtatctg ccatactgca tggtgtgta 180 tgtgtgtgta tgtgtgtctg tgtgtgtgcc cagacagaaa taccatgaag gaaaaaaaca 240 cttcaaagac aggagag agtgacctgg gaaggactcc ccaatgagat gagaactgag 300 cacatgccag aggaggtgag gactgaacca ttcaacaa gtggtgaata gtcctgcaga 360 420. cacagagagg gccagaagca ctcagaactc cagggggtca ggagtggttc tctggaggct tctgcccttg gaggttcctg aggaggaggc ttccatattg aaaatgtagt tagtggccgt 480 ttccattagt acagtgacta gagagagctg agggaccact ggactgaggc ctagatgctc agtcagatgg ccatgaaagc ctagacaagc acttccgggt ggaaaggaa cagcaggtgt gaggggtcag gggcaagtta gtgggagagg tcttccagat gaagtagcag gaacggagac 660 gcactggatg gccccacttg tcaaccagca aaagcttgga tcttgttcta agaggccagg 720 gacatgacaa gggtgatctc ggtttttaaa aggctttgtg ttacctaatc acttctatta 840. gtcagatact ttgtaacaca aatgagtact tggcctgtat tttagaaact tctgggatcc tgaaaaaaca caatgacatt ctggctgcaa cacctggaga ctcccagcca ggccctggac ccgggtccat tcatgcaaat actcagggac agattcttca ctaggtactg atgagctgtc ttggatgcaa atgtggcctc ttcattttac tacaagtcac catgagtcag gaggtgctgt ttgcacagtg tgactaagtg atggagtgtt gactgcagcc attcccggcc ccagcttgtg 1080 agagagatcc ttttaaattg aaagtaagct caaagttacc acgaagccac acatgtataa 1140 actgtgtgaa taatctgtgc acatacacaa accatgtgaa taatctgtgt acatgtataa 1200 actgtgtgaa taatctgtgt gcagcctttc cttacctact accttccagt gatcaggttt 1260 ggactgcctg tgtgctactg gaccctgaat gtccccaccg ctgtcccctg tcttttacga 1320 ttctgacatt tttaataaat tcagcggctt cccctctgct ctgtgcctag ctataccttg 1380 gtactctgca ttttggtttc tgtgacattt ctctgtgact ctgctacatt ctcagatgac 1440 atgtgacaca gaaggtgttc cctctggaga catgtgatgt ccctgtcatt agtggaatca 1500 gatgccccca aactgttgtc cagtgtttgg gaaagtgaca cgtgaaggag gatcaggaaa 1560 agaggggtgg aaatcaagat gtgtctgagt atctcatgtc cctgagtggt ccaggctgct 1620 gacttcactc ccccaagtga gggaggccat ggtgagtaca cacacctcac acatactata 1680 tccaacacac acacacacac acacacacac acgcacgcac gcacgcacgc acgcacacat 1740 gcacacacac gaactacatt tcacaaacca catacgcata ttacacccca aacgtatcac 1800 ctatacatac cacacataca cacccctcca cacacacacac cacacacacacacacacaca gcacacacat acataggcac acatcacac accacata tacatttgtg tatgcataca 1920 tgcatacaca cacaggcaca cagacaccac accatgcat tgtgtacgca cacatgcata 1980 cacacacata ggcacacatt gagcacacac atacatttgt gtacgcacac cacatagaca 2040 tatatgcatt tgtatatgca cacatgcatg cacacataca taggcaca tagcacac 2100 acatacattt gtgtatgcac acatgcacac accatcaca tgggagact caggttctc 2160 actaaggttc acatgaactt agcagttcct gttatctcg tgaacttgg agattgctg 2220 tggagagag gaagcgttgg cttgagccct ggcagcaat aaccccgccc agagaagta 2280 ggtttaaaaa tgagagggtc tcaatgtgga acccgcaggg cgccagttca gagagagac 2340 ctacccaagc caacgagag caaggcaga gggatgaacc tgggatgtag ttgaacctc 2400 tgtaccagct gggcttcatg ctattttgtt atatctttat taaatattct tttagtttta 2460 tgtgcgtgaa taccttgctt gcataaatgt atgggcactg tatgtgttct tggtgccggt 2520 ggaggccagg agagggcatg gatcctccgg agctggcgtt tgagacagtt gtgacccaca 2580 gtgtggggtc tgggaactgg gtcttagtgt tccgcaagtg cagctggggc tcttaacctc 2640 tgagccatcc ctccagcttc aagaaactta ttttcttagg acatggggga agggatccag 2700 ggctttaggc ttgtttgttc agcaaatact cttttcgtgt attttgaatt ttattttatt 2760 ttactttttt gggatagaat cacattctgc agctcaggct gggcctgaac tcatcaaaat 2820 cctcctgtct cagtctacca ggtgataaga ttactgatgt gagcctggct ttgacaagca 2880 ctttagagtc cccagccctt ctggacactt gttccaagta taatatatat atatatatat 2940 atatatatat atatatatat atatattgtg tgtgtgtgtt tgtgtgtgta tgagacactt 3000 gctctaaggg tatcatatat atccttgatt tgcttttaat ttatttttta attaaaaatg 3060 attagctaca tgtcacctgt atgcgtctgt atcatctata tatccttcct tccttctctc 3120 tctttctctc ttcttcttct cacccccaag catctatttt caaatccttg tgccgaggag 3180 atgccaagag tctcgttggg ggagatggtg agggggcgat acaggggaag agcaggagga 3240 aagggggaca gactggtgtg ggtctttgga gagctcagga gaatagcagc gatcttccct 3300 gtccctggtg tcacctctta cagccaacac cattttgtgg cctggcagaa gagttgtcaa 3360 gctggtcgca ggtctgccac acaaccccaa tctggcccca agaaaaggca cctgtgtgtg 3420 actctggggt taaaggcgct gcctggtcgt ctccagctgg acttgaaact cccgtttaat 3480 aaagagttct gcaaaataat acccgcagag tcacagtgcc aggttcccgt gctttcctga 3540 agcgccaggc acgggttccc taggaaatgg ggccttgctt gccaagctcc cacggcttgc 3600 cctgcaaacg gcctgaatga tctggcactc tgcgttgcca ctgggatgaa atggaaaaaa 3660 gaaaaagaag aagtgtctct ggaagcgggc gcgctcacac aaacccgcaa cgattgtgta 3720 aacactctcc attgagaatc tggagtgcgg ttgccctcta ctggggagct gaagacagct 3780 agtgggggcg gggggaggac cgtgctagca tccttccacg gtgctcgctg gctgtggtgc 3840 atgccgggaa ccgaaacgcg gaactaaagt caagtcttgc tttggtggaa ctgacaatca 3900 acgaaatcac ttcgattgtt ttcctctttt tactggaatt cttggatttg atagatgggg 3960 gaggatcaga ggggagggg aggggagggg aggagggg aggagggagggagggagg 4020 ggggagggg ggaggagggg aagggatgga ggaaatact aacttttcta attcacatg 4080 acaaagattc ggagaagtg caccgctagt gaccgggagg aggaatgccc tattggggcat 4140 tatattccct gtcgtctaat ggaatcaaac tctgttcc agcaccaagg attctgagcc 4200 tatcctattc aagacagtaa ctacagccca cacggaagg gctatacaac tgaagaata 4260 aaatttcac tttatttcat ttctgtgact gcatgttcac atgtagagag cacctgtgt 4320 ctaggggctg atgtgctggg cagtagagtt ctgagcccgt taactggaac aacccagaac 4380 tcccaccaca gttagagctt gctgagagag ggaggcctt ggtgagattt ctttgtgtat 4440 ttatttagag acagggtctc atactgtagt ccaagctagc ctccagctca cagaaatttct 4500 cctgttccgg ttccaagt actggagtta tgagtgtgtg ttaattgaac gctaagaatt 4560 tgctgattga agaaaacctc aagtggtttt ggctaatccc cacgacccca gaggctgagg 4620 caggaggaat gagagaatttc aaggttgcc agagccacag ggtgagctca atgtggagac 4680 tgtgaggtg agctcaatgt ggagactgtg agggtgagct caatgtggag actgtgaggg 4740 tgagctcaat gtggagactg tgaggtgag ctcaatgtgg agactgtgag ggtgagctca 4800 atgtggagac ctgtatcaag ataatatag tagtagtac atgcaggcg agggtgtggt 4860 tgagtgtag agcagttagt tgattgaca tgcttgaggt ctcccggtcc atctgtggcc 4920 ctgcaacagg agggaggga gggagggggg gagagagagagagagag agahagaagc 4980 taggatagggg atgagag gagggaa acgggaa attcagactc cttcctgagt 5040 tccgccaacg cctagtgaca tcctgtgcac accctaggt ggccttgtg tgcactgc 5100 ttgggtggtc gggaaaggca tttcagctt gttgcagaac tgccacagta gcatgctggg 5160 tccgtgaaag ttctcccg ttaacaagaa gtcttacta cttgtgacct caccagtgaa 5220 aatttcttta attgtctcct gttgttctgg gttttgcatt ttgtttcta aggatacatt 5280 cctgggtgat gtcatgaagt ccccaagac acagtggggc tgtgttgat tgggaagat 5340 gatttatctg gggtgtcaa aggaaagaa gggaacagg cacttgggaa atgtcctcc 5400 cgcccacccg aattttggct tggcaccgt ggtggaggag cagaacac gtggacgttt 5460 gaggaggcat ggggtcctag gaggacagga agcagaagga gagagctggg ctgacagcct 5520 gcaggcattg cacagttca gaggagatt acagcatgac tgagtttta gggatccaac 5580 agggacctgg gtagagattc tgtgggctct gaggcaactt gacctcagcc agatgtatt 5640 tgaataacct gctcttagag ggaaaacaga catagcaac agagccacgt ttagtgag 5700 aactctcact ttgcctgagt catgtgcggc catgcccagg gtcaggctg acacctcact 5760 caaaaacaag tgagaattg agacaatcc gtggtggcag ctactggaag ggccaccaca 5820 tccccagaaa gagtggagct gctaaaaagc cattgtgat aggcacagtt atcttgaatg 5880 catggagcag agattacgga aaatcgaga atgttaatga ggcaacattc gagttgagtc 5940 attcagtgtg ggaaacccag acgctccat ccctaaag gaacatctg ctctcagtca 6000 aaatggaat aaaattggg gcttgaattt ggcaatgat tcagactct gtgtaggtat 6060 tttcacacgc acagtggata atttcatgt tggatttat tgtgctaaa aggcagaaaa 6120 gggtaaaag cacatcttaa gagttatgag gttctacgaa taaaataat gttacttaca 6180 gctattcctt aattagtacc cccttccacc tgtgtatt tcctgagata gtcagtgggg 6240 aaagatctc tccttctct ctccctcccc ctccccct ctccctccct ccctccctcc 6300 ctcctcctc tccctccctc tctttcttg ctccttctc tctgcctcct 6360 tctccctttc ttctcattt attctaagta gctttaaca gcacaccaat tacctgtgta 6420 taacgggaaa accaggctc aagcagctta gagagattg atctgtgttc actagcgtgc 6480 aattcagagg tgggtgaaga taaaaggcaa acatttgagg ccattcctt atttggcacg 6540 gcacttagga agtggaacat gcctaatcta ctgtttgta ccaccttttcc ctataatgga 6600 ctgtttggga agctcctggg caaccgattc tggcatctca tggtcagag gcctgttaaa 6660 tggtactctt atttgcaag aaggctgtaa cttgtagctt taaaagccctc tcctcaagaa 6720 agaagggaga aaggatatgg ctagacatat ctaatagact taaccactgt gaaagcctt 6780 agtatgaatc agatagaacc tattttac tcagttttga aaaaaataat ctttatattt 6840 atttgtgtgt gtgtgtgtgt gtgtgtgtgt gtgtgtgtgt gtgtgtgtgt gaaccacatg 6900 tagcaggtgc tggaggaggc cagaagaggg caccagatct cctggaactg acaccacaca 6960 tggttatgag ctgcctgatg tgggtgctgg gaactgaact ctcgtgttct gcaagagcag 7020 caactgttct cttaactgat gagccatctc tccagccccc cccataattt taattgttca 7080 ttttagtaaa ttttatcat aatcaattat cacagtataa aacaatgatt ttatatatatat 7140 catatacata tcaaggatga cagtgagggg gatatgtgtg tgtgtgtgtg tgtgtgtgtg 7200 tgtgtgtgtg tgtgttattt gtgtgtgtgc tttttaagaa ggtgccatag tcactgcatt 7260 tctctgaagg atttcaaagg aatgagacat gtctgtctgc caggaaccct atcttcctct 7320 ttgggaatct gacccaaatg aggtattctg aggaactgaa tgaagagctc aagtagcagt 7380 gtcttaaacc caaatgtgct gtctagagaa agtcaacgtc atcagtgagc tgaggagaga 7440 tttactgagc ggaagacaag cgctctttga tttaagtggc tcgaacagtc acggctgtgg 7500 agtggagcct gtgctcaggt ctgaggcagt ctttgctagc cagctgtgat gagcagtgaa 7560 gaaagggtgg agatggaggc agggtgggag cagggctatg gttcagacta ggtatcgtga 7620 gcacaccagc tggttgactt gtggtctgtg ggtcaggcgt tgtaaacgcc ctcagggtca 7680 ggcagtcaca ttgcttgaag ctgaatgggt gaggcaacac agagagtgca aagaaggcaa 7740 agtaccacct cttccccgac ccaggtcact tctgggttat agctgagact ccggacagca 7800 tgcaaccagc tggttagagc ttcagggaaa acttgatgtc tgcatgttgc tatgaaatgt 7860 gattcggtac atctggagaa aatttataat gctggctcag tcaagcactg aaaaggta 7920 ccttggcttt gggagctaca tgacattgac ttgtaggcag actttttttt ttctgcccgc 7980 caattcccag ataaccaata tggaggctca atattaatta taaatgctcg gctgatagct 8040 caggcttgtt actagctaac tcttccaact taaatgaacc catttctatt atctacattc 8100 tgccacgtga cttaccttg tacttcctgt ttcctctcct tgtctgactc tgcccttctg 8160 cttcccagag tccttagtct ggttctcctg cttaacctta tcctgcccag ctgctgacca 8220 agcatttata attaatatta agtctcccag tgagactctc atccagggag gacttgggtg 8280 ctccccccctc ctcattgcca tccgtgtctt cctcttccct cgcttccccc tcctcttcct 8340 gctcttctc ctccacccct cctttcatag tattgatggc aagggtgttc tagaatggag 8400 gagtgcccat aggcatgcaa agaaaccagt taggatgctc tgtgaggggt tgtaatcata 8460 agcgatggac acaattcaag ccacagagtg aagacggaag gatgcactgt gctctagagc 8520 aacttctggg gcagaatcac agggtgagtt tctgacttga gggcgaagag gccacgagga 8580 agggagtgag tttgtctgag ctagaagcta cggcccacct cttggtagca gacctgccca 8640 caagcatgct ttgttaatca tgtgggatct gattttcctc taaatctatg ttcaactctt 8700 aagaaaatgt gaattctcac attaaaattt agatatacgt cttttggtgg ggggggtgta 8760 aaaaatcctc aagaatatgg atttctgggg gccggagaga tggctcagag gttaagagaa 8820 ctggttgctc ttctagacat tctgagttca attcccagca accacatggt ggctcacaac 8880 catctgtaat gcgacctggt gccatcttct gacatgcatg gatacatgca ggcagaaagc 8940 tgtatacata gtaaattgat aaatcttttt ttaaaaagag tatggattct gccgggtgtt 9000 ggtggcgcac gcctttaatc ccagcactct ggaggcagag gcaggtggat ctctgtgagt 9060 tcgagaccag cctggtctat aagagctagt tccaggacag cctccaaagc cacagaagaa 9120 ccctgtctcg aaaaaccaaa aaaaaaaaaa aaaaaaaaaa aaaaaaaaga gtatggattc 9180 tagaaagcc gtaacagctg gagctgtgta cggagttcag cgtggtacta gaacaga 9240 cattcatgat gaaacacccc aggattttta cttagtatct agtttccatt gttgttttga 9300 gaccggctct tatgctctcc aggctggcct caaactgctg atcttcccgc ctctacctct 9360 caagtcctgg gactacttgg ctcataaaac agttttgtc gggctccctg aagttatggt 9420 tgtacaaacc gtgggggtca atatactcac ttgggcagag agaaaggtc tgaatcccag 9480 acaatgactg catctcagga cagttgggaa gaggacaatg ccagaaggac ttagaaaaga 9540 tagactggag ggtggaaaag cagcaggaac agaaacaa aacaggaagc ttgctatcca 9600 gggccactct ggagtcctgt ggcaagatgg aagcgggcta ggggaataca tttgtgctac 9660 tgtgtgtgtg tgtgtgtgtg tgtgtgtgtg tgtgtgtgat caatgcctat caatgttgaa 9720 ggggaaatat gtataccaca ttgattctgg gagcaattct cagtatctgg cctagagaaa 9780 ggaatggccc ctgcagaata gacagagtga atggtgccct ttatcatttg ctaaagtgaa 9840 ggagaaataa acatccttcc atagagtttc aggtaaatga accccacagt tcatctgtgc 9900 cgtggtggag gcctggccaa cagttaaaaa gattagacac ggacaaagtc tgaaggaaac 9960 acctcgaata ggaagaggag agccacctca ttctgtaact ttcctcaagg ggaagatgtt 10020 ccaagagtgg gaataaatgg tcaaaggggg gatttttaat taggaaaacg atttcctgta 10080 tcacttgtga aactggaggt tgatttgggg cataggacaa tagatttgat gctttgcaaa 10140 aagctgtttc aaagcagaga aatggaatag agacaattat gtagcgagga gggagggtgg 10200 ggcgaagatg gagacagaga agtggaagct gactttaggg aagaggaaca tagaccacag 10260 gggcggggcg gggggcaggg gcggggggcg gggctcaaag gaggcagtgg gaacgttgct 10320 agtgtcgca gcgtaagcgt gaatgtgcaa gcgtctttgt ggtgtgtgac caggagtagc 10380 gtggctggct tgtgtgctgc ttgtaatccc agtctttgag gtttccacac tgttccacag 10440 tgggtgtgat tttccctcgg agagcatgag ggctctgctt tccccacatc ctccccagcg 10500 ttcgttggta tttgtttcca agatgttagt gggtgagaca aagcctctct gttgatttgc 10560 ctttaacagg tgacaaaaaa agctcaacca ggagacattt ttgccttctt ggaaggtaat 10620 gctcccatgt agagcaatgg gacccatctc taaggtgagg ctactcttgc agtttgcacc 10680 cagctcttct gatgcaggaa ggaagttggt gggcaagcaa gactgtttgc ttcttgcgat 10740 ggacacattc tgcacacaaa ggctcaggag gggagaaggc tgtttgatgt ttagcactca 10800 ggaaggcccc tgatgcatct gtgattagct gtctccatct gtggagcaga cacggactaa 10860 ctaaaaacca gtgtttttaa attgtcaagc ctttaaggtg aggaaattga cttattgtgc 10920 tgggccatac gtagagcaag tgctctgcat tgggccaacc cccggctctg gtttctaggc 10980 accagaatgg cctagaacta actcacaatc ctcccattcc aggtctcagg tgctagaatg 11040 aaccactata ccagcctgcc tgcctgccta cctgccttcc taaattttaa atcatgggga 11100 gtaggggaga atacacttat cttagttagg gtttctattg ctgtgaagag acaccatgag 11160 catggcaact cttataaagg aaaacattta gttgggtggc agtttcagag gttttagtac 11220 attgtcatca tggctgggaa catgatggca tgcagacaga catggtgctg gagaaaggga 11280 tgagagtcct acatcttgca ggcaacagga cctcagctga gacactggct ggtaccctga 11340 gcataggaaa cctcacagcc caccctcaca gtgacatatt tccttcaaca aagccatacc 11400 tcctaatagt gccactccct atgagatgac agggccaatt acattcaaac tgctataaca 11460 ctttaaagta ttttattttt attattgtaa attatgtatg tagctgggtg gtggcagccg 11520 aggtgcacgc cttaatccc agcacttggg aggcagaggc agatggatct ctgtgagttc 11580 aagaccagcc tggtctataa gagctagttg caaggaagga tatacaaaga acagttctag 11640 gatagccttc aaagccacag agaagtgctg tcttgaaaac caaaaattgt gctgggacct 11700 gtctctgctt tggttgcttc cactccccc agagctggac tcttggtcaa cactgaatca 11760 gctgcaaaat aaactcctgg attcctctct tgtaacagga gcccgaagtc aggcgcccac 11820 ttgtcttctc gcaggattgc catagacttt ttctgtgtgc ccaccattcc agactgaagt 11880 agagatggca gtggcagaga ctgggaaggc tgcaacgaaa acaggaagtt attgcaccct 11940 gggaatagtc tggaaatgaa gcttcaaaac ttgcttcatg ttcagttgta cacagactca 12000 ctcccaggtt gactcacacg tgtaaatatt cctgactatg tctgcactgc ttttatctga 12060 tgcttccttc ccaaaatgcc aagtgtacaa ggtgagggaa tcacccttgg attcagagcc 12120 cagggtcgtc ctccttaacc tggacttgtc tttctccggc agcctctgac acccctcccc 12180 ccattttctc tatcagaagg tctgagcaga gttggggcac gctcatgtcc tgatacactc 12240 cttgtcttcc tgaagatcta acttctgacc cagaaagatg gctaaggtgg tgaagtgttt 12300 gacatgaaga cttggtctta agaactggag caggggaaaa aagtcggatg tggcagcatg 12360 tacccgaaat cccagaactg gggaggtaga gacggatgag tgcccggggc tagctggctg 12420 ctcagccagc ctagctgaat tgccaaattc caactcctat tgaaaaacct ttaccaaaca 12480 aacaaacaaa caaataataa caacaacaac aacaacaaac taccccatac aaggtgggcg 12540 gctcttggct cttgaggaat gactcaccca aacccaaagc ttgccacagc tgttctctgg 12600 cctaaatggg gtgggggtgg ggcagagaca gagacagaga gagacatgac ttcctgggct 12660 gggctgtgtg ctctaggcca ccaggaactt tcctgtcttg ctctctgtct ggcacagcca 12720 gagcaccagc acccagcagg tgcacacacc tccctccgtg cttcttgagc aaacacaggt 12780 gccttggtct gtctattgaa ccggagtaag ttcttgcaga tgtatgcatg gaaacaacat 12840 tgtcctggtt ttatttctac tgttgtgata aaaaccgggg aactccagga agcagctgag 12900 gcagaggcaa atgcaaggaa tgctgcctcc tagcttgctc cccatggctt gccgggcctg 12960 ctttctgcaa gcccttctct ccccattggc atgcctgaca tgaacagcgt ttgaaatgct 13020 ctcaaatgtc actttcaaag aaggcttctc tgatcttgct aactaaatca gaccatgttt 13080 caccgtgcat tatctttctg ctgtctgtct gtctgtctgt ctgtctatct gtctatcatc 13140 tatcaatcat ctatctatct atcttctatt tatctaccta tcattcaatc atctatcttc 13200 taactagtta tcatttattt atttgtttac ttactttttt tatttgagac agtatttctc 13260 tgagtgacag ccttggctgt cctggaaccc attctgtaac caggctgtcc tcaaactcac 13320 agagatccaa ctgcctctgc ctctctggtg ctggggttaa agacgtgcac caccaacgcc 13380 ccgctctatc atctatttat gtacttatta ttcagtcatt atctatcctc taactatcca 13440 tcatctgtct atccatcatc tatctatcta tctatctatc tatctatcta tctatcatcc 13500 atctataatc aattg 13515 <211>14553 <212>DNA <213>Mus musculus <400>6 (SEQ ID NO: 10) cttgaagaac acatgttttc caagagggag cacccatgtt ggaatgacaa tgtagttagt 60 gctcctctcc tgtaggttag tgctcctttg ctataggtaa gtgctcctct cctataggtc 120 agtgctcctc tcctataggt tagtgctcct ctcctatagg ttagtgctcc tctcctacag 180 gttagtgctc ctctgctcta ggttagtcct gctctcctat agtacctaga gagctagggc 240 aaatgggcta ggcccgaagt gcagagacaa acagctatgg aagactgggt aagcacttcc 300 aagctacgaa agagcagtgt gaagggtcag ggcttgtgca gttagtaggg gagatcttcc 360 agttgaagaa agaagaac tgagagccac tgggtatcat cctcctgcgc catgccttcc 420 tggatactgc catgctccca ccttgatgat aatggaatga acctctgaac ctgtaagcca 480 gccccaatga aatattgttt ttatgagagt tgccttggtc atgctgtctg ttcacagcag 540 taaaacccta ataaggcag aagttggtac foottttgct gtgatagacc tgaccatgct 600 ttcctttgaa agaatgtgga tttggtgact ttggatttgc aacacagtgg aatgctttaa 660 atggagatta atgggtcatc aattcctagt aggaatatgg aagactttgt tgctgggagt 720 atttgaactg tgttgacctg gcctaagaga tttcaaagga gaagaatttc agaatgtggc 780 ataaagacag tttttgtggt attttggtga agaatgtggc tactttttgc ccttgtctga 840 900 gtggctgcgc tatctggaaa cttacagcca gcctcttgga cctcgggtga cttacgcaaa 960 tactcaggga cagagatgct tgactctgta ctgatgagtt gtcttggatg caaatatggg 1020 ctcttcattt gactacatgt cacgatgagt caggagctgc tctctccaga gtgtgacaaa 1080 gcgaggggat gctgacggta gctgttctag ctttgaaggt aagcctgcac ttatgctaaa 1140 gtcacacata cacgagccgg gtggagaacc tgtctgtgtg gagacacctt tcattacctg 1200 tggcatccag cctctcaagc ttggactgcc tgtgtgctcc tggactctgg aggtcccact 1260 gctctgtcct ctgctgctta tgatactgac attttaaaag aatccagtgg ttcccccctg 1320 tactcggtgt ctacttctac ctggatgttc ctcatttatg ttctgtgaca cttctctgtg 1380 actctgctgc attcctgggt gacatgtgga caccctgtcc ctttgcagac catgatgtca 1440 ctgtcactag tggaatcaga tgccccaagt gttgtcctgt gtttgggaac gtgacaggca 1500 gtacagaagc agaagaggaa gggtgaaaac ggaaatgtca cagcagcatc tgatgtgtgc 1560 ctcagtcacg catgctgctg attggaacta ctcagcatga gagagggcca tggtgaatac 1620 acaaccctat acacactgtg tccatttctc tctctctctt acacagagag agagggagga 1680 gggggagggg gaggcggagg gggaggggga gggagaggga gtgggagagg gagagggaga 1740 gggagaggga gagggagagg gagagggaga gggagagttt aatgtctgtg aagagatacc 1800 atgaccaaag caactcttat aaagcaac atttaattgg ggctggctta caggttcaga 1860 aattcagtcc attctcacca tggtgggaag catgcaggta gatgtggtgc tggaggaacc 1920 aagagttcta tatcctgatc tgaaggcagc your foot ctgcctcttc tgcacagggc 1980 agagcttgag catagaacat caaagccctt ccccacactt cctccaacaa ggtcatacat 2040 acttcaacaa agacacacct cctaacggtg ccactccctg tggaccaacc atttaaacgc 2100 atgagtctat gagggtcaaa gctcttcaaa ccaccacact catgtacaca cacacacaca 2160 cacacaca ctctcataca cacacaca cacacaca cacacaca cacacaca 2220 cacacaca ccacacacac acacacac agagttctat tttgcactgt ttcactgtca 2280 caaggttcta cttatctcag acacactgcc aggaattgtg tgggaagaact ttcagtttct 2340 ttgggttcac atggacttag cagttcttgg tgatcctgaa agatttctgc agaaagaagc 2400 caaagtgttg agcccaaggc ctggccacac attagtcctg tctagatgaa caggggttta 2460 aaaataaggg ggcatcaagg tgaagccagc aggggctgac ttagagagga gacccaccca 2520 agccaactgc tcgaagtcaa aagcgatgaa tccccatatc cagctgtgcc cggtgctgtc 2580 ttgctacatc tttagtaaat gttcttttag ttgtatgcgt atgaatattt tgcttgcata 2640 tatttgtgta caccataggt gttcctaggg cctatggagg ccagaagagg gcatcagatc 2700 ctttggaact ggaattatag acacttgtta cccatagagt agattgtggg aaatgagcct 2760 ttagtcttcg agagcggcca gtgctcttaa cctttggtcg tttctccagg tctttgagac 2820 tttattttct tggacatcag gacaggatcc agggctttga gcttgtttct tcagccagct 2880 ttcttttcat gtatattaaa ttttatgtta ttttgctttc tttttcccca agacagaatc 2940 acactctata tagctcaggc tgggtttgaa ttcagtttcc ctgtctcagt ctaccgggta 3000 atatgattac agatgtgagt ctgactttgg tatcaaagtc cccagccctt ctggatatgt 3060 gttttaagga tatcagatat atccttgatt tgctttgaat tttcttttta gttacaacat 3120 aattagttcc gtgtcacctg aatatgtgta tgtcacctac atagtcttcc ttcttctctt 3180 cttccctctc ccaccttccc aggtacctgt ctgtcttcat atccttgtgc tgagagtctt 3240 gttgagggag atgatgaccg agacagagcc actggggaag ggagatgggc tagtgcaggt 3300 cttcagagag gagctcgtga atattgtagc ccctttagtc cctggcatgt cctcttgtat 3360 agccaccgcc atgctgtggc ctggcagaag tgaataagtt gtccagctgt tgacaggcct 3420 gccctccaga cccagtctga tcccaagaaa gggcatctgt gtctgtctct gaggccgtaa 3480 gtgctgcctg gttgtctcca gcttgacttg acactccctc cttaataaga gtaccacaga 3540 acagggtctg cagagtccct gggccaggtc cctgtgctgt cctggaatgc caggcgtgaa 3600 tttcctgtga agtaggactt tgctcgccaa gctcccacgg cttgcccttc agatagccag 3660 aattatctgg taccctgcat tgccgttcaa tacgcagagt atcactggaa gcgcgcgcgc 3720 gcacacacac acacacacac acacacacac acacacacac acacgcccac tccatcttta 3780 aaccccaccc cccagcaacg gcggtgtaaa cactctccat caggaagctg aaacgcagtt 3840 gccctctgct ggggagatga aggcagcttg ctgggggcga ggaccgtgct agcaaccttc 3900 cctggtgcac acgggctctg gtgcatgacg ggaacggaaa cgcggaacta aagtcagtcc 3960 tgctttttttttttttttttttttttttttttttttttttttttttgggcgttggtg 4020 gtggactgag tgacaatcag tgaatcact taggtttttttctctctt cgttggtttt 4080 gatgacggt gggagaggggt cagagagaa gggagggat gggagagag ggaggaggga 4140 ggggcgggag gcggggggcg aggaaacgt gctactct ccaatcctac agaaaagg 4200 tttggagaaa gccgcactga gtgacccagc agaggaatc caggaatgtc cgctggaatc 4260 tgactgttga ttccagcgcc atgcagagaa tctaggctgg taggaacatt ctttgtccta 4320 tccgacataa taactccaac siacacggaa aagaaggct atacaagtga agaatggca 4380 ttttcacttt catgactata caatcactc caggtagtaa cacgtgtcta gcacagcggt 4440 tctcaacctg gggtcacga tcccactt ttctgcat cagacatttt tacgttgtta 4500 ttcataacag tagcaaatt gcagctatga agtaacaatg aaatgcattt atggtgcgtg 4560 tgtgtgtgtg tgggggta tcaccttaac atttactgta agaaggttga gatactgct 4620 ccagcagcta gtgtgttgga cttaggttct gggtatatta ccagcaatag ccaaccagaa 4680 tcccaccca ccacagcatt gaggccccat gcagggcttg ctgggagagg cactgataag 4740 acttctttat gtatttatt agagacgaat actcattagg taggccaagc tagcgtcaa 4800 ctcatggcaa ttctcctcct ccagttctct aagtactgga ctcaggagtg tgttgccatc 4860 atacagta aggatttatt gactgaagaa atctcaagt ggctttggtt aatccctact 4920 acgccagagg ctgaggcagg aggcgcgcaa ggtcaggct tgcctgggct acatagag 4980 tgagctcaat ttgacactt ggtgcggtgt tagtagtaat agtaagatg aaggtgtggc 5040 tcaggtgggg ccggtgatgzgacacttg gggtctcctg gtccatctgc agctgtgcaa 5100 shaggagagc gggaatgag gggaagag gggaagag 5160 agaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa has 5220 agggaaagg agggaaaaaaaaaaagggaaaaaagggaaaaaaagggaaaaaaaagggaaagg 5280 Agaaaaaaaaaaaaaaaaaaaaaaaaaaaa out 5340 cgaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa roamed 5400 acagaatgcg aggagggag gagagagaaaaaaagg aggagggaaaaaagg 5460 aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa haveaaaaaaaaaaaaaaaah.gaaaaaaaaaaaaaaaaaaaaaaaggggs 5520 agaaaaaaaaaagaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaagg has 5580 Agggaaaaaaaaaaaaaaaaaagggaaaaaaaaaaaaaggggaaaaaaahaaaaaaahaaaaaahaaaaaaaa Haaaaaaaaahaaaah from 5640 gaatgagag aggaaga aagaagaaa tgcgagagag gaggagag agaaaaaagg 5700 aaagagaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaagized 5760 gaaaaaaaaaaaaaaaaaaaaaaaaaaa say out gagaaggga gaaaaaa 5820 Gaaaaaaaaaaaaaaaaaaaaggageaaaaaaaaaaaaaaaaaaggaagg 5880 agaatgaaa tgagagag aggagacac aaagagccag aggagaga aaaaaggga 5940 aagagaaaga gaaagaggaa ggctcctt ggacacatct tccttctct tccctgggg 6000 accgccaag cctggtggca tactgtacat tctgtacact gttcattca aacaggctct 6060 gtcttaaaga tggtctgagc ggtcagaaaa gggtattgtt aacttgtttg caaaactgcc 6120 tcaggagagt gctgagtgcg tgaaagttgc tgcccgttaa ggagaagtct ctactacttg 6180 tgatctcacc atcgaaaatt tctttaattg tctcctggtg ttctgggttt tgcagttttg 6240 tttctaagga tacattcttg ggtgatgtca caaagtcccc aaagacacgg tggagctgtg 6300 ttagatgggg aaagacagtc tgctgaggat ttatctggaa ctgtcagaag gaaaagaagg 6360 taaatggggc acttgggaaa gtggcctcta gtttgacttc tggcttagca aaggttgtgg 6420 ggagataagg catacacagt agttagcagg aggcaacagg gtcctgggag gacgcgaggc 6480 agaaggagag gctgggctga cagcatgcaa tcattgcata gtctccaaag gagattgcaa 6540 catggctgag ttttcagagg tcctacagag cccgtggtag agattctgtg ggttctgaga 6600 caacttgact ttagccagat ggtatttgag taatctggga gagagaaaac agctacagca 6660 aacagggcca catttagtga cgaaactctc actttgactg ttgagtcatt tgcagtgggc 6720 cctgaggtca ggctggccct cagctcaaaa acaagcgagg aactgaagca attactcaga 6780 taatccacag ccacagccac tggaaagggc cacatcccca gagacagcac agcaggggtg 6840 ggggtggggc tatgagaaag ttagtgattg tagcagttat ctagaatgtg cggagcagag 6900 gaggttacac aaaaacctag aatgtcattc aatgtgggaa accgagaggc tcccaagccc 6960 taaaaggaac agtttgcttt cagccaaaat ggaaataaaa tttggggctt aaatctggca 7020 aatgattcag accttctgtg taggtgtctt taaatgcaca gcagattgat tttcatgttg 7080 gagtttattt gaactaaaag acagaaatgg tgaaaagcac acctgaagaa attgagatgc 7140 tatgaataaa atcatttact tacagctatc acttaattag tacctccttc caccttgctg 7200 atttattggg ctagtcaagg aagaaaagat cttccctcct ccttctctcc tcctccccct 7260 cctctcctcc tcccctcccc tccttgacct tcctctcctc cttttccctc ctccccctct 7320 tcttctcttc accccctcct cccctcccct cctctgtact cctccccttt cctcccaatc 7380 tcttttttct cccccttctt ctctttctcc cccctcctct tccctcctct tcctccctcc 7440 ctccctcctc ctcctcatcc tcctcttcct cttcatcctc ttctccttcc tccctctcct 7500 cctcctcctt ttccagccct acctaccttc cctttcttct tcatttattc aaagtagctt 7560 tgaacagcac tactcggttt agttgtgtat aaaaggaaaa tgcaggtcca agcagcttgg 7620 ggaagattgc tttttgctct ctggaggcag atgatgacag ttcaagatca ttccttttgc 7680 tccatgtcac aggaaggggg acatgccgaa tctaccagtt tgcagccacc tacacaggat 7740 ccaccttcac ttctaaggaa atgtttggga agctacctac caaccacttc tggcatctca 7800 tgggctagag gactcttaaa tggcactctt atttgtttaa taaaggaggt tgtgacgtgt 7860 agttttaaat cccttccaca caacaattgc tactctctga ccaaaaaaga agggagacag 7920 gatacggcta ggtgtctagt agactttacc actttgaaaa gccttaatat aaatcaggta 7980 gatacatctt tttaacttat tcttgtaaag acaaaaacaa aactttattt ttatttgtgt 8040 gtatgcttgt gtgtgtgtgc ctgtgtgtat accacatgtc gctggtgccg gagaacacca 8100 gaagagggga cctgatctcc tggagctaaa gctatccatg gttctgagct gcctgatgtg 8160 ggtgctggga acagaactct ggtcttctgc aagagcaaca agcctcctct taactacgaa 8220 tctcctcccc atccccccaa atacatttaa ttattcattt tagcagcttt atttcgtaac 8280 tacttatcac agcataaaac aaggatttta tatatattac atgcaatcga ggataagagt 8340 tgaggggaga tgcgtgtgct ccttctgggt gtctgtgcttt ttgaagaatg taagcagtgc 8400 acaagggacc gaggcgtgcc tgtctgccag gagctgtctt cttcccttgg actctgagct 8460 gagtgcagtg ctccgaagaa gtaaaagacg acctcatgaa gcaatgtctt caacccaaac 8520 atgctgtcca gacaaagtcc agcttcatta gtgctctgag gagagactta ctgagcctca 8580 ggaaagcccc cctcagcatg gcgaaagtcc actttgattg aagtgactcg aaagccatgg 8640 cagtgcggcg gcggccgcgt ggagcttgtg ctcgagtcgg aagcggcatc tttgtcaggc 8700 ggctgtgatt agcacgggga ggcaggactg gagtgaagga agagttgggg gcggggctta 8760 gcgctctggt ctcctaagct gtagtcagcg cctcaagatt tgtaacctgc cttctgcctt 8820 cccagccagg cagtcaagtg gctccaagct gaagactgca aagtgcccct aaccttttgg 8880 ttatagcgag gctgaagaca ccgtgctctt tcatgaaagc cggatgtctg aaatccgatt 8940 tgataaatat ggataaaacg tataacgctc gatcaatcga atcgaaggag ctcacgattg 9000 gcaccacggc tttggggaca acagagtact gactcgttgg gaggacttgg atacttcccc 9060 tcctcttcca tctcttcccc tttcctcact tcctcctcct tccttctcca ttttctccct 9120 cttcactgtt tcttactatt tttacaaaag attttattta tttatttatt tatttattta 9180 tttatttatt tatttattta tttatttaat gtatgcgagt acactgtagc tgtcttcaga 9240 cacaccagaa gagggcgtca agttccatta gagatggttt cgagccacca tgtggttgct 9300 ggggcctctg gaaggaccgc cagtgctctt aacccctgag ccatttctcc agtacccttc 9360 tcaccgtttc tcttcaatct tcttcctctt ccttctccac tttccttgtc ttcttggttt 9420 cattatcttt ctccctttct tcctcttctc cccttcttcc tcctccactg tagttttcct 9480 tccctactct tttcctgcct ccctcctcct cccctctcat tccccctcct ctttcctcct 9540 tctccctcct cctccttcct tctccctctc ccctctcccc tctcccttct cccttctccc 9600 cctcctcttc ctctttctcc ttctccaccc ctcctgtcac agtatcaatg gcaagggtgt 9660 tctagaatgg aggagtgtcc cctaggcact aacgaagcc agttaggatg ctctgagacg 9720 ggtacaattc agggaggggcc gtgggatgg aagggttgtg ctgcgattca ttctggaggca 9780 accccaggc agaatcatga ggttggttcc ggattcgcag ggcacaatc agaagaggaa 9840 ggtttcagga tgttcagga tgtcagga tgtcagga tgtcagga 9900 gccactgtac aagcgtgctt tattaaccac gtgggattaa atctctttt aaatttatt 9960 tcaactctta aggaacgtg aacttcaca ttcaattta gacttgcagc tcttatgggg 10020 aaaaaaaggg gatcttaaga atattaagca taggcggctg gagagatggc tcagcggtta 10080 agagcactct ctgctctccc agaggtcctg agttcattc ctagcaacca cataagtt 10140 aacaacagtc tttaatgaat tctaatgccc tcttctggtg tgtctgaaga cagttacagt 10200 gtactcatat aaaaaata aaaaattta aaaaaatgaa tattaggcat agattcctgg 10260 atcctaagaa agccatcaga gctggagcca tgtgtgggat cctgcttggt gctggagggg 10320 cagagttcat gcccccgggg ttttactta ttcacatt ttcatcgttg tttgaaca 10380 gggtcttgtg tggtccaggc tggccttgaa ctcatctttc agcctctacc tcacaggttc 10440 tgggattact tggttcctaa aagtatctcc gtcaagctcc ctggtgttat ggctgtgcca 10500 accaggagg tctatacact cgctcaggta gagggagaag atccgaatct ctgacaggga 10560 ctgctgcctc tcggggcaaa tggagtgaag gagagcggca gaaggattta ggaagatgg 10620 acgggagagt ggaatgctg cagaagccag aaaacaaagc aggaagcctg ctgtccagtg 10680 gggctcaaga gcggagggat gcgaggggc tgcgcaggaa catttagcgt ctgcgtctat 10740 ggggggtaggg gcggggtgcc agcacctagt cacctgaagg ggaaatgctt gcccagggag 10800 caggtctcag tagctgacct agaaagga gcggccccta cagaggagac acgggtcact 10860 gtttgttaaa gtgaaggaga aataatatt ctttcaaaga atcttaggtg agcccagttc 10920 atctgcgctg tggaggcctg gggaacagtt aaaaagaccc tgacacacac ccaaggcaaa 10980 caacacac acggctcctt ccgtaagggt ccatgattct ctgaagaatc agccccggaa 11040 tcagccccgg aatcaggtag tccgtaaaca caatgagtgt tttactctgc agaagtccag 11100 cctgctggcg tctcccatta ccaaataga gggatagtca cgtgagctca ccggctcgat 11160 ttaaggcacg tggttttcca gggtagatga gctttggctt ctggaaccat tatggggcac 11220 gaaggatgga gccaggatttt tttttttttt ttttttttc tattagcaat tgatttgctt 11280 gggcttggct ggacttgccc agttcttagg cccagctttc ttaactgccg atctgaagtc 11340 tgtcatggag tcagcctagc cttctcactt cccttcagct cgaataggaa gaggaggtgc 11400 acaccagatg gtctgagagc agggataaat ggtgtgcctt tgtctttcag tatttcgtta 11460 ttttaagtag gaagatgctt ttctgtatta cattgcttgt gaaaccggaa gttgattcgg 11520 ggcacaggac aatggatttg gtgttttgca aggactgttt cagaagagag aggagtggaa 11580 gggtggttag agtgaggagt ggggtggggac gggatggggg aagagaagga agggccagac 11640 aggctaggta gggctgagag gaggcggtgg gaacttcttg agttagcgca gcagtaaact 11700 tggatgtgcg tgtatctttg tgatatatga cccggagccg tgtagctggc tccgatagta 11760 ctgctaatgt cagtgtcggg gggggggtt cccatactgt tccacagggg ctgcacattc 11820 ccatcgagag caggaggct cctctctcca tacatcctcg ccagcattcc ttgttgttc 11880 tgtgatgaca gggggtggga tgaatctct ctgttggttt gagagaccgt gagaagctc 11940 aaccccagga cattttgcag tcttggagg cagtgcctcc atgtggagcc gtggagccca 12000 tctctgagtc caggtcactc tgcagttcg cactcagctc ttcagatgca ggagagacgt 12060 tggtgggaaa gcaagattgt ttgctgttg agatagacac attctccaca caaggctca 12120 cgtggggcaa agctgattg acgtacagcg ttcaggaacg cctgtggtag agctatgatt 12180 agctgtctcc atctatgaag cagacaaga gttataaaa aaatcaatgt ttcaattg 12240 tcaactttt aacccgacag caagcgctct gtccctgggc taatccctag cccctggttc 12300 ttgagatggg gtctttgtg cactagactg gcctagaact cacgatctta gtgttccagc 12360 ctcccagctg ctgggatgag ccgctataac cagctgcct gccttcctaa attttaagtg 12420 atgggaagtg ggggagaata cagtttaag tatgcagatc tgagagcagg aacctggcaa 12480 agccaagggg ccggagttac aggcggctaa catggggtgct gggaactgac ccaggtcctt gagaggagca gtgtgtactc ttgaccaaac aggtccgtct ctccagtccc cgtagtatta aaaataggta ctacgggcat ggtggtgcac acctttaatc ccagcactag ggaggcagag gcaggtggat ttctgagttt gaggccagcc tggtctacaa aatgagttcc aggacagcca cggctataca gagaaccct gtcttgaaaa caaaacaaca acaaaatagg tactacaaag cgatgtaatt gtgctcaaac atgcaaaccg aggggactgt atgcataaga aagagaaaga cggccacact ggttctatct gggtgacagg aaatcagtat ttttattttt cacattcatt tttttgttgt tgttgttgac acagtgattt ttctatcaaa aacattttt cttttatagt tccctgagg agctgttttt aaagccgtgc tttgaaaaac cattgaagga gcagaggcag ggagactcct gtgtggcagt cggtgaagca ggccctctgc aggcaggctg gccctggact 13080 tgggagtctc tttccctccc tcctgtgctc aaatagcaaa tgtcaggctt caatgtagct agaaggttct agaatgatta agttccaag gctgaagagc ttccctttt gccttcact 13200 tccctggaga gtcgttgtg tgttccggag tctgcaggt gccttttgtg atgcggttgg 13260 ttcatctcgg gagattccgc ctggaggacc caagttcaag ccctgcctga gctacagagt 13320 gactttcagg tcttctgcgc aattcagtga gacccagtct acaataaa agtaaaaga 13380 aggctgtgga tggaactcgg tggtagagtt ctggttttac tcctagagg agggag 13440 Gaggagggagggagggagggagggagggagggagggagggagggg 13500 aggagggg ctgachagaagagagagggagggagggagggagggagggaagg 13560 gagggggggggggggagggagggagggagggagggagggagggagggagggagggagg 13620 ggaaagaga gaagggtaag aagaaactgt tccaatggtc tgggccacag agtgatggcc 13680 tttgtggtg atcagctgta atccttgatt tgacacaacc tagaatctgg gaagcgagtt 13740 tctgtgaagg agcattcaca ctggctggcc tgtgggcgtg catgtgggag actgtcataa 13800 ttaggttcat taatacagga agtcccagcc cactacaaat ggctcgttc catacccaag 13860 agatgctaac tgtagacggt tggagaaagc aagcaagctg tggatacccc acgctctttc 13920 acctcggctc ctggggggtg ggtgcactgt gtctcttggt attttaaagt cctgccttga 13980 cgtccctgct gtgacagact gtaactggaa ttgtgagctt tagtccttta gttttctacg 14040 ttggtttttc tcaggatatt ttatcgcagt aacagaaaca agaccaggac acttgatctc 14100 ctctgatcaa cactgaagag ttacaaaaca ggctgaggaa acaaactttc ttctccctct 14160 cccccttctg tccctcccct tccttctcgc tccctccctt gccccctctc tccctgtctc 14220 tgtctctgtc tctgtctctg tctctgtctc tgtctctgcc tctcccctcc cctcccctcc 14280 ctctgtctct gtctctgtct ctgtctctgt ctctgtctct gtctctgtcc ctttctcctc 14340 tatctcctaa atggctggag gccatgctag ctcaatgttg aactttgaac acgtatttag 14400 gaaatctttg ttcttaacag ttctgaagtg ctgaagtggt ggtttagtct ctcggcctga 14460 caagctcact tcctctcact ctgtcttaat gaccaaatct gccatttccc taaaacagca 14520 caggctccag ctccaggttg ctccggagcg gag 14553 Example 12 - CHO Stable Site 2 Sequence - U.S. Patent No. 9,816,110 <211>4001 <212>DNA <213>Cricetulus griseus <400>1 (SEQ ID NO: 11) ccaagatgcc catcaactga ttaatagatg ataaaattat tgtacatttc agtgtaatat 60 tattcagttt ttaagaaaaa tgaaattatg taataagcat gtaaatggat atatcttgaa 120 acaaccattc cccattatat tacctaaaca ttgaaagtcc aaaatcatat gatcttttta 180 gtggatctac taatcttttg ctatatgtat tttattgaac tacccatgga tgtgagataa 240 ttggtaacaa cagcacatgg gagagcatgg gatcattcaa ggaagattag agagaatgca 300 ttttttagga gataatggag gagcaataga aaggattaaa tgaggttact gatgaaagtg 360 atggttagag aaggcaatat gaggagggat aactagcact tagggccttt tgaaaaagac 420 atagagaaaa tactattgta gaaacttcct ataattggtg tatagttata tacaccaaag 480 agctcagatg gagttaccct ataatggaaa tattaactac tttttatcac tgtgataaaa 540 catcctgaac agagcaacat agattgggaa gcatttactt tggcttacag ttctaacggg 600 ataaaaattc atgatgaaag aatgaatatg tcagcaaaca gcagtagcaa tggcctgaga 660 agcaggtgag agctcacatc ttgaagtgta agaatgtagc agagaagaaca aactgcaaat 720 gaccagaaaa tgcttttgga tcagagccca tacccctctg actgacttct ccagaaattc 780 tgaacaaata aaactcccca aacagagcca taactgaagg tccagtgtct gagactacta 840 ggggtatttc ttatcaaac cactacaatg gggtgggggg agcaatcctc caagtaggca 900 ctacacacag acaaataaaaa actctagtaa ctggaatgga ttgacttatt tgaattactt 960 gccagtggag ctacatagag cacaattat gtatttaaat taccctttat gatcttacaa 1020 aacttgacag taagatcata ttgctaaaga aaccacatat ttgaatcagg gaacatggtg 1080 atatctagtt gttcttcaac tggaaacttc atgctttctg cccagcattc atgttgctgg 1140 aaagagcaat gtacactacc agtgtagaaa ttaaatcatc aatcttatca agatgtggat 1200 cctataagtt acaataaaaa ttagcctgat aagatatccc caccagaaga atattcacat 1260 aaatgctatg ggagcaacaa gctattttct aaattagctt taatcctatt ctacaagaga 1320 gatccatat ctagaatagt tatagggatc aagaacccat ggcttgattg gtcataggcc 1380 caatgggaga tcctaatatt attgttctac aaaatgaaaa taactcctaa tgacttgttg 1440 ctgcagtaat aagttagtat gttgctcaac tctcacaga gaagttttgt cttacaataa 1500 atggcaatta aagcagcccc acaagattta tatcataccg atctcctcat ggcctatgca 1560 tctagaagct aggaacaaa gaggacccta aggagacat acatggtccc cctggagaag 1620 gggaaggggg cagacctcc aaagctatt gggaggcatgg gggagggggg agggagttag 1680 aagaagaga agggataa agggaggagagaggacaag aggagagagg aagatctagt 1740 chaaggaggaggaggaggaggaggaggagagagaggaggaggaggc cttgtatgtt 1800 taaatagaaa actggcacta gggaattgtc caagatcca caagtccaa ctaataatct 1860 aagcaatagt cgagaggcta ccttaaaagc ctttctctga taatgagat gatgactacc 1920 ttatatacca tcctagagcc ttcatccagt agctgatgga agcagaagca gatactaca 1980 gctaacact gagctagttg cagacaggga ggagtgatga gcaagtcaa gaccaggctg 2040 gagaaacaca cagaaacagc agacctgaa aaaatgttgc acatggaccc cagactgata 2100 gctgggagtc cagcatagga ctttctaga aaccctgaat gaggatatca gtttggaggt 2160 ctggttaatc tatggggaca ctggtagtgg atcatattt atccctagtt catgactgga 2220 atttgggtac ccattccaca tggaggattt ctctgtcagc ctagacacat gggggaggtt 2280 ctaggtcctg ctccaaataa tgtgttagac tttgagac tcccttgaga agactcaccc 2340 tccctgggga gcagaaaggg gatgggatga gggttggtga gggacaggag aggagggg 2400 ggtgagggaa ctgggattga caagtaaatg atgcttgttt ctaatttaaa tgataaagg 2460 aaaagtaaaaaaaaaaacaggcca aaagattata aagacagag gtggtgggtg 2520 actataaaga aacactatta tctaataa aacatgtcag agcacacat gaactatag 2580 tgtttatgaa agtatgtata attackacat atctcaagc CAagaaaaaa attackcatctt 2640 tcagtgatga aggtgatttt atttcccca gattaaagc caagaccta atgaaagtaa 2700 ttatctcaa aaggttgaaa atacatactt tgcaatacac agatctgcct agaaatctca 2760 tgttcacaat acacatgatg ctcaattgaa ttccattcaa tgttacagtt tagataaaca 2820 gtttgtagat aaactcacaa tgtatcattt ctttttattt tttgaccaaa cagcttctca 2880 tctgttattc agaataattc ctcgatggca ggatatccat cccaattggg ggaaggggag 2940 aatttgaaga aaacctagac cacatacata tttgccattg ggaaacaaag tctaaaatga 3000 tgttgttcac atcttctcta ctagtcctct ccccgtccca aagaaccttg gtatatgtgc 3060 ctcattttac agagagagga aagcaggaac tgagcatccc ttacttgcca tcctcaaccc 3120 aaaatttgca tcattgctca gctctgccct tctcatatga cagttacaag tcaaggcttc 3180 caaagtccct ctgtcatgtt tggtgtcaat agtttataca gatgacttca tgtcttcata 3240 tctaatgtct tatatagatt aatattaaac aatgttattt ctctaaccac atttaaatt 3300 aatttaaaaa tccattaatt gtgtctataa aatgcagaca gagtgctgag acacaatata 3360 agcctgatga tctgaatttg aaactcacac ccaccacatg gagaatcaac ttccaaaaat 3420 tttcctatta cttccacact tacaccattg tacaaacaca ataataatga acaaaatgaa 3480 atgaaataaa aaattaagtc tctgtaggta atgctactgt gcagcaaaag taaaaatggc 3540 agcttaagct tgctttatgg ttacacttta ccatcttcca ttattataa ggacttcaat 3600 catggcagaa ctatgctgtt attgtctcag tgtaacctaa ccaggtgttc cagatgttct 3660 taatgtggac acctaaacta tttgatattt gggttaagat ctttccctct ttcagaagaa 3720 acctcaggac agagggaatc ttgtctttta attttgagtc tgtagacttt ttccatttca 3780 aatatacatg aaacaagtga tgaagaaaat taatcaaaag gtgggaattg caatgatatt 3840 aggttcaata ttaagcttca atattatcat ggaatcgcct gttatacact gagtgtttgg 3900 caataaggga tttttagaag aaggagtttt tattctcaac aggttcctta agtttagctc 3960 aaataaatct aagcaatcca ctctagaatt aaatagtttc c 4001 <211> 14931 <212> DNA <213> Cricetulus griseus <220> <221> misc_feature <222> (2176)..(2239) <223> n is a, c, g, t, or a missing nucleotide <400> 4 (SEQ ID NO: 12) catgtacact tatgcaagta tgatatggcc siacacagta ttttacacca atttttatct 60 aaaatata catgtacatc aaaatatatt attaatata catcattat tctttctttc 120 caagtataa acacatacac tgaatttg gttcttgtgg ataattta tgaaacagga 180 aatgcaatt tatcttagca tgtttactc actttcttg catagataac cagtaatcac 240 attgatggat catgtagtga aatgtatttt taggtatcta aggaatttg gcttcgtttt 300 gtgcttgttg acactgaatt ctattcctaa cacagtgtg taaggattct gtctgatttc 360 ttttaccagt atttgtccat ttgcattttc atttattc atggctgctg ttcagaaag 420 tggaaggtag tgtgtcaagt ctgtttaca tgttccctg atgatcagtg tcttacacc 480 tctctgagta catgttggcc aatgtcgttt ctagacccat ctattcttgc ttgacttatc 540 ctggtacatg cctgccaaga aatttctct catcctttct gtctctcac tgatttactt 600 gatgtgtgga ttcacattg atcatatgga atagagat acaatttct ttattcacag 660 tttggaagac ttcaatctc atagatcatc attatttt gctactgttc cctatgctat 720 ggtgaaattt ccatttgaat aattgcttaa acaattaca agaaagaatc tatttttact 780 tgcaataact tccattcag aacatttact acactgttac tatatccaa aactagtttt 840 atatatcatg tgagaaatga ctaattcata atttggccat gatatttt tcagaaacag 900 aaaagtgac caatacatac acatgctat aaatattaag acttcagcaa atttaatatt 960 tattcatgat atcacataaa attcatttat tatgtttat ttaaatgtgt ttttaaaaca 1020 gtggtatcac taaatattaa gttagatgtg tttagtgct taatgaattt atattttaga 1080 atgttataag ttgtatatag tcaaatatgt aaaattttt atttttagg tctttctcat 1140 taaggtattt taatttggg tcccttttcc agagtgactc tagctcatga tgagttgaca 1200 taaaaactaa acagtacaaa atgtacattg cattcagtat tgcactgat ctttgcactg 1260 aagtttgagt cagttcatac atttagtact tgggagtac attaagcta ctttcattgc 1320 tctggcaaaa tgctcgataa gataagagtc tattgtggaa agccatggca gcaggaagt 1380 aagactgctg atgatgttta atccatagtc aagacgcaga aggagatgaa tgctggtatc 1440 caacatttt tgctgttcat tttctctaga accctagtcc ataaagatgt atgacttgca 1500 ttcaaaatgc gtccccttca gttgttcaac ttttctgtaa atacctttc aggcatgtct 1560 agaagattgt ttcgcaaata cttctcaatc cattcaagtt gatagtgcag attaatcact 1620 gcagaataaa agcctgtaac ttggctcacg tgccaaggaa tatgcacact cctgacacat 1680 caataagtaa atcaaagtgt agcttttgcc tttaacattg ccagacttat gtaatgttct 1740 gcacgttctt cctccatcac ttttattct aatggtgttt ccttgacatt gaatcacgct 1800 gtggaagctg cttagaatta acattgaaat ctactgatat attatgatg cagcaattta 1860 gatttactat tttacttaga attttttata attgagagaa tataatattt tcacagttat 1920 ctatctgctg taatagagg attttaaaaa aaatctctat aacttttttt tacaacacac 1980 agtaaaatta agttaaaatt tataaagtc actatgttga tttcaaagtg tgctacgccc 2040 acggtggtca cgcaggtgta gcagaagatg ccactaaggt gggctaaggc cgatgggttg 2100 gggtctgcgc tccctggaga tgagccccag gcggttccct ggcaatcagc tgcgatcatg 2160 atgcccgatg agccannnnn nnnnnnnnn nnnnnnnnn nnnnnnnnn nnnnnnnnn 2220 nnnnnnnnn nnnnnnnnc tgggtgactt tatggaaaga atttgataga tttcatgatg 2280 tagaagaatt ttattaggct tattttacag gagactaaga ccctgggacc taaagatatc 2340 tgggtcctga gaatcaggaa atgggtagag acgtggttga tggtatgaga cagattttag 2400 agaactctta gatcatgggc aatgaccgca atctgatgct tagatagat catctataaa 2460 caattatgct gttctttttc tttctgttgt atgatctgat gatgtagccc ccttgccaag 2520 ttccctgatc ccccttgcca agttccctga ttgtaacagt atataagcat tgcttgagag 2580 catattcaac tacattgagt gtgtctgtct gtcatttcct cgccgattcc tgatttctcc 2640 ttgagcctt tcccttgttc tccctcggtc ggtggtctcc acgagaggcg gtccgtggca 2700 aaagtgtata aatgttctaa aacatttgaa ctctaaaaca tgcaaaatga aaaattaaaa 2760 taaataaca tgaaaattaa aatatattag ctgctaaaag ttaaacaata ctatatata 2820 ttttgttatt agaattcaaa atcacattag ttggatttaa tttgaacatt gcattctttc 2880 aaataatt tcataaaaaagtttcccc atgatagtag aaataata catatgtatc 2940 tatctattta tttaactaca catatatagc atttgtttca actaaaataa atgaatgagc 3000 aaagcaccta agtaattggt gtctatatta tttatgaagc caatagtttc aataaatta 3060 tcatgcataa ggaggtattg caatgttaa accttttg aaacagatat tcccagttac 3120 agaaattata atttctaatc tttcctataa gtagaatgat gataattaat ataggccatt 3180 tgtaaataat gttcagatta aaatattctc tattcacta gagaagaatg atattaatg 3240 tattatattt tattcccat ttgtttgca tattcctct attackccctca gcagtttaaa 3300 tttgttcac catatgtgtg tgtgtttgta tcttaatat ggcactaaaa ttagaataat 3360 daddyaa tctttagg aaagattatt gatttatttt atgttgatag gaaatatct 3420 tttaattgtc ciegaatact ttttctcta ttttaggact gatcagaccc aggactaata 3480 ttttatatgt actaattcta tgtaccaaaa tatgttatta tctcatgaat tctgtctcaa 3540 tattgaggta aaaaatag tccatcatga actttaaaat taaaataatg attattaat 3600 ttttattcat attttgtttg tatgaatggt tatacatcac atgtgtgcct ggtgactgtg 3660 aatgtcagga gaaggtatga aagccactgg aattggaata agagataata tttgagatgt 3720 tatgtgggtg ctgagaatta gacgcaagcc atcttcaaga atagccagca tactatacca 3780 ctgagtaatc cattcatccc tcaataatta tctttgtaga cagtaaatat atttctaaac 3840 tataaatgac cagaaaaatt aatgtattat taatgaagac attcatctca tgtgacacac 3900 ttcacctgtc taaatcagta acactctctc cactaattaa gattttctaa gtgcatgaca 3960 cttactattt ctaaagctgt ccaatggggg ccagtcccca gtcagcaccc agtgagataa 4020 tccatgaatg cattatatc ttaggaaaaa ttcttatcta tgtagtattt agaacatttt 4080 catgtgaggg gataaacaag gaagcacaga tgctttctga tagaaacttt ctctttaatt 4140 catctagaaa aaaaaaacct ctcaggaaaa tctctcttgc tctcctccca atgctctatt 4200 cagcatcttc tccctactta attctagatc ttttctcta tgcctccttg ctgctgccct 4260 gctggctctg cctctatgcct ccccatgtca cttttctttg ctatctcacc gttaccttct 4320 ctgcctcact ctctgccttc ttctctgctt ctcacatggc caggctctgg acaattatag 4380 ttatatgtta cattctcata acacatgata tgtcacatag tttctctcag gctagggata 4440 tcacaatgac tggccaatga gcaagtggcc ttgcatgtag ctctaagttg gtgatggttc 4500 ccagacagta agtagccatt tggttgaaat ttgaggttgg gtagtacatg aagactgaat 4560 tttcttcaaa ctctggcctt gaaatagtaa aacaacacct atgaaaatga cgacctgtat 4620 ttgtctttag aggcaaccac atattgtctg cagggcctgc tttgaatttg ctctgaagtt 4680 agcttgtttg tgtaaaagga agaatcctat atcagcctga gaaatgtaaa atatcctagc 4740 atttcaagtc atcaaaatta tatggagagt ataaatcatc cttctgacta ttcatagtca 4800 tatttgtgtc caccaagtat aaaacacact accaaagggc tgtggaaaaa atcgccataa 4860 ctgttcttat tagggaggca tagcagtggt acctgaggaa gttacagcaa caaccagtca 4920 tccagtcaat aaccccatgg ctttgccact tggaggtacc caataatgtt tggctttgcc 4980 gagtaggact ccaacaaatt cagagggtca atttttaaat gctggttgtc actgctgaac 5040 agtcccattg ccctctgcat aattccacat tggaaagctt tttacactga ttgccaatca 5100 ttaaacagcc tactcagcat aaacaggtat gatattc tgcattttgt tacattacta 5160 gatgaatttcc tattctctcc tacaatagtg gaactgaaaa aagatacaca atcatactac 5220 ccctctacta atcttatgac ttatatcatt tcaatttca gaccataatg caactattg 5280 accaaaacat gtgagatga aaatagaa tgtagaata tattacatat aaaaagaaaa 5340 ggcggactta tttgtttta ttcttagca tgcatagca tacatgatt gaggtttata 5400 tataaaggg acaataatc ttcaagaac tacccctac tgaattaaaa tattaaagaa 5460 ggtcacacat ttactcaaat atattagact actgggcaaa tagacatgaa aagtagagtt 5520 atattgagg taggccttct gtgaaatgtc taggaatt atgtttcata cagtgtgtaa 5580 ccaagtggga atcatatcag aaagcagtca aaagcttata ttacaagtaa cagatgcttg 5640 gttatatgac ctcccagagc ttgactgtct atacacaaaa agtgtgtta ataaaactgt 5700 aatttgggct atgttttttt aaatggctc accacatga aaggaaggga atgagcatgt 5760 catggatgct tagagattat gcttccagca agaagaattg agctttggct cttattacag 5820 aaacatgaca aggtgtgagt tttatttatt agaaattata tatatttta agctggggac 5880 taaaaattt attgaaacaa acaggcaagg gataggcatg tactagaagc aaaaatagga 5940 tgtcaatgct gtaatgttat ttttggacc aaatagtat ttcctataga aatgacaatg 6000 atcttaggtt attattcttc ataaagatga caagttcaca agatatccta gttcattaaa 6060 atcgttttag tcatttaata gagtgctgtg atagattaca caaaggaaag cacttacgat 6120 gagaaataat gatatccaca attattttct taattcttag aaacattcta ttgttatatc 6180 tcaatctcag aagccactta ttgctttatt attgaaacat atgaaattgt aagttatata 6240 ttgtctatgg tgacatttca aagaacatgt gacgtacagt gtagcacaga taaagaacat 6300 aactgcagct gaatcagtaa ctaaacttac atacattaaa tctgccatgt tggcaacagt 6360 gtgtgcacta ccaaaggatg tactaatgct cacgacactc ccctatgtca cccttttgttc 6420 atcattacat cataggtcta ttttgtttgc ttttgaaatc tagaccaagt cttttgtgtc 6480 tttccaagca cagagctcat taatttacct catagacttg tttaacttct tctggttcat 6540 aattgaata gaatactca ctactaatta tgtgagaccc tgccagtacc atagcacatg 6600 gataatttt acataaaca tgcatacaag windowgatt cagactgaac atgaattttta 6660 gagaaatcag gaggagtat atgggagtgg tggagtgag actagagaaa tgtattaaa 6720 ctataatctc atacaagc tctactaagc aaaaaacatg aaacattgtc attcaagtga 6780 aacatcagtc ttcaattgg aagatattt ttactaggaa atgtctggt agatggttat 6840 tacttagaaaaaaaat tagaaaacgg taacttta taaaaagaat atacaatga 6900 gactacatga aaagttctta actaatgaaa caataatctt gaactttt tcttaaagt 6960 ttaatca taaccatcat ggaattcaa atttaaacta tttacatatt acccctgaaa 7020 taataactaa tacctaa aaatatata aaaaaaat ggcaatgcat gccatcatgg 7080 atttgggaga gagaatgttc attgcagttc tgaatggata ctggtgccac cacggtgaaa 7140 atctctgtat aggtccttcc aaagctgaa atagacata tcacagacc tgccacacat 7200 tttcaagca ataccaccaa ggactctacc tgactgcaga gandactttct cataaaatat 7260 tattgttgat ctattcataa tattctggaaa atagaacag ccagatgcc catcaactga 7320 ttaatagatg aaaattat tgtacatttc agtgtatatat tattcagtttt ttaagaaaaa 7380 tgaaattatg taataagcat gtaaatggat atatcttgaa acaaccattc cccattatat 7440 tacctaaca ttgaagtcc aaaatcatat gatcttttta gtggatctac taatctttg 7500 ctatatgtat tttattgaac tacccatgga tgtgagataa ttggtaacaa cagcacatgg 7560 gagagcatgg gatcattca ggagattag agagaatgca ttttgga gatatggag 7620 gagcataga aaggattaaa tgaggttact gatgaagtg atggttagag aaggcaat 7680 gaggagggat aactagcact taggcctttt tgaaaaagac atagagaaaa tactattgta 7740 gaaacttccct atattggtg tatagttata tacaccaag agctcagatg gagttaccct 7800 ataatggaaa tattaactac ttttcac tgtgataaaa catcctgaac agagcaacat 7860 agattgggaa gcatttactt tggctcag ttctaacgggg aaaaatttc atgatgaaag 7920 aatgaatatg tcagcaaaca gcagtagcaa tggcctgaga agcaggtgag agctcacatc 7980 ttgaagtgta agaatgtagc agagaagaaca aactgcaaat gaccagaaaa tgcttttgga 8040 tcagagccca tacccctctg actgacttct ccagaaattc tgaacaaata aaactcccca 8100 aacagagcca taactgaagg tccagtgtct gagactacta ggggtatttc ttattcaaac 8160 cactacaatg gggtgggggg agcaatcctc caagtaggca ctacacacag acaaataaaaa 8220 actctagtaa ctggaatgga ttgacttatt tgaattactt gccagtggag ctacatagag 8280 cacaattatt gtatttaaat taccctttat gatcttacaa aacttgacag taagatcata 8340 ttgctaaaga aaccacatat ttgaatcagg gaacatggtg atatctagtt gttcttcaac 8400 tggaaacttc atgctttctg cccagcattc atgttgctgg aaagagcaat gtacactacc 8460 agtgtagaaa ttaaatcatc aatcttatca agatgtggat cctataagtt acaataaaaa 8520 ttagcctgat aagatatccc caccagaaga atattcacat aaatgctatg ggagcaacaa 8580 gctattttct aaattagctt taatcctatt ctacaagaga gaatccatat ctagaatagt 8640 tatagggatc aagaacccat ggcttgattg gtcataggcc caatgggaga tcctaatatt 8700 attgttctac aaaatgaaaa taactcctaa tgacttgttg ctgcagtaat aagttagtat 8760 gttgctcaac tctcacaaga gaagttttgt cttacaataa atggcaatta aagcagcccc 8820 acaagattta tatcataccg atctcctcat ggcctatgca tctagagct aggaacaaa 8880 gaggacccta agagagacat acatggtccc cctggagaag gggaaggggg caaccctcc 8940 aaagctaatt gggagcatgg gggagggg agggagttag agggaga agggataaa 9000 agggagggagggagggagggagggagggagggagggagggagggagggagg 9060 caagaaaaga gatacatag taggaggagc cttgtatgtt taaatagaa actggcacta 9120 gggaattgtc caagatcca caagtcca ctaataatct aagcaatgt cgagaggcta 9180 ccttaaaagc ctttctctga taatgagat gatgactacc ttatatacca tcctagagcc 9240 ttcatccagt agctgatgga agcagaagca gatactaca gctaacact gagctagttg 9300 cagahaggga ggagtgatga gcaagtcaa gaccaggctg gagaacaca cagaaacagc 9360 agacctgaaa aaatgttgc acatggaccc cagactgata gctgggagtc cagcatagga 9420 cttttctga aaccctgaat gaggatatca gtttggaggt ctggttaatc tatggggaca 9480 ctggtagtgg atcatattt atccctagtt catgactgga atttgggtac ccattccaca 9540 tggaggaatt ctctgtcagc ctagacacat gggggaggtt ctaggtcctg ctccaaataa 9600 tgtgttagac tttgagac tcccttgaga agactcaccc tccctgggga gcagaaaggg 9660 gatgggatga gggttggtga gggacaggag aggagggg gggtgagggaa ctgggattga 9720 caagtaatg atgcttgttt ctaatttaaa tgaataagg aaagtaaa gagaaaaga 9780 aaacaggcca aagattatta aagacagag gtggtgggtg actataaaga aacactatta 9840 tctaataaaatgtcag agcacacat gaactatag tgtttatgaa agtatgtata 9900 ataactacat aatctcaagc caaaaaaa attackatctt tcagtgatga aggtgattt 9960 atttctcccca gattaagc caaaccta atgaagtaa ttatctca aaggttgaa 10020 atacatactt tgcaatacac agatctgcct agaatctca tgttcacaat acacatgatg 10080 ctcaattgaa ttccattcaa tgttacagtt tagataaaca gtttgtagat aaactcacaa 10140 tgtatcattt ctttttattt tttgaccaaa cagcttctca tctgttattc agaataattc 10200 ctcgatggca ggatatccat cccaattggg ggaaggggag aatttgaaga aaacctagac 10260 cacatacata tttgccattg ggaaacaaag tctaaaatga tgttgttcac atcttctcta 10320 ctagtcctct ccccgtccca aagaaccttg gtatatgtgc ctcattttac agagagagga 10380 aagcaggaac tgagcatccc ttacttgcca tcctcaaccc aaaatttgca tcattgctca 10440 gctctgccct tctcatatga cagttacaag tcaaggcttc caaagtccct ctgtcatgtt 10500 tggtgtcaat agtttataca gatgacttca tgtcttcata tctaatgtct tatatagatt 10560 aatattaaac aatgttattt ctctaaccac atttaaatt aatttaaaaa tccattaatt 10620 gtgtctataa aatgcagaca gagtgctgag acacaatata agcctgatga tctgaatttg 10680 aaactcacac ccaccacatg gagaatcaac ttccaaaaat tttcctatta cttccacact 10740 tacaccattg tacaaacaca ataataatga acaaaatgaa atgaaataaa aaattaagtc 10800 tctgtaggta atgctactgt gcaagcaaaag taaaaatggc agcttaagct tgctttatgg 10860 ttacacttta ccatcttcca ttaattataa ggacttcaat catggcagaa ctatgctgtt 10920 attgtctcag tgtaacctaa ccaggtgttc cagatgttct taatgtggac acctaacta 10980 tttgatattt gggttaagat ctttccctct ttcagaagaa acctcaggac agagggaatc 11040 ttgtctttta attttgagtc tgtagacttt ttccatttca aatatacatg aaacaagtga 11100 tgaagaaaat taatcaaaag gtgggaattg caatgatatt aggttcaata ttaagcttca 11160 atattatcat ggaatcgcct gttatacact gagtgtttgg caataaggga tttttagaag 11220 aaggagtttt tattctcaac aggttcctta agtttagctc aaataaatct aagcaatcca 11280 ctctagaatt aaatagtttc ctaagggcac agctatgaat agagctcaat ttacatataa 11340 aattttgttc accattttatg tcattccagt tttcattagt acaaggaaaa tacaaaatat 11400 ttagatgtca atatcaagtg aatagttcat ctcctttttt aatatatatc acctaaatca 11460 ccattttctc agaaaaatct ggctgaagt tctgtctgga acttcacat gaaaaatatg 11520 cacagcttgc tattataat cctagttgat ttttaagatt catgtctggt gtctgactca 11580 gaggggccag aggctagaca atatttt gatcttcat tgtgaagatt tttaatgatt 11640 attttatat aaatacaa gatgatgat atgtaactt tgtacagttc atagacgctg 11700 aactactttg tgcttaaaat gttagttccc tatcataaat gataggtgat aagtgtagt 11760 ttaatacttt ccctctgagc tatattcatg tactgagaa ttattaaa catgaaaga 11820 ctgtgtttat agtctcagct cctgagaact ggtcaacct taggcaggtg atgccagga 11880 gcaacgtttt tctctacag aggatgcttt gctgccaagc aacctggttg tgtggaatg 11940 ttccttttt aatcaagttt aaagggtctt catcatgctg tgctccaca tattttcagg 12000 ttagagcttg gtccttggag tattacttt taccagaaaa ttcatagtat tctttcaata 12060 actaacaact aaacttttcg aaaaaaga attggaattt caatttaaa gcctgagtaa 12120 aattcttgtg aatcaggata ttttattt agtcttatct tttaaaagt tattttatt 12180 tttaaaaaat tataatatac ttcatatt tccctccttc acttttcttt acaacactt 12240 ctatagatca ccatgtgtttt ttttttac atttatggcc tctttctgtt cattgttatt 12300 acatacaaat agtcttgcct atagagaac accacaattt gttacctgat aacaaattat 12360 cacccttaa aacctacaa ctattgaat tactgaaag actatactta tagtaaa 12420 gatatagtg tgtgcacata tatagataca catatatgta ggatttttaa ttttagattt 12480 tagacatca attattat atgactgaga aactagacac tataaatgag cattcagtat 12540 tcaaccgt gattttagat attgtcacaa tgacagaaaa tttcttata gaaaatttta 12600 agttttgtga ttgctctgtg cacttagtga agtctcacag aaaagaatc atagtatttt 12660 tagtttata taaaaagtac atattaa atggttggc aaaaacac atttgagcat 12720 tttcctatt tactatcaag tagtatcatt ttgaaataat atttgacta gtttcaaaaa 12780 tgaaaacaaa atttaacta atgcctaat ctagcctgat aacatttta tgaatgaat 12840 tattcaatag tgttatcaat taggggccca aaacttttcc taaaataaaa cttttaattt 12900 ttttccattt ttatttaaat tagaaacaaa attgttttac atgtaaatca gagtttcctc 12960 accctcccct tctccctgtc cctcactaac accctacttg tcccatacca tttctgctcc 13020 ccagggaggg tgaggccttc catggggaaa cttcagagtc tgtctatcct ttcggatagg 13080 gcctaggccc tcacccattt gtctaggcta aggctcacaa agtttactcc tatgctagtg 13140 ataagtactg atctactaca agagacacca tagatttcct aggcttcctc actgacaccc 13200 atgttcatgg ggtctggaac aatcatatgc tagtttccta ggtatcagtc tggggaccat 13260 gagctccccc ttgttcaggt caactgtttc tgtgggtttc accaccctgg tcttgactgc 13320 tttgctcatc actctccct ttctgtaact gggttccagt acaattccgt gtttagctgt 13380 gggtgtctac ttctactttc atcagcttct gggatggagc ctctaggata gcatacaatt 13440 agtcatcatc tcattatcag ggaagggcat ttaaaggtc ctctccattg ttgcttggat 13500 tgttagttgg tgtcatcttt gtagatctct ggacatttcc ctagtgccag atatctcttt 13560 aaacctacaa gactacctct attatggtat ctctttcttt gctctcgtct attcttccag 13620 acaaaatctt cctgctccct tatattttcc tctcccctcc tctctcccc ttctcattct 13680 cctagatcca tcttcccttc ccccatgctc ccaagagaga tgttgctcag gagatctttgt 13740 tccttaaccc ttttcttggg gatctgtctc tcttagggtt gtccttgttt ctagcttct 13800 ctggaagtgt ggattgtaag ctggtaatca tttgctccat gtctaaaatc catatatgag 13860 tgatgtttgt cttttttgga ctgggttacc tcactcaaaa tggttctttc catatgtctg 13920 tggatttcaa tagcacaaac aacatacagt atcttggggc aacactaacc aaacaagtga 13980 aagaccagta tagcaagaac tttgagttta aagaaagaaa ttaaagaaga taccagaaaa 14040 tggaaagatc tcccatgctc tttgataggc agaatcaaca tagtaaaaat ggcaatcttg 14100 ccaaaatcca tctacagact caatgcaatc cccattaaat accagcacac ttcttcacag 14160 acctgaaaga ataatactta actttatatg gagaaacaaa agacccagga taggccaaac 14220 aaccctgtac atgaggca cttccagagg catccccatc cctgacttca agctctatta 14280 taggtaata atcctgaaaa cagcttggta atggcacaaa atagacagg tagaccaatg 14340 gattgagtt gaaaaccctg atattaaccc acatacttat gaacacctga ctttgacaaa 14400 gaagctagg ttatacaatg tagaagaa agcatctca acaaatcgtg ctggcataac 14460 tggatgctgg catgtagaag actgcagata gatccatgtc taatgccatg cacaaactt 14520 aagtccaaat ggatcaaaa cctcacata aatccagcca cactgaacct catagaagag 14580 aaagtgggaa gtatccttga aataattggt acaggacc acatcttgaa cttaacacca 14640 gtagcacaga caatcagatc aaatcaat aaatgggacc tcctgaact gagaagcttc 14700 tgtaggcaa tggataagtc aacaggacaaatggcagcc cacggaatgg gaaagatat 14760 tcaccaatcc tatatctgac agagggctgc tctctttg caagaacac ataagctag 14820 tttttaaaac accattaat ccgattataa agttgggtag agaactaaat aagaattgt 14880 taacagagca atctacttg gcagaaagac acatagaaa gtgctcacca t 14931 Example 13 - Guide sequence in AAVS1-like region sequence in CHO (The following guides can be sense guide sequences or antisense guide sequences) [Table 3-1] [Table 3-2] [Table 3-3] [Table 3-4] [Table 3-5] [Table 3-6] [Table 3-7] [Table 3-8] [Table 3-9] [Table 3-10] [Table 3-11]

[0120] It should be understood that the description, specific examples, and data, while indicating exemplary embodiments, are given by way of illustration and are not intended to limit the invention. Various changes and modifications within the invention, including combining the embodiments in whole or in part, will become apparent to those skilled in the art from the discussion, disclosure, and data contained herein and are therefore considered part of the invention.

Claims

1. A mammalian cell comprising a first stable integration site located in a genomic safe harbor and a second stable integration site not located in the genomic safe harbor, wherein the first stable integration site comprises a first reporter gene encoding a first reporter protein, and the second stable integration site comprises a second reporter gene encoding a second reporter protein, and the first reporter protein and the second reporter protein are different.

2. 2. The mammalian cell of claim 1, wherein the first stable integration site and the second stable integration site comprise a recombinase recognition site.

3. The mammalian cell of claim 1 , wherein the mammalian cell is a human cell.

4. The mammalian cell of claim 3 , wherein the human cell is a HEK293 cell.

5. 2. The mammalian cell of claim 1, wherein the second stable integration site is located in a second genomic safe harbor that is different from the first genomic safe harbor.

6. The mammalian cell of claim 1 , wherein the second stable integration site is located in a region that is not a genomic safe harbor.

7. A mammalian cell comprising a first stable integration site located in a genomic safe harbor and a second stable integration site not located in the genomic safe harbor, wherein the first stable integration site comprises a first polynucleotide encoding a first protein and the second stable integration site comprises a second polynucleotide encoding a second protein.

8. 8. The mammalian cell of claim 7, wherein the second stable integration site is located in a second genomic safe harbor that is different from the first genomic safe harbor.

9. 8. The mammalian cell of claim 7, wherein the second stable integration site is located in a region that is not a genomic safe harbor.

10. The mammalian cell of claim 7 , wherein the mammalian cell is a human cell.

11. 1. A mammalian cell comprising a first stable integration site located in a genomic safe harbor and a second stable integration site not located in the genomic safe harbor, wherein the first stable integration site comprises a polynucleotide encoding a first reporter gene that encodes a first reporter protein, and the second stable integration site comprises a polynucleotide encoding a Cas9 gene and a second reporter gene that encodes a second reporter protein, wherein the first reporter protein and the second reporter protein are different.

12. 12. The mammalian cell of claim 11, wherein the first stable integration site and the second stable integration site comprise a recombinase recognition site.

13. The mammalian cell of claim 11, wherein the mammalian cell is a HEK293 cell.

14. The mammalian cell of claim 11 , wherein the mammalian cell is a CHO cell.

15. A mammalian cell as described in claim 11, wherein the Cas9 gene can be removed from the second stable integration site using a recombinase.

16. The mammalian cell of claim 11, wherein the Cas9 gene encodes a Cas9 protein, and the Cas9 protein can be used to generate a mammalian cell comprising at least one stable integration site for stable integration of a polynucleotide.

17. The mammalian cell described in claim 16, wherein the mammalian cell may have multiple stable integration sites for stable integration.

18. The mammalian cell described in claim 11, wherein the mammalian cell comprises a second stable integration site, the second stable integration site comprising a selectable marker gene and an internal ribosome entry site (IRES).

19. The mammalian cell described in claim 11, wherein the mammalian cell further contains a polynucleotide encoding a repressor under the control of a promoter.

20. The mammalian cell described in claim 19, wherein the repressor is a Tet repressor.

21. The mammalian cell described in claim 11, wherein the mammalian cell comprises a polynucleotide encoding a repressor protein under the transcriptional control of a promoter and a polyadenylation signal, and the polynucleotide encoding the repressor protein is inserted randomly or site-specifically into the cell genome.

22. The mammalian cell of claim 11, wherein the mammalian cell is modified by inserting, either randomly or site-specifically, a DNA cassette that is inserted into the cell genome, the DNA cassette comprising flanking lox sites, a promoter, a reporter gene, an IRES, a selectable marker gene and a polyadenylation signal.

23. A mammalian cell as described in claim 11, wherein the stably integrated Cas9 gene is adjacent to lox sites and under the control of a promoter.

24. A mammalian cell as described in claim 11, wherein expression of the Cas9 gene increases the efficiency of homology arm integration into a genomic safe harbor by increasing the occurrence of cuts in genomic DNA caused by the Cas9 endonuclease.

25. The mammalian cell described in claim 16, wherein the stably integrated Cas9 gene provides a higher homologous recombination repair efficiency than homologous recombination repair without a stably integrated Cas9 gene.

26. The mammalian cell of claim 25, wherein the stably integrated Cas9 gene provides a homologous recombination repair efficiency that is 10, 10, 10, 10, 10, 10, 10, 10, 10, 10, or 10 higher than homologous recombination repair without the stably integrated Cas9 gene.

27. ​​A mammalian cell described in claim 11, wherein the CMV promoter is operably linked to a Tet operator and controls transcription of the Cas9 gene.