Compositions and methods for making antibodies based on use of expression-enhancing loci
Site-specific integration of exogenous nucleic acids into expression-enhanced loci in eukaryotic cells using distinct RRSs addresses the instability of multispecific antibody expression, achieving stable and efficient production of antigen-binding proteins.
Patent Information
- Application Number
- JP2025170326
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2016-04-20
- Filing Date
- 2025-10-08
- Publication Date
- 2025-12-25
Smart Images

Figure 2025188169000016 
Figure 2025188169000017 
Figure 2025188169000018
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 325,400, filed April 20, 2016, the entire contents of which are incorporated herein by reference.
[0002] This disclosure relates to site-specific integration and expression of recombinant proteins in eukaryotic cells. In particular, this disclosure relates to compositions and methods for improved expression of antigen-binding proteins (including monospecific and bispecific antibodies) in eukaryotic cells, particularly Chinese hamster (Cricetulus griseus) cell lines, by utilizing expression-enhancing sites. [Background technology]
[0003] Cellular expression systems aim to provide a reliable and efficient source for the production of a given protein, both for research and therapeutic applications. For example, mammalian expression systems are capable of appropriate post-translational modification of recombinant proteins, and therefore, for therapeutic protein production purposes, expression of recombinant proteins in mammalian cells is the preferred method.
[0004] Despite the availability of various expression systems, efficient gene transfer and integrated gene stability for recombinant protein expression remain challenges. One of the concerns for long-term expression of a target transgene is minimizing disruption of cellular genes to avoid altering the phenotype of the cell line.
[0005] Engineering stable cell lines to accommodate the expression of multiple genes, such as multiple antibody chains found in multispecific antibodies, is particularly challenging. Wide variations in expression levels of integrated genes can occur. Integration of additional genes can result in even greater expression variation and instability due to the local genetic environment (i.e., position effects). Expression systems for producing multispecific antigen-binding proteins often require the expression of two or more different immunoglobulin chains intended to pair into unique multimeric forms, and often result in the predominance of homodimers over the desired heterodimer or multimeric combinations. Thus, there is a need in the art for improved mammalian expression systems. Summary of the Invention [Means for solving the problem]
[0006] In one embodiment, the disclosure provides a cell containing a plurality of exogenous nucleic acids site-specifically integrated into two expression-enhanced loci, wherein the plurality of exogenous nucleic acids together encode antigen binding proteins, which may be bispecific or conventional monospecific antigen binding proteins.
[0007] In some embodiments, a cell is provided that contains a first exogenous nucleic acid integrated into a first expression-enhanced locus and a second exogenous nucleic acid integrated into a second expression-enhanced locus, wherein the first and second exogenous nucleic acids together encode an antigen binding protein.
[0008] In some embodiments, the first exogenous nucleic acid contains a nucleotide sequence encoding a first heavy chain fragment (HCF), and the second exogenous nucleic acid contains a nucleotide sequence encoding a first light chain fragment (LCF).
[0009] In some embodiments, the second exogenous nucleic acid further contains a nucleotide sequence encoding a second HCF (or HCF*). The first and second HCFs may be the same or different, such as in a bispecific antigen-binding protein. The nucleotide sequence encoding each HCF or LCF may encode amino acids from the constant region. In some embodiments, the nucleotide sequence encoding the first HCF encodes a first CH3 domain, and the nucleotide sequence encoding the second HCF (HCF*) encodes a second CH3 domain. In some embodiments, the first and second CH3 domains may differ at at least one amino acid position, e.g., a position that results in different Protein A binding properties. In other embodiments, the nucleotide sequences encoding the first and second CH domains differ from each other in that one of the nucleotide sequences has been codon-modified.
[0010] In some embodiments, the first exogenous nucleic acid (containing the first HCF-encoding nucleotide sequence) further comprises a nucleotide sequence encoding a second LCF, which may be the same as or different from the first LCF in the second exogenous nucleic acid.
[0011] In many of the cell embodiments provided herein, each of the nucleotides encoding HCF or LCF is independently operably linked to a promoter, thereby separately controlling transcription of each of the HCF or LCF coding sequences.
[0012] In some embodiments, the first and second RRSs are located 5' and 3' relative to the first exogenous nucleic acid, respectively, and the third and fourth RRSs are located 5' and 3' relative to the second exogenous nucleic acid, respectively, where the first and second RRSs are different, and the third and fourth RRSs are different. Generally, the RRSs in a pair of RRSs flanking an exogenous nucleic acid are different, thereby avoiding unintended recombination and removal of the exogenous nucleic acid. In some embodiments, the first, second, third, and fourth RRSs are all different from one another.
[0013] The first exogenous nucleic acid in the first locus comprises a nucleotide sequence encoding a first HCF, and the second exogenous nucleic acid in the second locus comprises both a sequence encoding a first LCF and a sequence encoding a second HCF, and a first additional RRS may be present between the nucleotide sequence encoding the first LCF and the nucleotide sequence encoding the second HCF. The additional RRS may be different from each of the first, second, third, and fourth RRSs. In some embodiments, the first additional RRS may be present between a promoter to which the selectable marker gene is operably linked and the selectable marker gene, or the additional RRS may be present within the selectable marker gene or within an intron of the selectable marker gene present between the first LCF and HCF coding sequences, or between the first HCF and second HCF coding sequences.
[0014] The first exogenous nucleic acid in the first locus comprises a first HCF-encoding nucleotide sequence and a second HCF-encoding nucleotide sequence, and the second exogenous nucleic acid at the second locus comprises both a first LCF-encoding sequence and a second HCF-encoding sequence, wherein the first RRS and the second RRS may be 5' and 3' relative to the first exogenous nucleic acid, respectively, and the third RRS and the fourth RRS may be 5' and 3' relative to the second exogenous nucleic acid, respectively, wherein the first and second RRSs are different and the third and fourth RRSs are different. In some embodiments, the first and second HCFs are the same, and the first and second LCFs are the same, in which case the RRSs may be engineered such that the first and third RRSs are the same and the second and fourth RRSs are the same. In some embodiments, the first and second HCFs are different, and the first and second LCFs are the same, in which case the RRSs may be engineered such that all of the first, second, third, and fourth RRSs are different from one another. Whether the two HCFs are the same or different, an additional RRS may be present between the first and second HCF coding sequences and / or between the second and first HCF coding sequences. The additional (intermediate) RRS is different from each of the first, second, third, and fourth RRSs. The additional RRS may be contained within a selectable marker gene located between the two HCF / LCF coding sequences or within an intron of the selectable marker gene.
[0015] In another embodiment, a cell is provided that contains an RRS pair integrated into two expression-enhancing loci that can be used for RMCE-mediated integration of nucleic acids encoding antigen-binding proteins.
[0016] In one embodiment, a cell is provided that contains an RRS pair integrated into two expression-enhancing loci that can be used for co-integration of nucleic acids encoding antigen binding proteins via RMCE in the presence of a recombinase.
[0017] In some embodiments, a cell is provided that contains, from 5' to 3', a first RRS, a first exogenous nucleic acid, and a second RRS integrated into a first expression-enhancing locus, and contains, from 5' to 3', a third RRS, a second exogenous nucleic acid, and a fourth RRS integrated into a second expression-enhancing locus, wherein the first and second RRSs are different, and the third and fourth RRSs are different.
[0018] In some embodiments, the first exogenous nucleic acid comprises a first selectable marker and the second exogenous nucleic acid comprises a second selectable marker, wherein the first and second selectable marker genes are different.
[0019] In some embodiments, one or both of the first and second exogenous nucleic acids may contain an additional RRS, i.e., an additional RRS between the first and second RRSs in the first locus and / or an additional RRS between the third and fourth RRSs. The additional middle RRS is different from the RRSs at the 5' and 3' ends. When an additional RRS is contained between the 5' RRS and the 3' RRS (e.g., between the first RRS and the second RRS), one selectable marker gene may be contained between the 5' RRS and the additional (middle) RRS, and another, different selectable marker gene may be contained between the additional RRS and the 3' RRS.
[0020] In another embodiment, the cell provides a first exogenous nucleic acid comprising a third RRS, i.e., an additional RRS between the first and second RRSs in the first locus, wherein the first and second RRSs are flanked by two selectable markers at the 5' and 3' ends of the expression cassette. In other embodiments, the second exogenous nucleic acid may comprise the same third RRS, i.e., an additional RRS between the first and second RRSs in the second locus, wherein the first and second RRSs are flanked by two selectable markers at the 5' and 3' ends of the expression cassette. Four selectable marker genes are contained between the first, third, and second RRSs and are distinct from one another.
[0021] In another embodiment, the cell provides a first exogenous nucleic acid comprising a third RRS, i.e., an additional RRS between the first and second RRSs in the first locus, wherein the first and second RRSs flank two selectable markers at the 5' and 3' ends of the expression cassette. In other embodiments, the second exogenous nucleic acid may comprise a sixth RRS, i.e., an additional RRS between the fourth and fifth RRSs in the second locus, wherein the fourth and fifth RRSs flank two selectable markers at the 5' and 3' ends of the expression cassette. The four selectable marker genes are contained between the RRSs and are distinct from one another.
[0022] In many embodiments, the cells provided herein are cells of a CHO cell line.
[0023] In various embodiments, the two expression-enhancing loci used are selected from the group consisting of a locus containing a nucleotide sequence at least 90% identical to SEQ ID NO:1, a locus containing a nucleotide sequence at least 90% identical to SEQ ID NO:2, and a locus containing a nucleotide sequence at least 90% identical to SEQ ID NO:3.
[0024] In a further aspect, a vector set is provided for integrating and expressing the bispecific antigen binding protein in a cell.
[0025] In some embodiments, the vector set comprises, from 5' to 3', a first RRS, a first nucleic acid containing a nucleotide sequence encoding the first HCF, and a first vector containing a second RRS; from 5' to 3', a second nucleic acid containing a nucleotide sequence encoding the third RRS, the second HCF, and a second vector containing a fourth RRS; and a nucleotide sequence encoding a first LCF either within the first nucleic acid in the first vector or in a third vector different from the first and second vectors, wherein the first, second, third, and fourth RRS are different, and wherein the bispecific antigen binding protein contains a first HCF, a second HCF, and a first LCF, wherein the first HCF and second HCF are different.
[0026] In some embodiments, the nucleotide sequence encoding the first LCF is in a first nucleic acid in a first vector. In some embodiments, the first nucleic acid further comprises a first selectable marker gene.
[0027] In some embodiments, the nucleotide sequence encoding the first LCF is provided in a third vector and is flanked by a 5' RRS and a 3' RRS, where (i) the 3' RRS is the same as the first RRS and the 5' RRS is different from the first RRS and the second RRS, or (ii) the 5' RRS is the same as the second RRS and the 3' RRS is different from the first RRS and the second RRS. In some embodiments, the vectors may be designed so that the common RRS shared by the first and third vectors is provided in a split selectable marker gene format (or split-intron format), e.g., located at the 3' end of the 5' portion of the selectable marker gene of one of the first and third vectors and located at the 5' end of the remaining 3' portion of the selectable marker gene of the other vector.
[0028] In some embodiments, the vector set further comprises a nucleotide sequence encoding a second LCF provided either within a second nucleic acid in the second vector or in a fourth vector separate from the first, second, and third vectors.
[0029] In some embodiments, the first LCF and the second LCF are the same.
[0030] In some embodiments, the nucleotide sequence encoding the first LCF is contained within a first nucleic acid in a first vector, and the nucleotide sequence encoding the second VL is provided on a fourth vector. In some embodiments, the nucleotide sequence encoding the second LCF on the fourth vector is flanked by a 5' RRS and a 3' RRS, where (i) the 3' RRS is identical to the third RRS and the 5' RRS is different from the third and fourth RRSs, or (ii) the 5' RRS is identical to the fourth RRS and the 3' RRS is different from the third and fourth RRSs. In certain embodiments, the vectors are designed such that the common RRS shared by the second and fourth vectors is located in a split marker format (e.g., via an intron), e.g., at the 3' end of the 5' portion of the selectable marker gene on one of the second and fourth vectors and at the 5' end of the remaining 3' portion of the selectable marker gene on the other vector.
[0031] In some embodiments, the nucleotide sequence encoding the first LCF is in a first nucleic acid in a first vector, and the nucleotide sequence encoding the second VL is in a second nucleic acid on a second vector.
[0032] In some embodiments, the nucleotide sequence encoding the first LCF is on a third vector and the nucleotide sequence encoding the second VL is on a fourth vector. In some embodiments, the nucleotide sequence encoding the first LCF on the third vector is flanked by a 5' RRS and a 3' RRS, wherein (i) the 3' RRS on the third vector is identical to the first RRS and the 5' RRS on the third vector is different from the first and second RRS, or (ii) the 5' RRS on the third vector is identical to the second RRS and the 3' RRS on the third vector is different from the first and second RRS; and the nucleotide sequence encoding the second LCF on the fourth vector is flanked by a 5' RRS and a 3' RRS, wherein (iii) the 3' RRS on the fourth vector is identical to the third RRS and the 5' RRS on the fourth vector is different from the third and fourth RRS, or (iv) the 5' RRS on the fourth vector is identical to the fourth RRS and the 3' RRS on the fourth vector is different from the third and fourth RRS.
[0033] In many embodiments of the vector set provided herein, the nucleotide sequence encoding the first HCF may encode the first CH3 domain, and the nucleotide sequence encoding the second HCF may encode the second CH3 domain.In some embodiments, the first and second CH3 domains differ in at least one amino acid.In some embodiments, the nucleotide sequences encoding the first and second CH3 domains differ in that one of the nucleotide sequences is codon-modified.
[0034] In many embodiments of the vector sets provided herein, each of the nucleotides encoding HCF or LCF is independently linked to a promoter.
[0035] In some embodiments, the vector set may further include nucleotide sequences encoding one or more recombinases that recognize one or more of the RRSs, which may be included in one of the LCF-encoding vectors or the HCF-encoding vectors, or may be provided on a separate vector.
[0036] In yet another embodiment, a vector set is provided comprising: a first vector containing a first nucleic acid flanked by 5' and 3' homology arms for integration into a first expression-enhanced locus in a cell; and a second vector containing a second nucleic acid flanked by 5' and 3' homology arms for integration into a second expression-enhanced locus in a cell; wherein the first and second nucleic acids together encode an antigen binding protein.
[0037] In a further aspect, the present disclosure provides a system that can be used to generate cells comprising a combination of one or more vectors and cells (e.g., CHO cells) and having exogenous nucleic acids integrated into two expression-enhancing loci, where the exogenous nucleic acids together encode an antigen-binding protein, either a monospecific or bispecific protein. The system may be provided, for example, in the form of a kit.
[0038] In one embodiment, a system is provided comprising a cell and a vector set, wherein the cell contains an RRS set integrated into two separate expression-enhancing loci in its genome, the RRSs being different from each other and positioned between one or more exogenous nucleic acids, such as selectable markers, for recombination exchange with a gene of interest in the vector set; and wherein the RRSs in the vector set have the same positioning as the RRSs in the cell.
[0039] In some embodiments, a system is provided that includes a cell and a vector set, wherein the cell contains, from 5' to 3', a first RRS, a first exogenous nucleic acid, and a second RRS integrated into a first expression-enhancing locus, and contains, from 5' to 3', a third RRS, a second exogenous nucleic acid, and a fourth RRS integrated into a second expression-enhancing locus; wherein the first and second RRSs are different and the third and fourth RRSs are different; and wherein the first and second expression-enhancing loci are different; wherein the vector set includes: (i) a first vector containing, from 5' to 3', a 5' RRS of the first vector, a first nucleic acid, and a 3' RRS of the first vector, wherein the 5' RRS of the first vector and 3' RRS are different; (ii) a second vector containing, from 5' to 3', the 5' RRS of the second vector, a second nucleic acid, and the 3' RRS of the second vector, wherein the 5' RRS and 3' RRS of the second vector are different; and (iii) a nucleotide sequence encoding a first HCF and a nucleotide sequence encoding a first LCF, wherein one of the two heavy chain-encoding nucleotide sequences is the first nucleic acid and the other nucleotide sequence is the second nucleic acid; in this case, the first HCF and the first LCF are regions of an antigen-binding protein; and when the vector is introduced into a cell, the first and second nucleic acids in the vector are integrated into the first expression-enhancing locus and the second expression-enhancing locus, respectively, via RRS-mediated recombination.
[0040] In some embodiments, the antigen binding protein is a monospecific antigen binding protein.
[0041] In some embodiments, the first and third RRSs are identical, and the second and fourth RRSs are identical. In some embodiments, a first additional RRS is present between the first RRS and the second RRS in the first locus. In some embodiments, the 5' RRS of the first vector is identical to the first RRS and the third RRS; the 3' RRS of the first vector, the 5' RRS of the second vector, and the first additional RRS are identical; and the 3' RRS of the second vector is identical to the second RRS and the fourth RRS. In some embodiments, the LCF-encoding nucleotide sequence is in the first vector, and the HCF-encoding nucleotide sequence is in the second vector. In some embodiments, the 3' RRS of the first vector is located at the 3' end of the 5' portion of the selectable marker gene, and the 5' RRS of the second vector is located at the 5' end of the remaining 3' portion of the selectable marker gene. In other embodiments, the 5' RRS of the first vector is identical to the first RRS and the 3' RRS of the first vector is identical to the second RRS; and in this case the 5' RRS of the second vector is identical to the third RRS and the 3' RRS of the second vector is identical to the fourth RRS.
[0042] In various embodiments, the antigen binding protein is a bispecific antigen binding protein.
[0043] In some embodiments, the vector set of the system further comprises a nucleotide sequence encoding a second HCF that is different from the first HCF.
[0044] In some embodiments, the nucleotide sequence encoding the first LCF and the nucleotide sequence encoding the second HCF are both contained within a first nucleic acid in a first vector, and the nucleotide sequence encoding the first HCF is in a second vector. In some embodiments, the 5' RRS of the first vector is identical to the first RRS, the 3' RRS of the first vector is identical to the second RRS, the 5' RRS of the second vector is identical to the third RRS, and the 3' RRS of the second vector is identical to the fourth RRS.
[0045] In some embodiments, the nucleotide sequence encoding the second HCF is on a separate, third vector and is flanked by the 5' RRS of the third vector and the 3' RRS of the third vector. In some embodiments, the nucleotide sequence encoding the first LC is in a first vector and the nucleotide sequence encoding the first HCF is in a second vector, wherein the 5' RRS of the first vector is identical to the first RRS, the 3' RRS of the first vector is identical to the 5' RRS of the second vector and a first additional RRS, the 3' RRS of the second vector is identical to the second RRS, the 5' RRS of the third vector is identical to the third RRS, and the 3' RRS of the third vector is identical to the fourth RRS, wherein the first additional RRS is contained in a first locus between the first RRS and the second RRS. In some embodiments, the vectors are designed to provide a common RRS in a split marker format, e.g., the 3' RRS of a first vector is located at the 3' end of the 5' portion of the selectable marker gene contained in the first vector, and the 5' RRS of a second vector is located at the 5' end of the remaining selectable marker gene contained in the second vector.
[0046] In some embodiments, the vector set of the system further comprises a nucleotide sequence encoding a second LCF, which may be the same as or different from the first LCF.
[0047] In some embodiments, the nucleotide sequence encoding the second LCF is in a second nucleic acid of a second vector, wherein the 5' RRS of the first vector is identical to the first RRS, the 3' RRS of the first vector is identical to the second RRS, the 5' RRS of the second vector is identical to the third RRS, and the 3' RRS of the second vector is identical to the fourth RRS.
[0048] In some embodiments, the nucleotide sequence encoding the second LCF is in a separate, third vector and is flanked by the 5' RRS of the third vector and the 3' RRS of the third vector. In some embodiments, the 5' RRS and 3' RRS of the first vector are identical to the first RRS and the second RRS, respectively, in the first locus; the 5' RRS of the third vector is identical to the third RRS, the 3' RRS of the third vector is identical to the 5' RRS of the second vector and to an additional RRS present between the third RRS and the fourth RRS in the second locus, and the 3' RRS of the second vector is identical to the fourth RRS. In some embodiments, the common RRS is designed in a split marker format, for example, the 3' RRS of the third vector is located at the 3' end of the 5' portion of the selectable marker gene contained in the third vector, and the 5' RRS of the second vector is located at the 5' end of the remaining 3' portion of the selectable marker gene contained in the second vector.
[0049] In many embodiments of the system provided herein, the nucleotide sequence encoding HCF or LCF may encode amino acids from the constant region. In some embodiments, the nucleotide sequence encoding the first HCF may encode the first CH3 domain, and the nucleotide sequence encoding the second HCF may encode the second CH3 domain. In some embodiments, the first and second CH3 domains differ in at least one amino acid. In some embodiments, the nucleotide sequences encoding the first and second CH3 domains differ in that one of the nucleotide sequences is codon-modified.
[0050] In many embodiments of the systems provided herein, each of the nucleotides encoding HCF or LCF is independently linked to a promoter.
[0051] In some embodiments, the vector set of the system may further include nucleotide sequences encoding one or more recombinases that recognize one or more of the RRSs, which may be included in one of the LCF-encoding vectors or the HCF-encoding vectors, or may be provided on a separate vector.
[0052] In various embodiments, the cells in the systems provided herein are CHO cells.
[0053] In various embodiments, the two expression-enhanced loci are selected from the group consisting of a locus comprising the nucleotide sequence of SEQ ID NO: 1, a locus comprising the nucleotide sequence of SEQ ID NO: 2, and a locus comprising the nucleotide sequence of SEQ ID NO: 3.
[0054] In another aspect, the present disclosure also provides a method for producing a bispecific antigen-binding protein. In one embodiment, the method utilizes the system disclosed herein, and a vector of the system is introduced into cells of the system by transfection. Transfected cells in which exogenous nucleic acids have been properly integrated into two expression-enhancing loci of the cells via RMCE may be screened and identified. HCF-containing polypeptides and LCF-containing polypeptides can be expressed from the integrated nucleic acids, and the antigen-binding proteins of interest can be obtained from the identified transfected cells and purified using known methods.
[0055] In another embodiment, the method simply utilizes a cell as described herein above, which contains exogenous nucleic acids integrated at two expression-enhanced loci, the exogenous nucleic acids encoding antigen binding proteins, and which express the antigen binding proteins from the cell. In certain embodiments, for example, the following items are provided: (Item 1) A cell comprising a first exogenous nucleic acid integrated into a first expression-enhanced locus and a second exogenous nucleic acid integrated into a second expression-enhanced locus, wherein the first and second exogenous nucleic acids together encode an antigen binding protein. (Item 2) The cell of item 1, wherein the first exogenous nucleic acid comprises a nucleotide sequence encoding a first HCF and the second exogenous nucleic acid comprises a nucleotide sequence encoding an LCF. (Item 3) 3. The cell of item 2, wherein the second exogenous nucleic acid further comprises a nucleotide sequence encoding a second HCF. (Item 4) 4. The cell of item 3, wherein the first and second HCFs are different. (Item 5) The cell of item 3, wherein the nucleotide sequence encoding the first HCF encodes a first CH3 domain and the nucleotide sequence encoding the second HCF encodes a second CH3 domain. (Item 6) 6. The cell of item 5, wherein the first and second CH3 domains differ at at least one amino acid position. (Item 7) 6. The cell of item 5, wherein the nucleotide sequences encoding the first and second CH domains differ from each other in that one of the nucleotide sequences is codon modified. (Item 8) 4. The cell of item 3, wherein the first exogenous nucleic acid further comprises a nucleotide sequence encoding a second LCF. (Item 9) 9. The cell of item 8, wherein the first and second LCFs are the same. (Item 10) 9. The cell of any one of items 2, 3, and 8, wherein each of the nucleotide sequences encoding HCF or LCF is operably linked to a promoter. (Item 11) 3. The cell of item 2, wherein a first RRS and a second RRS are located 5' and 3', respectively, to the first exogenous nucleic acid, a third RRS and a fourth RRS are located 5' and 3', respectively, to the second exogenous nucleic acid, the first and second RRSs are different, and the third and fourth RRSs are different. (Item 12) 12. The cell of item 11, wherein the first, second, third, and fourth RRSs are different from each other. (Item 13) 4. The cell of claim 3, wherein a first additional RRS is present between the nucleotide sequence encoding the first LCF and the nucleotide sequence encoding the second HCF, and the additional RRS is different from each of the first, second, third, and fourth RRSs. (Item 14) 14. The cell of item 13, wherein the first additional RRS is contained within an intron of a selectable marker gene located between the nucleotide sequence encoding the first LCF and the adjacent nucleotide sequence encoding the second HCF. (Item 15) 9. The cell of item 8, wherein a first RRS and a second RRS are located 5' and 3' to the first exogenous nucleic acid, respectively, and a third RRS and a fourth RRS are located 5' and 3' to the second exogenous nucleic acid, respectively, and the first and second RRSs are different, and the third and fourth RRSs are different. (Item 16) 16. The cell of item 15, wherein the first and second HCFs are identical, the first and second LCFs are identical, the first and third RRSs are identical, and the second and fourth RRSs are identical. (Item 17) 16. The cell of item 15, wherein the first and second HCFs are different, the first and second LCFs are identical, and the first, second, third, and fourth RRSs are different from each other. (Item 18) 18. The cell of item 16 or 17, wherein a first additional RRS is present between the nucleotide sequence encoding the first LCF and the nucleotide sequence encoding the second HCF in the second exogenous nucleic acid, and the first additional RRS is different from the first, second, third, and fourth RRSs. (Item 19) 19. The cell of item 18, wherein the first additional RRS is contained within an intron of a first selectable marker gene located between the nucleotide sequence encoding the first LCF and the nucleotide sequence encoding the second HCF in the second exogenous nucleic acid. (Item 20) 19. The cell of item 18, wherein a second additional RRS is present between the nucleotide sequence encoding the second LCF and the nucleotide sequence encoding the first HCF, wherein the first and second RRS are the same or different and are each different from the first, second, third, and fourth RRSs. (Item 21) 21. The cell of Item 20, wherein the first additional RRS is contained within an intron of a first selectable marker gene located between the nucleotide sequence encoding the first LCF and the nucleotide sequence encoding the second HCF in the second exogenous nucleic acid, and the second additional RRS is contained within an intron of a second selectable marker gene located between the nucleotide sequence encoding the second LCF and the nucleotide sequence encoding the first HCF, and the first and second selectable marker genes are different. (Item 22) 3. The cell of item 2, wherein the antigen-binding protein is monospecific. (Item 23) 5. The cell of item 4, wherein the antigen-binding protein is bispecific. (Item 24) a first RRS, a first exogenous nucleic acid, and a second RRS, in a 5' to 3' direction, integrated within a first expression-enhancing locus; A cell comprising, in a 5' to 3' direction, a third RRS, a second exogenous nucleic acid, and a fourth RRS integrated within a second expression-enhancing locus, The cell, wherein the first and second RRS are different and the third and fourth RRS are different. (Item 25) 25. The cell of item 24, wherein the first exogenous nucleic acid comprises a first selectable marker gene and the second exogenous nucleic acid comprises a second selectable marker gene, and the first and second selectable marker genes are different. (Item 26) 26. The cell of item 25, wherein the first exogenous nucleic acid further comprises a first additional RRS, and the first additional RRS is different from the first and second RRSs. (Item 27) 27. The cell of item 26, wherein the second exogenous nucleic acid further comprises a second additional RRS, wherein the second additional RRS is different from the third and fourth RRSs. (Item 28) 25. The cell of item 24, wherein the first exogenous nucleic acid comprises a first selectable marker gene, a first additional RRS, and a first additional selectable marker gene, wherein the first selectable marker gene and the first additional selectable marker gene are different, and the first additional RRS is different from the first RRS and the second RRS. (Item 29) 28. The cell of Item 27, wherein the second exogenous nucleic acid comprises a second selectable marker gene, a second additional RRS, and a second additional selectable marker gene, wherein the second selectable marker gene and the second additional selectable marker gene are different from each other and from the first selectable marker gene and the first additional selectable marker gene, and the second additional RRS is different from the third and fourth RRSs. (Item 30) 30. The cell according to any one of items 1 to 29, wherein the cell is a CHO cell. (Item 31) 31. The cell of item 30, wherein one of the two expression-enhanced loci is selected from the group consisting of a nucleotide sequence that is at least 90% identical to SEQ ID NO:1, a nucleotide sequence that is at least 90% identical to SEQ ID NO:2, and a nucleotide sequence that is at least 90% identical to SEQ ID NO:3. (Item 32) A vector set for expressing a bispecific antigen-binding protein in a cell, comprising: a first vector comprising, in the 5' to 3' direction, a first RRS, a first nucleic acid comprising a nucleotide sequence encoding a first HCF, and a second RRS; a second nucleic acid comprising, in the 5' to 3' direction, a nucleotide sequence encoding a third RRS, a second HC, and a second vector comprising a fourth RRS; and a nucleotide sequence encoding a first LC, either within the first nucleic acid in the first vector or in a third vector distinct from the first and second vectors; wherein the first, second, third, and fourth RRSs are different; and a vector set wherein said bispecific antigen-binding protein comprises said first HCF, said second HCF, and said first LCF, and wherein said first and second HCFs are different. (Item 33) 33. The vector set of item 32, wherein the nucleotide sequence encoding the first LCF is located within the first nucleic acid in the first vector. (Item 34) 35. The vector set according to Item 34, wherein the first nucleic acid further comprises a first selectable marker gene. (Item 35) 33. The vector set of item 32, wherein the nucleotide sequence encoding the first LCF is in the third vector and is flanked by a 5' RRS and a 3' RRS, wherein (i) the 3' RRS is the same as the first RRS and the 5' RRS is different from the first RRS and the second RRS, or (ii) the 5' RRS is the same as the second RRS and the 3' RRS is different from the first RRS and the second RRS. (Item 36) 36. The vector set according to Item 35, wherein the common RRS between the first and third vectors is located at the 3' end of the 5' portion of the selectable marker gene of one of the first and third vectors and at the 5' end of the remaining 3' portion of the selectable marker gene of the other vector. (Item 37) 33. The vector set of claim 32, further comprising a nucleotide sequence encoding a second LCF, either within the second nucleic acid in a second vector or in a fourth vector separate from the first, second, and third vectors. (Item 38) 38. The vector set according to Item 37, wherein the first and second LCFs are identical. (Item 39) 39. The vector set of item 38, wherein the nucleotide sequence encoding the first LCF is within the first nucleic acid in the first vector, and the nucleotide sequence encoding the second VL is on the fourth vector. (Item 40) 40. The vector set of item 39, wherein the nucleotide sequence encoding the second LCF on the fourth vector is flanked by a 5' RRS and a 3' RRS, wherein (i) the 3' RRS is identical to the third RRS and the 5' RRS is different from the third RRS and the fourth RRS, or (ii) the 5' RRS is identical to the fourth RRS and the 3' RRS is different from the third RRS and the fourth RRS. (Item 41) 41. The vector set according to Item 40, wherein the common RRS between the second and fourth vectors is located at the 3' end of the 5' portion of the selectable marker gene of one of the second and fourth vectors and at the 5' end of the remaining 3' portion of the selectable marker gene of the other vector. (Item 42) 39. The vector set of Item 38, wherein the nucleotide sequence encoding the first LCF is within the first nucleic acid in the first vector, and the nucleotide sequence encoding the second VL is within the second nucleic acid on the second vector. (Item 43) 39. The vector set of item 38, wherein the nucleotide sequence encoding the first LCF is on the third vector and the nucleotide sequence encoding the second VL is on the fourth vector. (Item 44) The nucleotide sequence encoding the first LCF on the third vector is flanked by a 5' RRS and a 3' RRS, wherein (i) the 3' RRS on the third vector is identical to the first RRS and the 5' RRS on the third vector is different from the first and second RRSs, or (ii) the 5' RRS on the third vector is identical to the second RRS and the 3' RRS on the third vector is different from the first and second RRSs, and wherein the fourth 44. The vector set of item 43, wherein the nucleotide sequence encoding the second LCF on the vector is flanked by a 5' RRS and a 3' RRS, wherein (i) the 3' RRS on the fourth vector is identical to the third RRS and the 5' RRS on the fourth vector is different from the third and fourth RRSs, or (ii) the 5' RRS on the fourth vector is identical to the fourth RRS and the 3' RRS on the fourth vector is different from the third and fourth RRSs. (Item 45) 33. The vector set of Item 32, wherein the nucleotide sequence encoding the first HCF encodes a first CH3 domain, and the nucleotide sequence encoding the second HCF encodes the second CH3 domain. (Item 46) 46. The vector set of item 45, wherein the first and second CH3 domains differ by at least one amino acid. (Item 47) 46. The vector set of item 45, wherein the nucleotide sequences encoding the first and second CH3 domains differ in that one of the nucleotide sequences is codon-modified. (Item 48) 33. The vector set according to Item 32, wherein each of the nucleotide sequences encoding the variable regions is independently linked to a promoter. (Item 49) 33. The vector set of item 32, further comprising a nucleotide sequence encoding a recombinase that recognizes the first and second RRSs and / or the third and fourth RRSs. (Item 50) a first vector comprising a first nucleic acid flanked by 5' and 3' homology arms for integration into a first expression-enhancing locus in a cell; and a second vector comprising a second nucleic acid flanked by 5' homology arms and 3' homology arms for integration into a second expression-enhancing locus of the cell, A vector set, wherein the first and second nucleic acids together encode an antigen-binding protein. (Item 51) A system comprising a cell and a vector set, The cells a first RRS, a first exogenous nucleic acid, and a second RRS, in a 5' to 3' direction, integrated within a first expression-enhancing locus; a third RRS, a second exogenous nucleic acid, and a fourth RRS, integrated within the second expression-enhancing locus, in a 5' to 3' direction; wherein the first and second RRSs are different, the third and fourth RRSs are different, and wherein the first and second expression-enhancing loci are different; The vector set comprises: a first vector comprising, in a 5' to 3' direction, a 5' RRS of the first vector, a first nucleic acid, and a 3' RRS of the first vector, wherein the 5' RRS and the 3' RRS of the first vector are different; a second vector comprising, in 5' to 3' direction, a 5' RRS of the second vector, a second nucleic acid, and a 3' RRS of the second vector, wherein the 5' RRS and 3' RRS of the second vector are different; a nucleotide sequence encoding a first HCF and a nucleotide sequence encoding a first LCF, wherein one of the nucleotide sequences is in the first nucleic acid and the other nucleotide sequence is in the second nucleic acid, wherein the first HCF and the first LCF are regions of antigen binding proteins; In this case, when the vector is introduced into the cell, the first and second nucleic acids in the vector are integrated into the first expression-enhancing locus and the second expression-enhancing locus, respectively, via recombination mediated by the RRS. (Item 52) 52. The system of claim 51, wherein the antigen-binding protein is a monospecific antigen-binding protein. (Item 53) Item 53. The system of item 52, wherein the first and third RRSs are identical and the second and fourth RRSs are identical. (Item 54) 54. The system of Item 53, wherein a first additional RRS is present between the first RRS and a second RRS in the first locus. (Item 55) 55. The system of item 54, wherein the 5' RRS of the first vector is identical to the first RRS and third RRS, the 3' RRS of the first vector, the 5' RRS of the second vector, and the first additional RRS are identical, and the 3' RRS of the second vector is identical to the second and fourth RRS. (Item 56) 56. The system of item 55, wherein the nucleotide sequence encoding the VL is in the first vector and the nucleotide sequence encoding the HCF is in the second vector. (Item 57) 56. The system of Item 55, wherein the 3' RRS of the first vector is positioned at the 3' end of the 5' portion of the selectable marker gene, and the 5' RRS of the second vector is positioned at the 5' end of the remaining 3' portion of the selectable marker gene. (Item 58) 53. The system of claim 52, wherein the 5' RRS of the first vector is identical to the first RRS, the 3' RRS of the first vector is identical to the second RRS, and wherein the 5' RRS of the second vector is identical to the third RRS and the 3' RRS of the second vector is identical to the fourth RRS. (Item 59) 53. The system of claim 52, wherein the antigen-binding protein is a bispecific antigen-binding protein. (Item 60) 60. The system of item 59, further comprising a nucleotide sequence encoding a second HCF that is different from the first HCF. (Item 61) Item 61. The system of item 60, wherein the nucleotide sequence encoding the first LCF and the nucleotide sequence encoding the second HCF are both contained within the first nucleic acid in the first vector, and the nucleotide sequence encoding the first HCF is in the second vector. (Item 62) 62. The system of claim 61, wherein the 5' RRS of the first vector is identical to the first RRS, the 3' RRS of the first vector is identical to the second RRS, and wherein the 5' RRS of the second vector is identical to the third RRS and the 3' RRS of the second vector is identical to the fourth RRS. (Item 63) 61. The system of claim 60, wherein the nucleotide sequence encoding the second HCF is on a separate third vector and is flanked by a 5' RRS of the third vector and a 3' RRS of the third vector. (Item 64) 64. The system of claim 63, wherein the nucleotide sequence encoding the first LCF is in the first vector, the nucleotide sequence encoding the first HCF is in the second vector, the 5' RRS of the first vector is identical to the first RRS, the 3' RRS of the first vector is identical to the 5' RRS of the second vector and a first additional RRS, the 3' RRS of the second vector is identical to the second RRS, the 5' RRS of the third vector is identical to the third RRS, and the 3' RRS of the third vector is identical to the fourth RRS, wherein the first additional RRS is contained in the first locus between the first and second RRS. (Item 65) 65. The system of Item 64, wherein the 3' RRS of the first vector is positioned at the 3' end of the 5' portion of a selectable marker gene contained in the first vector, and the 5' RRS of the second vector is positioned at the 5' end of the remaining selectable marker gene contained in the second vector. (Item 66) Item 62. The system of item 61, further comprising a nucleotide sequence encoding a second LCF. (Item 67) Item 67. The system of item 66, wherein the first and second LCFs are the same. (Item 68) 68. The system of claim 67, wherein the nucleotide sequence encoding the second LCF is in the second nucleic acid of the second vector, wherein the 5' RRS of the first vector is identical to the first RRS, the 3' RRS of the first vector is identical to the second RRS, the 5' RRS of the second vector is identical to the third RRS, and the 3' RRS of the second vector is identical to the fourth RRS. (Item 69) 68. The system of item 67, wherein the nucleotide sequence encoding the second LCF is in a separate, third vector and is flanked by a 5' RRS of the third vector and a 3' RRS of the third vector. (Item 70) 70. The system of item 69, wherein the 5' RRS and 3' RRS of the first vector are identical to the first RRS and second RRS, respectively, in the first locus; the 5' RRS of the third vector is identical to the third RRS; the 3' RRS of the third vector is identical to the 5' RRS of the second vector and to an additional RRS present between the third RRS and fourth RRS in the second locus; and the 3' RRS of the second vector is identical to the fourth RRS. (Item 71) 71. The system of Item 70, wherein the 3' RRS of the third vector is positioned at the 3' end of the 5' portion of the selectable marker gene contained in the third vector, and the 5' RRS of the second vector is positioned at the 5' end of the remaining 3' portion of the selectable marker gene contained in the second vector. (Item 72) 61. The system of item 60, wherein the nucleotide sequence encoding the first HCF encodes a first CH3 domain and the nucleotide sequence encoding the second HCF encodes a second CH3 domain. (Item 73) 73. The system of claim 72, wherein the first and second CH3 domains differ by at least one amino acid. (Item 74) 73. The system of claim 72, wherein the nucleotide sequences encoding the first and second CH3 domains differ in that one of the nucleotide sequences is codon modified. (Item 75) 52. The system of item 51, wherein the nucleotide sequence encoding the first HCF and the nucleotide sequence encoding the first LCF are each operably linked to a promoter. (Item 76) 52. The system of claim 51, wherein the cells are CHO cells. (Item 77) 77. The system of Item 76, wherein the two expression-enhanced loci include a locus having the nucleotide sequence of SEQ ID NO: 1 and a locus having the nucleotide sequence of SEQ ID NO: 2. (Item 78) (i) providing a system according to any one of items 51 to 77; (ii) introducing the vector into the cell by transfection; and (iii) selecting transfected cells in which the nucleic acid in the vector has been integrated into the first and second expression-enhancing loci via recombination mediated by the RRS. (Item 79) (iv) expressing and obtaining the antigen binding protein from the selected transfected cells. (Item 80) 24. A method for producing an antigen-binding protein, the method comprising providing a cell according to any one of items 1 to 23, and expressing and obtaining the antigen-binding protein from the cell. [Brief explanation of the drawings]
[0056] [Figure 1]Figure 1. Example antibody cloning strategies comparing integration into multiple expression-enhancing loci versus integration into a single expression-enhancing locus for traditional monospecific antibodies. Two vectors were transfected into host cells. The first vector carried nucleic acid encoding antibody chain 1 (AbC1), e.g., the light chain, and the second vector carried nucleic acid encoding antibody chain 2 (AbC2), e.g., the heavy chain, and a selectable marker distinct from the marker integrated into the target locus in the host cell. From 5' to 3', the RRS1, middle RRS3, and RRS2 sites of the vector construct match the RRS sites flanking the selectable marker in the host cell locus. An additional vector transfected into the host cell encodes a recombinase. If the selectable marker on the second vector is an antibiotic resistance gene, the two vectors are engineered to link and express the marker, allowing selection of positive recombinant clones that grow in the antibiotic. Alternatively, a fluorescent marker allows selection of positive clones using fluorescence-activated cell sorting (FACS) analysis. The same vector may be utilized for site-specific integration at a single locus, such as the EESYR® locus (Locus 1).
[0057] [Figure 2] Figure 2. Using a two-vector cloning strategy for integration within two expression-enhancing loci and one expression-enhancing locus, antibody A and antibody B were cloned into the loci shown in Figure 1. Antibody-expressing cells were isolated and subjected to fed-batch culture for 12 days. They were then harvested and subjected to an Octet titer assay using a Protein A sensor. Cells were also observed to be isogenic and stable. An overall titer increase of 0.5-0.9-fold was observed using the two-site integration method.
[0058] [Figure 3]Figure 3. Example antibody cloning strategies comparing integration into two separate expression-enhancing sites versus integration into a single expression-enhancing locus for a bispecific antibody encoded by three antibody chains. Three vectors were utilized for this bispecific strategy. The first vector carries nucleic acid encoding antibody chain 1 (AbC1), e.g., the common light chain, flanked by RRS1 and RRS3 for integration into EESYR® (SEQ ID NO: 1; Locus 1). The second vector carries antibody chain 2 (AbC2), e.g., the heavy chain, with an upstream selectable marker flanked by RRS4 and RRS6 for integration into the locus of SEQ ID NO: 2. The third vector then carries a nucleic acid encoding a second copy of AbC1 linked to a selectable marker different from that in the second vector and flanked 5' by RRS4 and 3' by RRS5 in a vector cassette (the 5' and 3' RRSs matched the RRS sites in the host cell at the locus with SEQ ID NO:2). The two-vector system shown may also be utilized for site-specific integration at a single locus, such as the EESYR® locus (with SEQ ID NO:1; Locus 1). The titer of each producer cell line was analyzed. See Figure 5.
[0059] [Figure 4]Figure 4. Example antibody cloning strategies comparing integration into two separate expression-enhancing sites versus integration into a single expression-enhancing locus for bispecific antibodies encoded by three or four antibody chains. Four vectors are utilized in this bispecific strategy. The first vector carries nucleic acid encoding antibody chain 1 (AbC1), e.g., the first light chain, flanked by RRS1 and RRS3 for integration into EESYR® (SEQ ID NO: 1; Locus 1). The second vector carries antibody chain 2 (AbC2), e.g., the heavy chain, with an upstream selectable marker flanked 5' by RRS3 and 3' by RRS2, which is also integrated into the EESYR® locus (with SEQ ID NO: 1; Locus 1). The third vector then encodes a selectable marker different from the selectable marker in the second vector, and carries nucleic acid linked in a vector cassette to a second (different) heavy chain, such as antibody chain 3 (AbC3), flanked 5' by RRS6 and 3' by RRS5 (the 5' and 3' RRSs match the RRS sites in the host cell at locus 2 with SEQ ID NO:2). The fourth vector then carries nucleic acid encoding a second light chain, such as antibody chain 1 (AbC1), which may be the same as or different from the first light chain. The four-vector system shown may be utilized for site-specific integration at two loci, such as the EESYR® locus (with SEQ ID NO:1; locus 1) and a locus with SEQ ID NO:2 or SEQ ID NO:3. This four-vector system is compared to a two-vector system integration at one locus (the EESYR® locus; locus 1).
[0060] [Figure 5]Figure 5. Bispecific antibody C and bispecific antibody D were cloned using a two-vector cloning strategy to compare integration into two expression-enhancing loci versus one expression-enhancing locus. Cells were isolated and subjected to 12-day fed-batch culture, then harvested and subjected to Octet titer assays using immobilized anti-Fc and a secondary anti-Fc* (a modified Fc detection antibody; see US 2014-0134719 A1, published May 15, 2014). Cells were also observed to be isogenic and stable. Overall bispecific titers (heterodimer formation) were observed to increase 1.75- to 2-fold using the two-site integration method compared to bispecific (heterodimer) expression at a single integration site.
[0061] [Figure 6] Figures 6A and 6B show that bispecific antibody E (Ab E), bispecific antibody F (Ab F), bispecific antibody G (Ab G), and bispecific antibody H (Ab H) were cloned in either RSX or RSX2BP. Each bispecific-expressing cell line contains the (common) light chain nucleotide, heavy chain nucleotide (wild-type Fc), and modified heavy chain nucleotide (Fc*) at either the 1 expression-enhanced locus (RSX) or the 2 expression-enhanced locus (RSX2BP). Cells were isolated and subjected to 13-day fed-batch culture in a bioreactor, then harvested and subjected to HPLC elution to determine overall and bispecific antibody titers (Figure 6A). The ratio of bispecific antibody species titers (purified and separated from homodimeric species) per total antibody titer was determined as a percentage of total Ab (Figure 6B).
[0062] [Figure 7] Figure 7. Monospecific antibodies J and K (Ab J and Ab K, respectively) were expressed in RSX or RSX2. Cells were isolated and subjected to 13 days of fed-batch culture in a bioreactor, then harvested and subjected to HPLC elution to determine overall IgG titers. DETAILED DESCRIPTION OF THE INVENTION
[0063] Definition of Terms The term "antibody," as used herein, includes immunoglobulin molecules composed of four polypeptide chains: two heavy (H) chains and two light (L) chains inter-connected by disulfide bonds. Each heavy chain may comprise a heavy chain variable region (abbreviated herein as HCVR or VH) and a heavy chain constant region. The heavy chain constant region comprises three domains, CH1, CH2, and CH3, and a hinge. Each light chain comprises a light chain variable region (abbreviated herein as LCVR or VL) and a light chain constant region. The light chain constant region comprises one domain (CL). The VH and VL regions can be further subdivided into regions of hypervariability called complementarity-determining regions (CDRs), which are interrupted by more conserved regions called framework regions (FRs). Each VH and VL is composed of three CDRs and four FRs, arranged from the amino terminus to the carboxy terminus in the following order: FR1, CDR1, FR2, CDR2, FR3, CDR3, FR4 (heavy chain CDRs may also be abbreviated as HCDR1, HCDR2, and HCDR3, and light chain CDRs may also be abbreviated as LCDR1, LCDR2, and LCDR3).
[0064] The term "antigen binding protein" includes proteins having at least one CDR and capable of selectively recognizing an antigen, i.e., at least micromolar levels. KD Therapeutic antigen-binding proteins (e.g., therapeutic antibodies) are often capable of binding to antigens in nanomolar or picomolar ranges. KD Typically, an antigen-binding protein comprises two or more CDRs, e.g., two, three, four, five, or six CDRs. Examples of antigen-binding proteins include antibodies, antigen-binding fragments of antibodies, such as polypeptides containing the variable regions of the heavy and light chains of an antibody (e.g., Fab fragments (F(ab')2 fragments)), proteins containing the variable regions of the heavy and light chains of an antibody and additional amino acids from the constant regions of the heavy and / or light chains (e.g., one or more constant domains, i.e., one or more of the CL, CH1, hinge, CH2, and CH3 domains).
[0065] The term "bispecific antigen-binding protein" includes antigen-binding proteins that can selectively bind to, or have different specificities for, two or more epitopes, either on two different molecules (e.g., antigens) or on the same molecule (e.g., the same antigen). The antigen-binding portion, or antigen-binding fragment portion (Fab) of such a protein, confers specificity for a particular antigen and is typically composed of immunoglobulin heavy and light chain variable regions. In some circumstances, the heavy and light chain variable regions may not be a cognate pair, or may have different binding specificities.
[0066] An example of a bispecific antigen-binding protein is a "bispecific antibody," which includes antibodies that can selectively bind to two or more epitopes. Bispecific antibodies often comprise two different heavy chains, each of which specifically binds to a different epitope, either on two different molecules (e.g., antigens) or on the same molecule (e.g., the same antigen). When a bispecific antigen-binding protein can selectively bind to two different epitopes (a first epitope and a second epitope), the affinity of the variable region of the first heavy chain for the first epitope is generally at least one to two orders of magnitude, or three or four orders of magnitude, lower than the affinity of the variable region of the first heavy chain for the second epitope, or vice versa. Bispecific antigen-binding proteins, such as bispecific antibodies, may comprise heavy chain variable regions that recognize different epitopes of the same antigen. A typical bispecific antibody has two heavy chains, each containing three heavy-chain CDRs followed (from N- to C-terminus) by a CH1 domain, a hinge, a CH2 domain, and a CH3 domain. A typical bispecific antibody also has an immunoglobulin light chain, which does not confer antigen-binding specificity but can associate with each heavy chain, or can associate with each heavy chain and bind one or more epitopes bound by the heavy-chain antigen-binding region, or can associate with each heavy chain and bind one or both epitopes of the heavy chain. In one embodiment, the Fc domain comprises at least a CH2 and a CH3. The Fc domain may comprise a hinge, a CH2 domain, and a CH3 domain.
[0067] One embodied bispecific format comprises a first heavy chain (HC), a second heavy chain with a modified CH3 (HC*), and a common light chain (LC) (two copies of the same light chain). In another embodiment, it comprises a first heavy chain (HC), a common LC, and an HC-ScFv fusion polypeptide (where the second HC is fused to the N-terminus of the ScFv). In another embodiment, it comprises a first heavy chain (HC), a cognate LC, and an HC-ScFv fusion polypeptide (where the second HC is fused to the N-terminus of the ScFv). In another embodiment, it comprises a first heavy chain (HC), an LC, and an Fc domain. In another embodiment, it comprises a first HC, an LC, and an ScFv-Fc fusion polypeptide (where the Fc is fused to the C-terminus of the ScFv). In another embodiment, it comprises a first HC, a common LC, and an Fc-ScFv fusion polypeptide (where the Fc is fused to the N-terminus of the ScFv). In another embodiment, it comprises a first HC, an LC, and an ScFv-HC (where the second HC is fused to the C-terminus of the ScFv).
[0068] In some embodiments, one heavy chain (HC) may be the native or "wild-type" sequence and the second heavy chain may be Fc domain modified. In other embodiments, one heavy chain (HC) may be the native or "wild-type" sequence and the second heavy chain may be codon modified.
[0069] The term "cell" includes any cell suitable for expression of a recombinant nucleic acid sequence and having a locus that allows for stable integration and enhanced expression of exogenous nucleic acid. Cells include mammalian cells, such as non-human animal cells, human cells, or cell fusions, e.g., hybridomas or quadromas. In some embodiments, the cells are human, monkey, ape, hamster, rat, or mouse cells. In some embodiments, the cell is a mammalian cell selected from the following cells: CHO (e.g., CHO K1, DXB-11 CHO, Veggie-CHO), COS (e.g., COS-7), retinal cells, Vero, CV1, kidney (e.g., HEK293, 293 EBNA, MSR 293, MDCK, HaK, BHK), HeLa, HepG2, WI38, MRC 5, Colo205, HB 8065, HL-60, (e.g., BHK21), Jurkat, Daudi, A431 (epidermal), CV-1, U937, 3T3, L cells, C127 cells, SP2 / 0, NS-0, MMT 060562, Sertoli cells, BRL 3A cells, HT1080 cells, myeloma cells, tumor cells, and cell lines derived from the foregoing cells. In some embodiments, the cells are retinal cells expressing one or more viral genes, e.g., viral genes (e.g., PER.C6 TM cells).
[0070] "Cell density" refers to the number of cells per sample volume, e.g., total number of cells (live and dead) per mL. Cell counts may be performed manually or automatically, e.g., using a flow cytometer. Automated cell counters are adapted to count live or dead cells or both live and dead cells, e.g., using standard methods such as trypan blue uptake. The phrase "viable cell density" or "viable cell concentration" refers to the number of live cells per sample volume (also referred to as "viable cell count"). Many known manual or automated techniques may be used to determine cell density. Online biomass measurements of the culture may be taken, with capacitance or optical density correlating to the number of cells per volume. Final cell densities in cell cultures, e.g., production cultures, may range from, e.g., about 1.0 to 10 x 10, depending on the starting cell line. 6 In some embodiments, the final cell density varies between 1.0 and 10x10 cells / mL prior to harvesting the protein of interest from the production cell culture. 6 In other embodiments, the final cell density reaches 5.0 x 10 cells / mL. 6 >cells / mL, 6x10 6 >cells / mL, 7x10 6 >cells / mL, 8x10 6 >cells / mL, 9x10 6 cells / mL, or 10x10 6 cells / mL.
[0071] The term "codon-modified" means that a protein-encoding nucleotide sequence has been modified at one or more nucleotides, i.e., one or more codons, without changing the amino acid encoded by that codon, resulting in a codon-modified version of the nucleotide sequence. The codon modification of a nucleotide sequence can provide a convenient basis for distinguishing the nucleotide sequence from its codon-modified version in nucleic acid-based assays (e.g., hybridization-based assays, PCR, etc., among others). In some instances, codons of a nucleotide sequence are modified to improve or optimize expression of the encoded protein in a host cell using codon optimization techniques known in the art (Gustafsson, C., et al., 2004, Trends in Biotechnology, 22:346-353; Chung, BK-S., et al., 2013, Journal of Biotechnology, 167:326-333; Gustafsson, C., et al., 2012, Protein Expr Purif, 83(1):37-46). Sequence design software tools using such techniques are also known in the art and include, but are not limited to, Codon optimizer (Fuglsang A. 2003, Protein Expr Purif, 31:247-249), Gene Designer (Villalobos A, et al., 2006, BMC Bioinforma, 7:285), and OPTIMIZER (Puigbo P, et al. 2007, Nucleic Acids Research, 35:W126-W131), among others.
[0072] The term "complementarity-determining region" or "CDR" includes an amino acid sequence encoded by a nucleic acid sequence of an immunoglobulin gene of an organism, which amino acid sequence is normally (i.e., in wild-type animals) located between two framework regions in the variable region of the light or heavy chain of an immunoglobulin molecule (e.g., an antibody or T-cell receptor). CDRs can be encoded, for example, by germline sequences or rearranged or unrearranged sequences, and can be encoded, for example, by naive B cells or mature B cells or T cells. Under some circumstances (e.g., with respect to CDR3), a CDR can be encoded by two or more sequences (e.g., germline sequences) that are not contiguous (e.g., in the unrearranged nucleic acid sequence) but are contiguous in the B-cell nucleic acid sequence as a result of splicing or joining of sequences (e.g., VDJ rearrangement to form the heavy chain CDR3).
[0073] The term "enhanced expression locus" refers to a locus in a cellular genome that contains a sequence and exhibits increased expression levels relative to other regions or sequences in the genome when an appropriate gene or construct is exogenously added (i.e., integrated) into or near the sequence, or is "operably linked" to the sequence.
[0074] When used to describe enhanced expression, the term "enhanced" includes at least about 1.5-fold to at least about 3-fold enhanced expression over expression typically observed with random integration of an exogenous sequence into a genome or integration at another locus, e.g., compared to various random integrations of a single copy of the same expression construct. The fold enhanced expression observed using the sequences of the invention is compared to the expression level of the same gene, measured under substantially the same conditions in the absence of the sequences of the invention, e.g., when integrated at another locus in the homologous genome. Enhanced recombination efficiency includes enhanced recombination capacity of a locus (e.g., utilization of recombinase recognition sites (RRSs)). Enhancement refers to recombination efficiency over random recombination, typically 0.1% when no recombinase recognition sites or homologs are utilized. Preferred enhanced recombination efficiencies are about 10-fold over random or about 1%. Unless specified, the claimed inventions are not limited to a particular recombination efficiency. Enhanced expression loci often support higher production of a protein of interest by host cells. Enhanced expression therefore involves higher production of the protein of interest per cell (higher titer per gram of protein) rather than simply achieving higher titer through higher copy number of cells in culture. Specific productivity, Qp (pg / cell / day, or pcd), is considered a measure of sustainable productivity. Recombinant host cells that exhibit a Qp of greater than 5 pcd, greater than 10 pcd, or greater than 15 pcd, or greater than 20 pcd, or greater than 25 pcd, or even greater than 30 pcd are desirable. Host cells with inserted genes of interest at expression-enhancing loci or "hot spots" exhibit high specific productivity.
[0075] The terms "exogenously added gene," "exogenously added nucleic acid," or simply "exogenous nucleic acid," when used in reference to a locus of interest, refer to any DNA sequence or gene that is not present within the locus of interest when the locus exists in nature. For example, an "exogenous nucleic acid" within a CHO locus (e.g., a locus having the sequence of SEQ ID NO: 1 or SEQ ID NO: 2) can be a hamster gene that is not present within the particular CHO locus in nature (i.e., a hamster gene from another locus in the hamster genome), a gene from any other species (e.g., a human gene), a chimeric gene (e.g., human / mouse), or any other gene not known to be present within the CHO locus in nature.
[0076] The terms "heavy chain" or "immunoglobulin heavy chain" include immunoglobulin heavy chain constant region sequences from any organism, including heavy chain variable domains unless otherwise specified. Unless otherwise specified, heavy chain variable domains include three heavy chain CDRs and four FR regions. A typical heavy chain comprises (from N- to C-terminus) the variable domain followed by a CH1 domain, hinge, CH2 domain, and CH3 domain. The term "heavy chain fragment" includes a peptide of at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more amino acids of a heavy chain, and may include one or more CDRs combined with one or more CDRs, one or more FRs, CH1, hinge, CH2, or CH3, variable region, constant region, fragments of a constant region (e.g., CH1, CH2, CH3), or combinations thereof. Examples of HCFs include VHs and all or part of an Fc region. The phrase "nucleotide sequence encoding HCF" includes nucleotide sequences encoding polypeptides consisting of HCF and polypeptides containing HCF, including polypeptides that may contain additional amino acids in addition to the particular HCF. For example, nucleotide sequences encoding HCF specifically include nucleotide sequences encoding polypeptides consisting of VH, polypeptides consisting of VH linked to CH3, and polypeptides consisting of a full-length heavy chain.
[0077] "Homologous sequence," in the context of nucleic acid sequences, refers to a sequence that is substantially homologous to a reference nucleic acid sequence. In some embodiments, two sequences are considered to be substantially homologous if at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more of their corresponding nucleotides are identical over the relevant stretch of residues. In some embodiments, the relevant stretch is the complete (i.e., full-length) sequence.
[0078] The term "light chain" includes immunoglobulin light chain constant region sequences from any organism, including human kappa light chains and human lambda light chains, unless otherwise specified. A light chain variable (VL) domain typically contains three light chain CDRs and four framework (FR) regions, unless otherwise specified. A full-length light chain generally contains, from the amino terminus to the carboxyl terminus, a VL domain comprising FR1-CDR1-FR2-CDR2-FR3-CDR3-FR4 and a light chain constant domain. Light chains that can be used in the present invention include, for example, light chains that do not selectively bind to either the first or second epitope selectively bound by a bispecific antibody. Suitable light chains also include light chains that can bind to or contribute to binding to one or both epitopes bound by the antigen-binding region of an antibody. The terms "light chain fragment" or "light chain fragment" (or "LCF") include peptides of at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more amino acids of a light chain, and may include one or more CDRs combined with one or more CDRs, one or more FRs, a variable region, a constant region, a fragment of a constant region, or a combination thereof. Examples of LCFs include all or a portion of VLs and light chain constant regions (CLs). The term "nucleotide sequence encoding an LCF" includes nucleotide sequences that encode polypeptides consisting of an LCF as well as nucleotide sequences that encode polypeptides containing an LCF, including, for example, polypeptides that may contain additional amino acids in addition to the particular LCF. For example, a nucleotide sequence encoding an LCF specifically includes a nucleotide sequence that encodes a polypeptide consisting of a VL or a polypeptide consisting of a full-length light chain.
[0079] The phrase "operably linked" refers to the linkage of nucleic acids or proteins in a manner that allows the linked molecules to function as intended. DNA regions are operably linked when they are functionally related to each other. For example, a promoter is operably linked to a coding sequence if it is capable of mediating transcription of the sequence; a ribosome binding site is operably linked to a coding sequence if it is positioned to permit translation. Generally, operably linked can include, but does not require, contiguity. In the case of sequences such as secretory leaders, contiguity and proper placement in reading frame are typical characteristics. An expression-enhancing sequence of a locus of interest is operably linked to a gene of interest (GOI) if it is functionally associated with the GOI, e.g., if its presence results in enhanced expression of the GOI.
[0080] "Percent identity," when describing a subject locus, e.g., SEQ ID NO: 1 or SEQ ID NO: 2, or a fragment thereof, is meant to include homologous sequences that show identity along the contiguous region of homology, but non-homologous gaps, deletions, or insertions in the comparison sequence are not taken into account in calculating percent identity.
[0081] As used herein, the determination of "percent identity" between, for example, SEQ ID NO: 1 or a fragment thereof and a species homolog does not include comparisons of sequences where the species homolog does not have a homologous sequence to compare in an alignment (i.e., it does not include comparisons of sequences where SEQ ID NO: 1 or a fragment thereof has an insertion at that point, or where the species homolog possibly has a gap or deletion). Thus, "percent identity" does not include penalties for gaps, deletions, and insertions.
[0082] A "recognition site" or "recognition sequence" is a specific DNA sequence recognized by a nuclease or other enzyme that binds to the DNA backbone and directs site-specific cleavage. Endonucleases cleave DNA within a DNA molecule. Recognition sites are also referred to in the art as recognition target sites.
[0083] A "recombinase recognition site" ("RRS") is a specific DNA sequence recognized by a recombinase, such as Cre recombinase (Cre) or flippase (flp). Site-specific recombinases can carry out DNA rearrangements, including deletions, rearrangements, and translocations, when one or more of their target recognition sequences are strategically placed within an organism's genome. In one example, Cre specifically mediates recombination events with its DNA target recognition site, loxP, which consists of two 13-bp inverted repeats separated by an 8-bp spacer. Two or more recombinase recognition sites can be used to facilitate, for example, recombination-mediated DNA exchange. Variants or mutants of recombinase recognition sites, such as lox sites, can also be used (Araki, N. et al., 2002, Nucleic Acids Research, 30:19, e103).
[0084] "Recombinase-mediated cassette exchange" or "RMCE" refers to the process of precisely replacing a genomic target cassette with a donor cassette. The molecular configuration typically provided to carry out this process includes: 1) a genomic target cassette flanked on both the 5' and 3' ends by recognition target sites specific to a particular recombinase; 2) a donor cassette flanked by matching recognition target sites; and 3) a site-specific recombinase. Recombinase proteins are known in the art (Turan, S. and Bode J., 2011, FASEB J., 25, pp. 4088-4107) and are capable of precisely cleaving DNA within a specific recognition target site (DNA sequence) without adding or deleting nucleotides. Common recombinase / site combinations include, but are not limited to, Cre / lox and Flp / frt. Vectors containing the R4-attP site and encoding the phiC31 integrase for RMCE are also provided in commercially available kits (see, e.g., U.S. Patent Application Publication No. US20130004946).
[0085] "Site-specific integration" or "targeted insertion" refers to a gene targeting method used to direct the insertion or integration of a gene or nucleic acid sequence into a specific location in the genome, i.e., to move DNA to a specific site between two nucleotides in a continuous polynucleotide chain. Site-specific integration or targeted insertion may be performed on a particular nucleic acid that contains multiple expression units or cassettes, for example, multiple genes, each with its own regulatory elements (e.g., promoters, enhancers, and / or transcription termination sequences). "Insertion" and "integration" are used interchangeably. It is understood that insertion of a gene or nucleic acid sequence (e.g., a nucleic acid sequence with an expression cassette) may result in (or be engineered to result in) the replacement or deletion of one or more nucleic acids depending on the gene editing technique used.
[0086] "Stable integration" means that the exogenous nucleic acid integrated into the host cell genome remains integrated for an extended period of time in cell culture, such as at least 7 days, at least 10 days, at least 15 days, at least 20 days, at least 25 days, at least 30 days, at least 35 days, at least 40 days, at least 45 days, at least 50 days, at least 55 days, at least 60 days, or longer. It is understood that producing bispecific antigen-binding proteins for large-scale production and purification is challenging. Stability and clonality are essential for the reproducibility of any biomolecule, particularly a biomolecule used for therapeutic purposes. The stable clones expressing bispecific antibodies generated by the disclosed methods provide a consistent and reproducible means for generating therapeutic biomolecules.
[0087] overview The present disclosure provides compositions and methods for improving expression of multiple polypeptides in host cells, particularly Chinese hamster (Cricetulus griseus) cell lines, by using multiple (e.g., two) expression-enhancing loci in the host cells. More specifically, the present disclosure provides compositions and methods designed to integrate multiple exogenous nucleic acids encoding antigen-binding proteins into multiple expression-enhancing loci in a site-specific manner in host cells, such as CHO cells. In particular, the present disclosure provides cells containing multiple exogenous nucleic acids integrated into multiple expression-enhancing loci, where the multiple exogenous nucleic acids together encode antigen-binding proteins. The present disclosure also provides nucleic acid vectors designed for site-specific integration of multiple exogenous nucleic acids into multiple expression-enhancing loci. The present disclosure further provides a system for site-specific integration of multiple exogenous nucleic acids from the vectors into the multiple expression-enhancing loci, including a host cell containing multiple recombinase recognition sites (RRSs) at each of the multiple expression-enhancing loci, and a vector set containing matching RRSs and multiple exogenous nucleic acids. The present disclosure further provides methods of making antigen binding proteins using the cells, vectors, and systems disclosed herein.
[0088] Cells with multiple exogenous nucleic acids site-specifically integrated into multiple expression-enhancing loci In one embodiment, the disclosure provides a cell containing a plurality of exogenous nucleic acids site-specifically integrated into two expression-enhanced loci, the plurality of exogenous nucleic acids together encoding antigen binding proteins, which may be bispecific or conventional (i.e., monospecific) antigen binding proteins.
[0089] The cells provided herein are capable of producing desired antigen binding proteins at high titers and / or high specific productivity (pg / cell / day). In some embodiments, the cells produce antigen binding proteins at titers of at least 1 g / L, 1.5 g / L, 2.0 g / L, 2.5 g / L, 3.0 g / L, 3.5 g / L, 4.0 g / L, 4.5 g / L, 5.0 g / L, 10 g / L or more. In some embodiments, the cells producing the antigen binding protein have a specific productivity of at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 picograms / cell / day or more, as determined based on total antigen binding protein (pg) produced per cell per day.
[0090] Host cells harboring exogenous nucleic acids encoding antigen binding proteins and integrated into two expression-enhancing loci can be grown in production culture at high cell densities, e.g., 1-10x10 6 In other embodiments, the host cells encoding the antigen binding protein have a final cell density (in the production culture) of at least 5x10 6 cells / mL, 6x10 6 cells / mL, 7x10 6 cells / mL, 8x10 6 cells / mL, 9x10 6 cells / mL, or 10x10 6 cells / mL.
[0091] In some embodiments, cells are provided that are capable of producing bispecific antigen-binding proteins, wherein the ratio of bispecific antigen-binding protein titer to total antigen-binding protein titer is at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 50%, 60% or more. In some embodiments, cells are provided that are capable of producing bispecific antigen-binding proteins, wherein the ratio of bispecific antigen-binding protein is at least 50% of the total antigen-binding protein titer produced by the cells.
[0092] In other embodiments, cells capable of producing antigen binding proteins are provided, wherein the total antigen binding protein titer produced by expression at two loci is at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 50%, 60% or more higher than the total antigen binding protein titer produced by expression at a single locus. In certain embodiments, cells capable of producing antigen binding proteins are provided, wherein the total antigen binding protein titer produced by expression at two loci is at least 0.5-fold, 0.75-fold, 1-fold, 1.5-fold, 1.75-fold, 2-fold or more higher than the total antigen binding protein titer produced by expression at a single locus.
[0093] In some embodiments, the cell contains a first exogenous nucleic acid integrated at a specific site within a first expression-enhanced locus and a second exogenous nucleic acid integrated at a specific site within a second expression-enhanced locus, wherein the first and second exogenous nucleic acids together encode an antigen-binding protein. The first and second exogenous nucleic acids together contain multiple nucleotide sequences encoding the HCF or LCF (e.g., variable regions) of the antigen-binding protein. For example, with respect to a monospecific antibody, there may be one nucleotide sequence encoding the HCF (e.g., VH) and one nucleotide sequence encoding the LCF (e.g., VL) for the monospecific antigen-binding protein, or multiple copies (e.g., two copies) of each. With respect to a bispecific antibody, there may be two nucleotide sequences each encoding an HCF (typically, the two HCFs are different from each other), one or two copies of a nucleotide sequence encoding an LCF, or two nucleotide sequences encoding two different LCFs. Depending on whether the antigen-binding protein is monospecific or bispecific, the nucleotide sequences encoding the HCF or LCF (e.g., variable regions) may be integrated in different varieties or combinations at the two expression-enhanced loci. For example, with respect to a monospecific antigen-binding protein, in one example, a nucleotide sequence encoding an HCF and a nucleotide sequence encoding an LCF may be integrated separately at two loci, one at each locus. Alternatively, in another example, nucleic acids encoding an HCF and an LCF may be integrated at one locus, and separate nucleic acids encoding the same HCF and the same LCF may be integrated at another locus. With respect to a bispecific antigen-binding protein, in one example, a nucleotide sequence encoding a first HCF and an LCF may be integrated at a first locus, and a nucleotide sequence encoding a second HCF may be integrated at a second locus. In this case, the two HCFs are different, and the LCF is a common LCF for the bispecific antigen-binding protein. Alternatively, in another example, a nucleotide sequence encoding a first HCF and an LCF may be integrated at a first locus, and a nucleotide sequence encoding a second HCF and the same LCF may be integrated at a second locus.
[0094] In some embodiments, the nucleotide sequence encoding HCF or LCF may encode amino acids from the constant region. For example, the nucleotide sequence encoding HCF or LCF may encode one or more of CL, CH1, CH2, CH3, or a combination of CH1, CH2, or CH3, or may encode the entire constant region. In some embodiments, the nucleotide sequence encoding HCF encodes the CH3 domain. In certain embodiments, the nucleotide sequence encoding HHCF encodes a heavy chain. In some embodiments, the nucleotide sequence encoding LCF encodes a light chain.
[0095] In embodiments involving two HCFs, the nucleotide sequence encoding the first HCF may encode amino acids from a first constant region, and the nucleotide sequence encoding the second HCF may encode amino acids from a second constant region, where the amino acids from the two constant regions may be the same or different at at least one position (e.g., a position that results in different Protein A binding properties, or other positions described herein below for various bispecific antigen-binding proteins). Regardless of the amino acid difference, the two nucleotide sequences encoding amino acids from the two constant regions can be distinguished by modifying one or more codons in one nucleotide sequence, thereby providing a convenient basis for distinguishing the two nucleotide sequences in nucleic acid-based assays.
[0096] In some embodiments, the nucleotide sequences encoding each HCF or LCF are independently operably linked to a transcriptional control sequence, including a promoter. "Independently" means that each coding sequence is operably linked to a separate transcriptional control sequence, e.g., a promoter, such that transcription of the coding sequences is under separate control and management. In some embodiments, the promoters directing transcription of the two HCF-containing polypeptides are the same. In some embodiments, the promoters directing transcription of the two HCF-containing polypeptides and the promoter directing transcription of the VL-containing polypeptide are all the same, e.g., a CMV promoter. In some embodiments, the nucleotide sequences encoding each HCF or LCF are independently operably linked to an inducible or repressible promoter. Inducible and repressible promoters allow production to occur only during the production phase (fed-batch culture) and not during the growth phase (seed train culture). Alternatively, expression of antibody components (HCF and LCF) at different loci can be specifically and precisely controlled. Better control of the production (expression) of each gene product can be achieved by using different promoters.
[0097] In one such example, cells are first engineered to express the tetracycline repressor protein (TetR), and each HCF-encoding nucleotide sequence and each LCF-encoding nucleotide sequence are placed under the transcriptional control of a promoter whose activity is controlled by TetR. Two tandem TetR operators (tetO) are positioned immediately downstream of a CMV promoter. In some embodiments, each HCF- and / or LCF-encoding nucleotide sequence is independently operably linked to a promoter upstream of at least one TetR operator (TetO) or Arc operator (ArcO). In other embodiments, each HCF- and / or LCF-encoding nucleotide sequence is independently operably linked to a CMV / TetO or CMV / ArcO hybrid promoter. Additional suitable promoters are described herein below.
[0098] In some embodiments, multiple exogenous nucleic acids integrated at two loci are flanked by RRSs. For example, a first RRS and a second RRS are located 5' and 3', respectively, to a first exogenous nucleic acid integrated at a first locus, and a third RRS and a fourth RRS are located 5' and 3', respectively, to a second exogenous nucleic acid integrated at a second locus, where the first RRS and the second RRS are different, and the third RRS and the fourth RRS are different. In some embodiments, the first, second, third, and fourth RRSs are all different from one another. In other embodiments, the first RRS and the third RRS are identical, and the second RRS and the fourth RRS are identical, where the first exogenous nucleic acid encodes HCF and LCF, and the second exogenous nucleic acid encodes the same HCF and LCF.
[0099] In some embodiments, where the exogenous nucleic acid integrated at a locus contains nucleotide sequences encoding two HCFs or LCFs, an additional RRS may be contained between the two nucleotide sequences. Such additional RRS must be different from the two RRSs flanking the exogenous nucleic acid. In some embodiments, the additional RRS is inserted within an intron of a selectable marker gene contained within the integrated exogenous nucleic acid and positioned between the nucleotide sequences encoding the two HCFs or LCFs. After transcription and post-transcriptional processing, the intron is excised, resulting in an mRNA encoding the selectable marker. In embodiments, where the exogenous nucleic acid integrated at a first locus and the exogenous nucleic acid integrated at a second locus both contain sequences encoding two HCFs or LCFs (e.g., LCF-HCF1 and LCF-HCF2), an additional RRS may be contained in only one or both of the first and second exogenous nucleic acids at each locus, between the sequences encoding the two HCFs or LCFs. The additional RRS at the first locus may be the same as or different from the additional RRS at the second locus. Each additional RRS must be different from the two RRSs flanking the exogenous nucleic acid integrated at that locus. Optionally, each additional RRS may be inserted within an intron of a selectable marker gene, and the selectable marker genes with introns at the two loci may be different.
[0100] When multiple HCF or LCF-encoding sequences are contained within an exogenous nucleic acid integrated into a locus, the relative positions of the multiple coding sequences within the locus can vary. For example, in embodiments where the integrated exogenous nucleic acid contains a nucleotide sequence encoding LCF and a nucleotide sequence encoding HCF, the nucleotide sequence encoding LCF can be located upstream or downstream relative to the nucleotide sequence encoding HCF. In certain embodiments, the nucleotide sequence encoding LCF is located upstream relative to the nucleotide sequence encoding HCF. In certain embodiments where both loci contain a nucleotide sequence encoding LCF and a nucleotide sequence encoding HCF, the nucleotide sequence encoding LCF is located upstream or downstream relative to the nucleotide sequence encoding HCF at both loci.
[0101] In a further embodiment, a cell is provided that contains a first pair of RRSs integrated into a first expression-enhanced locus and a second pair of RRSs integrated into a second expression-enhanced locus, where the two RRSs in each pair are different. Such a cell is useful for receiving multiple integrated exogenous nucleic acids that together encode antigen binding proteins.
[0102] In some embodiments, the first exogenous nucleic acid is present between two RRSs of the first locus, and the second exogenous nucleic acid is present between two RRSs of the second locus. The first and second exogenous nucleic acids may each encode one or more selectable marker genes. The selectable marker genes may be different from each other.
[0103] In some embodiments, an additional RRS is present between two RRSs (i.e., the 5' RRS and the 3' RRS) in a pair of loci, where the additional RRS is different from both the 5' RRS and the 3' RRS of the locus. In some embodiments, the additional RRS is present between the 5' RRS and the 3' RRS at one of the two loci. In other embodiments, the additional RRS is present between the 5' RRS and the 3' RRS at each of the two loci. When an additional RRS is present between the 5' RRS and the 3' RRS, a selectable marker gene may be included between the 5' RRS and the additional RRS, and another selectable marker gene may be included between the additional RRS and the 3' RRS, where the two selectable markers are different.
[0104] In many of the described embodiments, the cell is a CHO cell, in which case one of the two expression-enhanced loci is selected from the group consisting of a nucleotide sequence that is at least 90% identical to SEQ ID NO:1, a nucleotide sequence that is at least 90% identical to SEQ ID NO:2, and a nucleotide sequence that is at least 90% identical to SEQ ID NO:3.
[0105] Bispecific antigen-binding proteins Bispecific antigen-binding proteins, e.g., bispecific antibodies, suitable for cloning and production in the cells, vectors and systems described in this disclosure are not limited to any particular format of the bispecific antigen-binding protein.
[0106] In various embodiments, the bispecific antigen-binding protein comprises two polypeptides, each polypeptide containing an antigen-binding portion (e.g., HC) and a CH3 domain, wherein the antigen-binding portions of the two polypeptides have different antigen specificities, and wherein the two CH3 domains are heterodimeric with respect to each other in that one of the CH3 domains is modified at at least one amino acid position to result in different Protein A binding properties between the two polypeptides. See, for example, the bispecific antibodies described in U.S. Patent No. 8,586,713. In this manner, different Protein A isolation schemes can be used to readily isolate heterodimeric bispecific antigen-binding proteins from homodimers.
[0107] In some embodiments, the bispecific antigen binding protein comprises two heavy chains that have different antigen specificities and differ at least one amino acid position in the CH3 domain, resulting in different Protein A binding properties between the two heavy chains.
[0108] In some embodiments, the two polypeptides contain a CH3 domain of a human IgG, wherein one of the two polypeptides contains a CH3 domain of a human IgG selected from IgG1, IgG2, and IgG4, and the other of the two polypeptides contains a modified CH3 domain of a human IgG selected from IgG1, IgG2, and IgG4, wherein the modification reduces or abolishes binding of the modified CH3 region to Protein A. In particular embodiments, one of the two polypeptides contains a CH3 domain of a human IgG1, and the other of the two polypeptides contains a modified CH3 domain of a human IgG1, wherein the modification is selected from the group consisting of (i) 95R and (ii) 95R and 96F according to the IMGT exon numbering system. In other particular embodiments, the modified CH3 domain comprises 1 to 5 additional modifications selected from the group consisting of 16E, 18M, 44S, 52N, 57M, and 82I according to the IMGT exon numbering system.
[0109] In various other embodiments, the two polypeptides contain CH3 domains of mouse IgG, wherein one of the two polypeptides contains an unmodified CH3 domain of mouse IgG and the other of the two polypeptides contains a modified CH3 domain of mouse IgG, wherein the modification reduces or eliminates binding of the modified CH3 region to Protein A. In various embodiments, the mouse IgG CH3 domain is modified to comprise specific amino acids at specific positions (EU numbering) selected from the group consisting of: 252T, 254T, and 256T; 252T, 254T, 256T, and 258K; 247P, 252T, 254T, 256T, and 258K; 435R and 436F; 252T, 254T, 256T, 435R, and 436F; 252T, 254T, 256T, 258K, 435R, and 436F; 24tP, 252T, 254T, 256T, 258K, 435R, and 436F; and 435R. In certain embodiments, a particular group of modifications selected from the group consisting of: M252T, S254T, S256T; M252T, S254T, S256T, I258K; I247P, M252T, S254T, S256T, I258K; H435R, H436F; M252T, S254T, S256T, H435R, H436F; M252T, S254T, S256T, I258K, H435R, H436F; I247P, M252T, S254T, S256T, I258K, H435R, H436F; and H435R.
[0110] In various embodiments, the bispecific antigen-binding protein is a hybrid of mouse and rat monoclonal antibodies or antigen-binding proteins, such as a hybrid of mouse IgG2a and rat IgG2b. According to these embodiments, the bispecific antibody is composed of a heterodimer of two antibodies, each with one heavy / light chain pair, linked via their Fc portions. The described heterodimer can be easily purified from a mixture of the two original antibody homodimers and bispecific heterodimers because the binding properties of the bispecific antibody for Protein A are different from those of the original antibodies. Rat IgG2b does not bind to Protein A, whereas mouse IgG2a does. Consequently, the mouse-rat heterodimer binds to Protein A but elutes at a higher pH than the mouse IgG2a homodimer. This allows for selective purification of the bispecific heterodimer.
[0111] In various other embodiments, the bispecific antigen-binding proteins are of a type referred to in the art as "knobs-into-holes" (see, e.g., U.S. Pat. No. 7,183,076). In these embodiments, the Fc portions of two antibodies are engineered to create a protruding "knob" in one and a complementary "hole" in the other. When produced in the same cell, the combination of the engineered "knob" and the engineered "hole" is said to cause the heavy chains to preferentially form heterodimers over homodimers.
[0112] In other embodiments, the first and second heavy chains comprise one or more amino acid modifications in the CH3 domains that allow the two heavy chains to interact. Amino acid residues at the CH3-CH3 interface are replaced with charged amino acids, thereby electrostatically disfavoring homodimer formation. (See, e.g., PCT International Publication No. WO2009089004 and European Publication No. EP1870459.)
[0113] In other embodiments, the first heavy chain comprises a CH3 domain of isotype IgA and the second heavy chain comprises a CH3 domain of IgG (or vice versa), promoting preferential formation of heterodimers (see, e.g., PCT International Patent Application Publication No. WO2007110205).
[0114] In other embodiments, various formats can be incorporated with immunoglobulin chains by engineering methods that promote heterodimer formation, such as Fab-arm exchange (PCT International Patent Application Publication No. WO2008119353; PCT International Patent Application Publication No. WO2011131746), coiled-coil domain interactions (PCT International Patent Application Publication No. WO2011034605), or leucine zipper peptides (Kostelny, et al. J. Immunol. 1992, 148(5):1547-1553).
[0115] Immunoglobulin heavy chain fragments (e.g., variable regions) that can be used to generate bispecific antigen-binding proteins can be generated using any method known in the art. For example, a first heavy chain comprises a variable region encoded by a nucleic acid derived from the genome of a mature B cell of a first animal, the first animal being immunized with a first antigen, and the first heavy chain specifically recognizes the first antigen. A second heavy chain comprises a variable region encoded by a nucleic acid derived from the genome of a mature B cell of a second animal, the second animal being immunized with a second antigen, and the second heavy chain specifically recognizes the second antigen. Immunoglobulin heavy chain variable region sequences can also be obtained by any other method known in the art, such as phage display. In other examples, nucleic acids encoding heavy chain variable regions include those of antibodies reported in the art or otherwise available. In some embodiments, one of the two heavy chain coding sequences is codon-modified, thereby providing a convenient basis for distinguishing between the two coding sequences in nucleic acid-based assays.
[0116] Bispecific antibodies, with two heavy chains that recognize two different epitopes (or two different antigens), are more easily isolated if they can pair with the same light chain (i.e., the light chains have identical variable and constant domains). Various methods are known in the art for generating light chains that can pair with two heavy chains of different specificities without interfering or substantially interfering with the heavy chain variable domains and their selectivity and / or affinity for their target antigens, such as the techniques described and disclosed in U.S. Pat. No. 8,586,713.
[0117] Bispecific antigen-binding proteins may have a variety of dual antigen specificities and associated useful applications.
[0118] In some examples, bispecific antigen-binding proteins can be generated that have binding specificity for a tumor antigen and a T cell antigen, targeting an antigen on a cell, e.g., CD20, and also targeting an antigen on a T cell, e.g., a T cell receptor such as CD3. In this case, the bispecific antigen-binding protein targets both the target cells in the patient (e.g., B cells in a lymphoma patient via CD20 binding) as well as the patient's T cells. In various embodiments, the bispecific antigen-binding protein is designed to activate T cells upon CD3 binding, thereby coupling T cell activation to specific selected tumor cells.
[0119] In the context of a bispecific antigen-binding protein, when one portion binds to a T cell receptor, e.g., binds to CD3, and the other portion binds to the target antigen, the target antigen is a tumor-associated antigen. Non-limiting examples of specific tumor-associated antigens include, for example, AFP, ALK, BAGE protein, BIRC5 (survivin), BIRC7, β-catenin, brc-abl, BRCA1, BCMA, BORIS, CA9, carbonic anhydrase IX, caspase-8, CALR, CCR5, CD19, CD20 (MS4A1), CD22, CD30, CD40, CDK4, CEA, CLEC-12, CTL A4, cyclin-B1, CYP1B1, EGFR, EGFRvIII, ErbB2 / Her2, ErbB3, ErbB4, ETV6-AML, EpCAM, EphA2, Fra-1, FOLR1, GAGE proteins (e.g., GAGE-1, -2), GD2, GD3, GloboH, glypican-3, GM3, gp100, Her2, HLA / B-raf, HLA / k-ras, HLA / MA GE-A3, hTERT, LMP2, MAGE proteins (e.g., MAGE-1, -2, -3, -4, -6, and -12), MART-1, mesothelin, ML-IAP, Muc1, Muc2, Muc3, Muc4, Muc5, Muc16 (CA-125), MUM1, NA17, NY-BR1, NY-BR62, NY-BR85, NY-ESO1, OX40, p15, p53, PAP, PA X3, PAX5, PCTA-1, PLAC1, PRLR, PRAME, PSMA (FOLH1), RAGE protein, Ras, RGS5, Rho, SART-1, SART-3, Steap-1, Steap-2, TAG-72, TGF-β, TMPRSS2, Thompson-nouvelle antigen (Tn), TRP-1, TRP-2, tyrosinase, and uroplakin-3. In some embodiments, the bispecific antigen-binding protein comprises one moiety that binds to CD3. Examples of anti-CD3 antibody moieties are described in U.S. Patent Application Publications US2014 / 0088295A1 and US20150266966A1, and International Patent Application Publication WO2017 / 053856, published March 30, 2017, all of which are incorporated by reference herein.In other embodiments, the bispecific antigen binding protein comprises one moiety that binds CD3 and one moiety that binds BCMA, CD19, CD20, CD28, CLEC-12, Her2, an HLA protein, a MAGE protein, Mucl6, PSMA, or Steap-2. In yet other embodiments, the bispecific antigen binding protein is selected from the group consisting of an anti-CD3x anti-CD20 bispecific antibody (described in U.S. Patent Application Publications US2014 / 0088295A1 and US20150266966A1, which are incorporated herein by reference), an anti-CD3x anti-Mucin16 bispecific antibody (e.g., an anti-CD3x anti-Muc16 bispecific antibody), and an anti-CD3x anti-prostate specific membrane antigen bispecific antibody (e.g., an anti-CD3x anti-PSMA bispecific antibody).
[0120] In the context of a bispecific antigen-binding protein, where one portion binds to a T cell receptor, e.g., binds to CD3, and the other portion binds to the target antigen, the target antigen may be an infectious disease-associated antigen. Non-limiting examples of infectious disease-associated antigens include, for example, antigens expressed on the surface of a viral particle or preferentially expressed on a cell infected with a virus, where the virus is selected from the group consisting of HIV, hepatitis virus (A, B, or C), herpesvirus (e.g., HSV-1, HSV-2, CMV, HAV-6, VZV, or Epstein-Barr virus), adenovirus, influenza virus, flavivirus, echovirus, rhinovirus, coxsackievirus, coronavirus, respiratory syncytial virus, mumps virus, rotavirus, measles virus, rubella virus, parvovirus, vaccinia virus, HTLV, dengue virus, papillomavirus, molluscum contagiosum virus, poliovirus, rabies virus, JC virus, and arboviral encephalitis virus. Alternatively, the target antigen may be an antigen expressed on the surface of a bacterium, or may be an antigen preferentially expressed on a cell infected with a bacterium, wherein the bacterium is selected from the group consisting of Chlamydia, Rickettsia, Mycobacterium, Staphylococcus, Streptococcus, Pneumococcus, Neisseria meningitidis, Neisseria gonorrhoeae, Klebsiella, Proteus, Serratia, Pseudomonas, Legionella, Corynebacterium diphtheriae, Salmonella, Bacillus, Vibrio cholerae, Clostridium tetani, Clostridium botulinum, Bacillus anthrax, Yersinia pestis, Leptospira, and Lyme disease bacteria.In certain embodiments, the target antigen is an antigen expressed on the surface of a fungus or is an antigen preferentially expressed on cells infected with a fungus, wherein the fungus is selected from the group consisting of Candida (e.g., albicans, krusei, glabrata, tropicalis), Cryptococcus neoformans, Aspergillus (e.g., fumigatus), Mucorales (e.g., Mucor, absidia, rhizopus), Sporothrix schenkii, Blastomyces dermatitidis, Paracoccidioides brasiliensis, Coccidioides immitis, and the like. In certain embodiments, the target antigen is selected from the group consisting of Entamoeba histolytica, Balantidium coli, Naegleria fowleri, Acanthamoeba sp., Giardia lambia, Cryptosporidium sp., Pneumocystis carinii, Plasmodium vivax, and Babesia murin. The antigens are selected from the group consisting of Trypanosoma microti, Trypanosoma brucei, Trypanosoma cruzi, Leishmania donovani, Toxoplasma gondii, Nippostrongylus brasiliensis, Taenia crassiceps, and Brugia malayi. Non-limiting examples of specific parasite-associated antigens include, for example, HIV gp120, HIV CD4, Hepatitis B virus glycoprotein L, Hepatitis B virus glycoprotein M, Hepatitis B virus glycoprotein S, Hepatitis C virus E1, Hepatitis C virus E2, hepatocyte-specific protein, herpes simplex virus gB, cytomegalovirus gB, and HTLV envelope protein.
[0121] Bispecific binding proteins can be created with two binding moieties, each directed to a binding partner on the same cell surface (i.e., each directed to a different target). This design is particularly suitable for targeting specific cells or cell types that express both targets on the same cell surface. Although the targets may also appear individually on other cells, the binding moieties of these binding proteins are selected so that each binding moiety binds to its target with relatively low affinity (e.g., low micromolar or high nanomolar, e.g., greater than 100 nanomolar kD, e.g., 500, 600, 700, 800 nanomolar). In this case, long-term target binding is preferred only when the two targets are in close proximity on the same cell.
[0122] Bispecific binding proteins can be created with two binding moieties that bind to the same target at different epitopes on the same target. This design is particularly suitable for maximizing the success rate of blocking a target using the binding protein. For example, multiple extracellular loops of a transmembrane channel or cell surface receptor can be targeted by the same bispecific binding molecule.
[0123] Bispecific binding proteins may be engineered with two binding molecules that cluster and activate negative regulators of immune signaling, resulting in immune suppression. In cis suppression can be achieved when the targets are on the same cell. In trans suppression can be achieved when the targets are on different cells. In cis suppression can be achieved, for example, using a bispecific binding protein with an anti-IgGRIIb binding moiety and an anti-FelD1 binding moiety, whereby IgGRIIb clusters only in the presence of FelD1, downregulating the immune response to FelD1. In trans suppression can be achieved, for example, using a bispecific binding protein with an anti-BTLA binding moiety and a binding moiety that specifically binds to a tissue-specific antigen of interest, whereby clustering of inhibitory BTLA molecules occurs only in selected target tissues, potentially providing a strategy for addressing autoimmune diseases.
[0124] Bispecific binding proteins may be engineered to activate multicomponent receptors. In this design, two binding moieties directed to two components of the receptor bind to and cross-link the receptor, activating signaling from the receptor. This can be achieved by using a bispecific binding protein with a binding moiety that binds to IFNAR1 and a binding moiety that binds to IFNAR2, where binding cross-links the receptors. Such bispecific binding proteins may provide an alternative to interferon therapy.
[0125] Bispecific binding proteins may be engineered to transport binding moieties across semipermeable barriers, such as the blood-brain barrier. In this design, one binding moiety binds to a target that can cross a particular selected barrier, while the other binding moiety targets a therapeutically active molecule, where the therapeutically active target molecule typically cannot cross the barrier. Bispecific binding proteins of this type are useful for delivering therapeutic agents to tissues that would otherwise be inaccessible to the therapeutic agent. Some examples include targeting the pIGR receptor to transport therapeutic agents to the gastrointestinal tract or lungs, or targeting the transferrin receptor to transport therapeutic agents across the blood-brain barrier.
[0126] Bispecific binding proteins can be generated that deliver binding moieties to specific cells or cell types. In this design, one binding moiety targets a cell surface protein (e.g., a receptor) that is readily internalized within the cell. The other binding moiety targets an intracellular protein, the binding of which produces a therapeutic effect.
[0127] Bispecific binding proteins that bind to surface receptors on phagocytic immune cells and to surface molecules on infectious pathogens (e.g., yeast or bacteria) deliver the infectious pathogen to the vicinity of the phagocytic immune cell, facilitating phagocytosis of the pathogen. An example of such a design is a bispecific antibody that targets the CD64 or CD89 molecule as well as the pathogen.
[0128] A bispecific binding protein has an antibody variable region as one binding moiety and a non-Ig moiety as the second binding moiety. The antibody variable region performs targeting, while the non-Ig moiety is an effector or toxin linked to Fc. In this case, a ligand (e.g., an effector or toxin) is delivered to the target bound to the antibody variable region.
[0129] Bispecific binding proteins with two moieties, each bound to an Ig region (e.g., an Ig sequence containing a CH2 and a CH3 region), can bring any two protein moieties near each other in an Fc environment. Examples of this design include traps, such as homodimeric or heterodimeric trap molecules.
[0130] Enhanced expression loci Expression-enhanced loci suitable for use in the present invention include, for example, a locus with a nucleotide sequence having substantial homology to SEQ ID NO:1 described in U.S. Patent No. 8,389,239 (also referred to herein as the "EESYR® locus" or "Locus 1"), a locus with a nucleotide sequence having substantial homology to SEQ ID NO:2 or SEQ ID NO:3 described in U.S. Patent Application Serial No. 14 / 919,300 (also referred to herein as the "YARS locus" or "Locus 2"), and other expression-enhanced loci and sequences reported in the art (e.g., US20150167020A1 and U.S. Patent No. 6,800,457).
[0131] In some embodiments, the two expression-enhancing loci used in the invention are selected from the group consisting of a locus comprising a nucleotide sequence having substantial homology to SEQ ID NO: 1, a locus comprising a nucleotide sequence having substantial homology to SEQ ID NO: 2, and a locus comprising a nucleotide sequence having substantial homology to SEQ ID NO: 3. These loci contain sequences that not only confer enhanced expression of the gene operably linked to and integrated into (i.e., within or adjacent to) the sequence, but also exhibit increased recombination efficiency and improved integration stability compared to other sequences in the genome.
[0132] SEQ ID NO:1, SEQ ID NO:2, and SEQ ID NO:3 were identified in CHO cells. Other mammalian species (e.g., humans or mice) have been found to have limited homology to the identified expression-enhancing regions. However, sequences may be found in cell lines derived from other tissue types of Chinese hamsters (Cricetulus griseus) or other isogenic species and can be isolated using techniques known in the art. For example, cross-species hybridization or PCR-based techniques may identify other homologous sequences. Furthermore, variations can be generated in the nucleotide sequences set forth in SEQ ID NO:1, SEQ ID NO:2, or SEQ ID NO:3 using site-directed or random mutagenesis techniques known in the art. The resulting sequence variants may then be tested for expression-enhancing activity. DNAs at least about 90% identical in nucleic acid identity to SEQ ID NO:1, SEQ ID NO:2, or SEQ ID NO:3 that have expression-enhancing activity are expected to be isolated by routine experimentation and exhibit expression-enhancing activity.
[0133] The integration site, site, or nucleotide position of the insert of one or more exogenous nucleic acids may be any location within or adjacent to any of the expression-enhancing sequences (e.g., SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3). Whether a particular chromosomal location within or adjacent to a locus of interest supports stable integration and efficient transcription of an integrated exogenous gene can be determined according to standard methods known in the art, such as those described in U.S. Pat. No. 8,389,239 and U.S. Patent Application Publication No. 14,919,300.
[0134] The integration sites contemplated herein are located within or within the vicinity of the expression-enhancing sequence, e.g., less than about 1 kb, 500 base pairs (bp), 250 bp, 100 bp, 50 bp, 25 bp, 10 bp, or less than about 5 bp upstream (5') or downstream (3') of the location of the expression-enhancing sequence on the chromosomal DNA. In yet some other embodiments, the integration sites used are located about 1000 base pairs, 2500 base pairs, 5000 base pairs, or more upstream (5') or downstream (3') of the location of the expression-enhancing sequence on the chromosomal DNA.
[0135] It is understood in the art that large genomic regions, such as scaffold / matrix attachment regions (S / MARs), also known as scaffold-attachment regions (SARs) or matrix-associated or matrix attachment regions (MARs), are regions of genomic DNA in eukaryotic cells to which the nuclear matrix attaches. Without being bound by any one theory, S / MARs often map to non-coding regions, separating a given transcriptional region (e.g., chromatin domain) from its neighbors and providing platforms for the structure and / or binding of factors that enable transcription, such as recognition sites for deoxyribonucleases or polymerases. Some S / MARs have been characterized at lengths of approximately 14-20 kb (Klar, et al. 2005, Gene 364:79-89). Therefore, integration of a gene at an enhanced expression locus (e.g., within or near SEQ ID NO:1 or SEQ ID NO:2 or SEQ ID NO:3) is predicted to result in enhanced expression. In some embodiments, host cells with an exogenous nucleic acid sequence encoding a bispecific antigen-binding protein integrated at a specific site within an enhanced expression locus exhibit high specific productivity. In other embodiments, host cells encoding the bispecific antigen-binding protein have a specific productivity of at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, or 30 picograms / cell / day (pcd).
[0136] In some embodiments, the exogenous nucleic acid is integrated at a site within a locus having the nucleotide sequence of SEQ ID NO: 1. In certain embodiments, the integration site is within or near the nucleotide sequence of SEQ ID NO: 1. In certain embodiments, the integration site is at a position within SEQ ID NO: 1 selected from nucleotides spanning positions numbered 10-13,515; 20-12,020; 1,020-11,020; 2,020-10,020; 3,020-9,020; 4,020-8,020; 5,020-7,020; 6,020-6,920; 6,120-6,820; 6,220-6,720; 6,320-6,620; 6,420-6,520; 6,460-6,500; 6,470-6,490; and 6,475-6,485. In other embodiments, the integration site is in a sequence selected from the group consisting of nucleotides 5,000-7,400, 5,000-6,500, 6,400-7,400 of SEQ ID NO: 1, and nucleotides 6,400-6,500 of SEQ ID NO: 1. In certain embodiments, the integration site is before, after, or within the "turn on" triplet nucleotides 6471-6473 of SEQ ID NO: 1.
[0137] In some embodiments, the exogenous nucleic acid is integrated at a site within a locus comprising the nucleotide sequence of SEQ ID NO:2 or SEQ ID NO:3. In particular embodiments, the integration site is within or near the nucleotide sequence of SEQ ID NO:2. In specific embodiments, the integration site is within or near the nucleotide sequence of SEQ ID NO:3. In some embodiments, the integration site is within or near nucleotides 1990-1991, 1991-1992, 1992-1993, 1993-1994, 1995-1996, 1996-1997, 1997-1998, 1999-2000, 2001-2002, 2002-2003, 2003-2004, 2004-2005, 2005-2006, 2006-2008, 2009-3010, 2010-2011, 2011-2012, 2012-2013, 2013-2014, 2014-2015, 2015-2016, 2016-2017, 2017-2018, 2018-2019, 2019-3020, 2019-3020, 2020-2021, 2020-2022, 2020-2023, 2020-2024, 2020-2025 ...6, 2020 In certain embodiments, the integration is at or within nucleotides 2001-2022 of SEQ ID NO:3. In some embodiments, the exogenous nucleic acid is inserted at or within nucleotides 2001-2002 or nucleotides 2021-2022 of SEQ ID NO:3, and as a result of the insertion, nucleotides 2002-2021 of SEQ ID NO:3 are deleted.
[0138] Site-specific integration into expression-enhancing loci Integration of one or more exogenous nucleic acids into an expression-enhancing locus in a site-specific manner, i.e., into one particular site within an expression-enhancing locus disclosed herein, can be achieved in several ways, including homologous recombination and recombinase-mediated cassette exchange, as described in the art (see, e.g., U.S. Pat. No. 8,389,239 and the techniques disclosed therein).
[0139] In some embodiments, cells are provided that contain at least two, i.e., two or more different recombinase recognition sequences (RRSs) within an expression-enhancing locus convenient for integration of a nucleic acid containing one or more exogenous nucleic acids or genes of interest. Such cells can be obtained by introducing an exogenous nucleic acid sequence containing two or more RRSs into the desired locus using a variety of means, including homologous recombination, as described herein below and in the art, e.g., U.S. Patent No. 8,389,239 and the techniques disclosed therein.
[0140] In certain embodiments, cells are provided that contain three or more different recombinase recognition sequences (RRS) within an expression-enhancing locus that is convenient for the integration of multiple exogenous nucleic acids. In certain embodiments, cells are provided that contain three different recombinase recognition sequences (RRS) within an expression-enhancing locus that can mediate the integration of two separate exogenous nucleic acids, for example, where the 5' and middle RRS in the genome match the 5' and 3' RRS flanking a first exogenous nucleic acid to be integrated, and the middle and 3' RRS in the genome match the 5' and 3' RRS flanking a second exogenous nucleic acid to be integrated.
[0141] A suitable RRS may be selected from the group containing LoxP, Lox511, Lox5171, Lox2272, Lox2372, Loxm2, Lox-FAS, Lox71, Lox66, and mutants thereof, in which case the site-specific recombinase is Cre recombinase or a derivative thereof, and recombinase-mediated cassette exchange (RMCE) is performed using it. In another example, a suitable RRS may be selected from the group containing FRT, F3, F5, FRT mutant-10, FRT mutant +10, and mutants thereof, in which case the site-specific recombinase Flp recombinase or a derivative thereof is used, and RMCE is performed using it. In yet another example, the RRS may be selected from the group containing attB, attP, and mutants thereof, in which case the site-specific recombinase is phiC31 integrase or a derivative thereof, and RMCE is performed using it.
[0142] In other embodiments, native cells are modified by homologous recombination techniques, whereby a nucleic acid sequence containing one or more exogenous nucleic acids is integrated into a specific site within the expression-enhancing locus.
[0143] With homologous recombination, homologous polynucleotide molecules (i.e., homologous arms) align and exchange their sequence sections. During this exchange, a transgene can be introduced if it is flanked by homologous genomic sequences. In one example, a recombinase recognition site can be introduced into the host cell genome at the integration site via homologous recombination. In another example, a nucleic acid sequence containing one or more exogenous nucleic acids of interest, e.g., one or more nucleic acids each encoding an HCF or LCF (e.g., variable region), where the nucleic acid sequence is flanked by sequences homologous to the sequence of the target locus (homologous arms), is inserted into the host genome.
[0144] Homologous recombination in eukaryotic cells can be facilitated by introducing a break at the integration site in the chromosomal DNA. This can be accomplished by targeting a specific nucleus to a specific integration site. DNA-binding proteins that recognize DNA sequences at the target locus are known in the art. Gene targeting vectors can also be used to facilitate homologous recombination.
[0145] Gene targeting vector construction and nuclease selection for homologous recombination are within the skill of those skilled in the art to which the present invention pertains. In some instances, zinc finger nucleases (ZFNs), which have a modular structure and contain individual zinc finger domains, identify specific 3-nucleotide sequences within a target sequence (e.g., the site of targeted integration). Some embodiments can utilize ZFNs with a combination of individual zinc finger domains to target multiple target sequences. Transcription activator-like (TAL) effector nucleases (TALENs) can also be employed for site-specific genome editing. TAL effector protein DNA binding domains are typically used in combination with the non-specific cleavage domain of a restriction nuclease such as FokI. In some embodiments, a fusion protein containing a TAL effector protein DNA binding domain and a restriction nuclease cleavage domain is employed to identify and cleave DNA at target sequences within the locus of the present invention (Boch J et al., 2009 Science 326:1509-1512). RNA-guided endonucleases (RGENs) are programmable genome engineering tools developed from bacterial adaptive immune mechanisms. In this system, i.e., clustered regularly interspaced short palindromic repeats (CRISPR) / CRISPR-associated (Cas) immune response, the protein Cas9 forms a sequence-specific endonuclease when complexed with two RNAs, one of which guides target selection. RGENs consist of components (Cas9 and tracrRNA) and target-specific CRISPRRNA (crRNA). Both the efficiency of DNA target cleavage and the location of the cleavage site vary based on the location of the protospacer adjacent motif (PAM), an additional requirement for target recognition (Chen, H. et al., J. Biol. Chem. Published online March 14, 2014, manuscript M113.539726).Sequences unique to specific targeting of loci (e.g., SEQ ID NO:1, SEQ ID NO:2, or SEQ ID NO:3) can be identified by aligning many of these sequences to the CHO genome, revealing potential off-target sites with 16-17 base pair matches.
[0146] In some embodiments, a targeting vector carrying a nucleic acid of interest (e.g., a nucleic acid containing one or more RRSs, optionally adjacent to one or more selectable marker genes, or a nucleic acid containing one or more exogenous nucleic acids encoding HCF or LCF (e.g., variable regions), each flanked by 5' and 3' homology arms) is introduced into a cell along with one or more additional vectors or mRNAs. In one embodiment, the one or more additional vectors or mRNAs contain nucleotide sequences encoding site-specific nucleases, including, but not limited to, zinc finger nucleases (ZFNs), ZFN dimers, transcription activator-like effector nucleases (TALENs), TAL effector domain fusion proteins, and RNA-guided DNA endonucleases. In certain embodiments, the one or more vectors or mRNAs include a first vector comprising nucleotide sequences encoding a guide RNA, a tracrRNA, and a Cas enzyme, and a second vector comprising a donor (exogenous) nucleotide sequence. Such donor sequences contain nucleotide sequences encoding a gene of interest, or a recognition sequence, or a gene cassette with any one of these exogenous elements intended for targeted insertion. When mRNA is used, the mRNA can be transfected into cells using common transfection methods known to those skilled in the art and may encode enzymes such as transposases or endonucleases. The mRNA introduced into cells may be transient and not integrated into the genome, but the mRNA may carry exogenous nucleic acids necessary or beneficial for integration to occur. In some cases, mRNA is selected to eliminate any risk of long-lasting side effects of the accompanying polynucleotide, in which case only short-term expression is required to achieve the desired integration of the nucleic acid.
[0147] Vectors for site-specific integration Provided herein are nucleic acid vectors for introducing exogenous nucleic acids into two expression-enhancing loci via site-specific integration. Suitable vectors include vectors designed to contain an exogenous nucleic acid sequence flanked by RRSs for integration via RMCE, and vectors designed to contain an exogenous nucleic acid sequence of interest flanked by homology arms for integration via homologous recombination.
[0148] In various embodiments, a vector is provided for carrying out site-specific integration via RMCE. In some embodiments, the vector is designed to carry out simultaneous integration of multiple nucleic acids into two target loci. In contrast to sequential integration, simultaneous integration allows for the efficient and rapid isolation of desired clones that produce antigen-binding proteins or other target multimeric proteins suitable for large-scale production (manufacturing).
[0149] In some embodiments, a vector set for expression of a bispecific antigen binding protein in a cell is provided.
[0150] In some embodiments, the vector set may contain two "HCF vectors," each vector containing nucleic acid flanked by 5' and 3' RRSs, which nucleic acid contains a nucleotide sequence encoding an HCF, and in this case the HCFs are different. The RRSs on the two HCF vectors are different from each other and are designed to integrate the HCF-encoding nucleotide sequence into two expression-enhancing loci. The vector set also contains a nucleotide sequence encoding an LCF, which may be included in one or both of the HCF vectors (thereby providing two copies of the same LCF), or alternatively, may be provided in a separate "LCF vector" and flanked by 5' and 3' RRSs.
[0151] In some embodiments, the nucleotide sequence encoding LCF is contained in one of the HCF vectors and is located between the 5' and 3' RRS on the HCF vector. The sequence encoding LCF may be located upstream or downstream of the sequence encoding HCF.
[0152] In some embodiments, the nucleotide sequence encoding LCF is contained in both HCF vectors and is located between the 5' RRS and 3' RRS on each HCF vector. Similarly, the sequence encoding LCF may be located upstream or downstream of the sequence encoding HCF in each vector.
[0153] In some embodiments, the nucleotide sequence encoding LCF is provided in a separate vector, and the "LCF" vector is flanked by a 5' RRS and a 3' RRS, where the two RRSs are different from each other. The RRSs in the vector set may be designed so that the sequence encoding LCF can be "linked" to one of the HCF-encoding sequences via a common RRS during RMCE with a target locus, where the target locus also contains a common RRS. For example, the 3' RRS of the LCF vector may be identical to the 5' RRS of one of the HCF vectors, thereby generating an LCF-HCF sequence after integration at the target locus via RMCE. In another example, the 3' RRS of the HCF vector may be identical to the 5' RRS of the LCF vector, thereby generating an HCF-LCF sequence after integration at the target locus via RMCE. In some embodiments, the common RRS is designed in a split selectable marker format. That is, a common RRS is contained at the 3' end of the 5' portion of a selectable marker gene contained in one vector and also at the 5' end of the remaining 3' portion of the same selectable marker gene contained in another vector, such that upon "ligation" and integration into the target locus, the properly integrated nucleic acid contains the entire gene for the selectable marker, allowing for convenient identification of transfectants. In some embodiments, the common RRS is designed in a split-gene format. That is, the common RRS is contained as part of the 5' portion of a gene or intron within such split gene on one vector, at the 3' end of the 5' portion of that gene, and at the 5' end of the remaining portion of the gene or intron within such split gene, as part of the remaining 3' portion of the split gene. In still other embodiments, a third or middle RRS in the first vector is designed to be between the promoter and selectable marker gene to which it is operably linked (but separated on the other vector). The third or middle RRS in the first vector is designed to be 3' of the promoter, and the third or middle RRS in the second vector is designed to be 5' of the selectable marker gene.
[0154] In some embodiments, the vector set may include an additional nucleotide sequence encoding an LCF. That is, the vector set may include two HCF vectors and two LCF-encoding nucleotide sequences. The two LCF-encoding sequences may encode the same or different LCFs. In some embodiments, the two LCF-encoding sequences may each be contained within an HCF vector, resulting in two vectors, each containing an HCF-encoding sequence and an LCF-encoding sequence. The two vectors may be designed with RRSs suitable for targeting the two vector sequences into two loci. In other embodiments, one of the two LCF-encoding sequences is contained within the HCF vector and located between the 5' and 3' RRSs on the HCF vector, and the other LCF-encoding sequence is provided on a separate vector. That is, one vector containing both LCF and HC (in the sequence LCF-HCF or HCF-LCF, or in short "LCF / HCF vector"), one HCF vector, and one LCF vector. In some of these other embodiments, the RRS of the vectors may be designed to allow the HCF coding sequence on the HCF vector and the LCF coding sequence on the LCF vector to join at the target locus via RMCE. For example, the 3' RRS of the LCF vector may be identical to the 5' RRS of the HCF vector, and the common RRS may be designed in a split selectable marker format or a split intron format, thereby facilitating selection and identification of transfectants. In still other embodiments, if the two LCFs are different, the nucleotide sequences encoding the two LCFs may each be provided on a separate vector. That is, the vector set contains two HCF vectors and two LCF vectors. The RRS is designed to allow the appropriate "joining" of one LCF coding sequence with one HCF coding sequence at one target locus and the other LCF coding sequence with the other HCF coding sequence at a second target locus. Figures 1, 3, and 4 are illustrative of different vector formats and RRS / locus combinations, but are not meant to be limiting.Each given vector system provides the means for simultaneous integration of each nucleotide sequence in the presence of a recombinase for rapid and convenient selection of positive integrants (desired clones).
[0155] Nucleotide sequences encoding HCF or LCF may encode amino acids or domains from the constant region, or may encode the entire constant region. In certain embodiments, nucleotide sequences encoding HCF or LCF may encode one or more constant domains, such as CL, CH1, hinge, CH2, CH3, or a combination thereof. In some embodiments, nucleotide sequences encoding HCF domains may encode a CH3 domain. For example, a nucleotide sequence encoding a first HCF may encode a first CH3 domain, and a nucleotide sequence encoding a second HCF may encode a second CH3 domain. The first and second CH3 domains may be identical or may differ by at least one amino acid. Differences in the CH3 domain or in the constant region can take any of the forms for the bispecific antigen-binding proteins described herein, such as differences that result in different protein A binding properties or differences that result in a "knob-and-hole" format. Regardless of any amino acid sequence differences, two HCF-encoding nucleotide sequences may differ in that one of the two nucleotide sequences has been codon-modified.
[0156] In some embodiments, the nucleotide sequence encoding each HCF or LCF is independently operably linked to a transcription control sequence, including, for example, a promoter. In some embodiments, the promoter directing transcription of the two HCF-containing polypeptides is the same. In some embodiments, the promoter directing transcription of the two HCF-containing polypeptides and the promoter directing transcription of the LCF-containing polypeptide are all the same (e.g., a CMV promoter or any other suitable promoter described herein). In some embodiments, the nucleotide sequence encoding each HCF or LCF is independently operably linked to an inducible or repressible promoter. Inducible and repressible promoters allow, for example, production to occur only during the production phase (fed-batch culture) and not during the growth phase (seed train culture). Inducible or repressible promoters also allow specific expression of one or more genes of interest. In some embodiments, the nucleotide sequence encoding each HCF and / or LCF is independently operably linked to a promoter upstream of at least one TetR operator (TetO) or Arc operator (ArcO). In yet other embodiments, the nucleotide sequences encoding each HCF and / or LCF are independently operably linked to a CMV / TetO or CMV / ArcO hybrid promoter. Examples of hybrid promoters (also referred to as regulatory fusion proteins) are found in International Patent Application Publication No. WO03101189A1, published December 11, 2003, which is incorporated herein by reference.
[0157] In some embodiments, the vector set includes nucleotide sequences encoding recombinases that recognize one or more RRSs, which may be included in one of the HCF-encoding vectors or the LCF-encoding vectors, or may be provided in separate vectors.
[0158] In various other embodiments, vectors are provided for performing site-specific integration via homologous recombination.
[0159] In some embodiments, a vector set is provided that includes two vectors, each vector containing an exogenous nucleic acid flanked by 5' and 3' homology arms for site-specific integration into two expression-enhancing loci in a cell, wherein the exogenous nucleic acids of the two vectors together encode antigen-binding proteins. Thus, the homology arms on one vector are designed for integration into one of the two loci, and the homology arms on the other vector are designed for integration into the other locus. In these embodiments, the antigen-binding proteins may be monospecific or bispecific.
[0160] It is within the skill of one in the art to select a sequence homologous to a sequence within an expression-enhancing locus and include the selected sequence as a homology arm in a targeting vector. In some embodiments, the vector or construct comprises a first homology arm and a second homology arm, wherein the combined first and second homology arms comprise a target sequence that replaces an endogenous sequence within the locus. In other embodiments, the first and second homology arms comprise a target sequence that is to be integrated or inserted into an endogenous sequence within the locus. In some embodiments, the homology arms contain nucleotide sequences homologous to nucleotide sequences present in SEQ ID NO:1, SEQ ID NO:2, or SEQ ID NO:3. In certain embodiments, the vector contains a 5' homology arm having a nucleotide sequence corresponding to nucleotides 1001-2001 of SEQ ID NO:3 and a 3' homology arm having nucleotides homologous to nucleotides 2022-3022 of SEQ ID NO:3. The homologous arms, e.g., a first homologous arm (also referred to as a 5' homologous arm) and a second homologous arm (also referred to as a 3' homologous arm), are homologous to a target sequence within the locus. In the 5' to 3' direction, the homologous arms may extend a region or target sequence within the locus comprising at least 1 kb, or at least 2 kb, or at least 3 kb, or at least 4 kb, or at least 5 kb, or at least 10 kb. In other embodiments, the total number of nucleotides of the target sequence selected for the first and second homologous arms contains at least 1 kb, or at least 2 kb, or at least 3 kb, or at least 4 kb, or at least 5 kb, or at least 10 kb. In some examples, the distance between the 5' homology arm and the 3' homology arm (homologous to the target sequence) contains at least 5 bp, 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, or at least 1 kb, or at least about 2 kb, or at least about 3 kb, or at least about 4 kb, or at least 5 kb, or at least about 10 kb.For example, when nucleotides 1001-2001 and 2022-3022 of SEQ ID NO: 3 are selected for the 5' and 3' homologous arms, the distance between the two homologous arms may be 20 nucleotides (corresponding to nucleotides 2002-2021 of SEQ ID NO: 3). Such homologous arms can mediate integration of an exogenous nucleic acid sequence into a locus containing SEQ ID NO: 3, e.g., within nucleotides 1990-2021 or 2002-2021 of SEQ ID NO: 3, and the simultaneous deletion of nucleotides 2002-2021 of SEQ ID NO: 3.
[0161] The vectors disclosed herein for introducing an exogenous nucleic acid for site-specific integration into an expression-enhancing locus may contain additional genes and sequences for directing expression of the exogenous nucleic acid of interest and the encoded polypeptide, and for selection and identification of cells into which the exogenous nucleic acid of interest has been successfully integrated. Such additional sequences include, for example, transcriptional and translational regulatory sequences, selectable marker genes, and the like, and are also described herein below.
[0162] Control arrays The vectors disclosed herein for introducing exogenous nucleic acids into expression-enhancing loci in a site-specific manner, and the cells resulting from site-specific integration, may contain control sequences for directing expression of the exogenous nucleic acid of interest and the encoded polypeptide. Control sequences include transcriptional promoters, enhancers, sequences encoding appropriate mRNA ribosomal binding sites, and sequences controlling transcription and translation termination. Transcriptional and translational control sequences may be provided by viral sources. For example, commonly used promoters and enhancers are derived from, for example, polyoma, adenovirus 2, simian virus 40 (SV40), murine or human cytomegalovirus (CMV), the CMV immediate-early (CMV-IE) or CMV major IE (CMV-MIE) promoters, as well as RSV, the SV40 late promoter, SL3-3, MMTV, ubiquitin (Ubi), ubiquitin C (UbC), and the HIV LTR promoter. Viral genomic promoters, control sequences, and / or signal sequences may be utilized to drive expression, provided such control sequences are compatible with the selected host cell. Depending on the cell type in which the protein of interest is to be expressed, non-viral cellular promoters (e.g., β-globin and EF-1α promoters) can also be used. DNA sequences derived from the SV40 viral genome, such as the early and late promoters, enhancers, splice, and polyadenylation sites of SV40 origin, may be used to provide other genetic elements useful for expressing exogenous DNA sequences. The early and late promoters are particularly useful because both promoters are readily obtained as fragments from the SV40 virus and also contain the SV40 viral origin of replication (Fiers et al., Nature 273:113, 1978). Smaller or larger SV40 fragments may also be used. Typically, an approximately 250-bp sequence extending from the Hind III site toward the Bgl I site located in the SV40 origin of replication is included.Inducible (e.g., induced by a compound, cofactor, or regulatory protein) and / or repressible (e.g., repressible by a compound, cofactor, or regulatory protein) promoters can be used and are particularly useful because they allow production of the antigen binding protein to occur only during the production phase (fed-batch culture) and not during the growth phase (seed train culture), or allow for precise control of the expression of antibody components at different loci specifically. Examples of inducible promoters include the alcohol dehydrogenase I gene promoter, tetracycline-responsive promoter systems, glucocorticoid receptor promoters, estrogen receptor promoters, ecdysone receptor promoters, metallothionein-based promoters, and T7 polymerase-based promoters. Examples of repressible promoters include hybrid promoters (also referred to as regulatory fusion proteins) containing a CMV promoter or other promoter operably linked to at least one TetR operator (TetO) or Arc operator (ArcO), as described in International Patent Application Publication No. WO03101189A1, published December 11, 2003 (incorporated herein by reference). Sequences suitable for expressing multiple transcripts via bicistronic vectors have previously been reported (Kim SK and Wold BJ, Cell 42:129, 1985) and can be used in the present invention. Examples of suitable strategies for multicistronic expression of proteins include the use of 2A peptides (Szymczak et al., Expert Opin Biol Ther 5: 627-638 (2005)) and internal ribosome entry sites ("IRES"), both of which are known in the art. Other types of expression vectors are also useful, such as those described in US Pat. No. 4,634,665 (Axel et al.) and US Pat. No. 4,656,134 (Ringold et al.). Selection Marker
[0163] The vectors disclosed herein for introducing exogenous nucleic acids into expression-enhancing loci in a site-specific manner, and the cells resulting from site-specific integration, may contain one or more selectable marker genes.
[0164] In some embodiments, the selectable marker gene confers drug resistance, such as those described in Kaufman, RJ (1988) Meth. Enzymology 185:537, including DHFR-MTX resistance, P-glycoprotein and multidrug resistance (MDR)-various lipophilic cytotoxic drugs (e.g., adriamycin, colchicine, vincristine), and adenosine deaminase (ADA)-Xyl-A, or adenosine and 2'-deoxycoformycin resistance. Other primary selectable markers include microbial antibiotic resistance genes, such as neomycin resistance, kanamycin resistance, or hygromycin resistance. Several suitable selection systems exist for mammalian hosts (Sambrook, supra, pp. 16.9-16.15). Protocols for cotransfection using two primary selectable markers have also been reported (Okayama and Berg, Mol. Cell Biol. 5:1136, 1985).
[0165] In other embodiments, the selectable marker gene encodes a polypeptide that provides a detectable signal for recognition of successful or unsuccessful insertion and / or replacement of the gene cassette, or that can generate a detectable signal. Suitable examples include, inter alia, fluorescent markers or proteins, and enzymes that catalyze chemical reactions that generate detectable signals. Examples of fluorescent markers are known in the art and include, but are not limited to, Discosoma coral (DsRed), green fluorescent protein (GFP), enhanced green fluorescent protein (eGFP), cyan fluorescent protein (CFP), enhanced cyan fluorescent protein (eCFP), yellow fluorescent protein (YFP), enhanced yellow fluorescent protein (eYFP), and near-infrared fluorescent proteins (e.g., mKate, mKate2, mPlum, mRaspberry, or E2-Crimson). See, for example, Nagai, T., et al. 2002 Nature Biotechnology 20:87-90; Heim, R. et al. 1995 February 23 Nature 373:663-664; and Strack, R. et al. 2009 Biochemistry 48:8279-81.
[0166] System for producing antigen-binding proteins In a further aspect, the present disclosure provides a system that can be used to generate cells comprising a combination of one or more vectors and cells (e.g., CHO cells) and having exogenous nucleic acids integrated into two expression-enhancing loci, where the exogenous nucleic acids together encode an antigen-binding protein, either a monospecific or bispecific protein. The system may be provided, for example, in the form of a kit.
[0167] In some embodiments, the system is designed to allow efficient vector construction and simultaneous integration of multiple exogenous nucleic acids via RMCE into specific sites within two expression-enhancing loci. Simultaneous integration allows for rapid isolation of desired clones, and the use of two expression-enhancing loci is also important for the generation of stable cell lines suitable for protein production (e.g., commercially available cell lines).
[0168] The system provided herein includes a cell and a vector set. The cell contains a pair of RRSs (5' RRS and 3' RRS) integrated into each of two expression-enhancing loci. In some embodiments, an exogenous nucleic acid is present between the 5' RRS and 3' RRS at each locus and may contain, for example, one or more selectable marker genes. The vector set includes at least two vectors, each of which contains a pair of RRSs (5' RRS and 3' RRS) flanking a nucleotide sequence encoding HCF or LCF, wherein the nucleotide sequence of one of the two vectors encodes HCF (HCF vector) and the nucleotide sequence of the other of the two vectors encodes LCF (LCF vector), where HCF and LCF are regions of antigen-binding proteins. The 5'RRS and 3'RRS in each pair of RRSs are different, and the RRSs of the system are designed so that when the vector is introduced into a cell, the HCF-encoding nucleotide sequence or LCF-encoding nucleotide sequence in the vector is integrated into the two expression-enhancing loci by RMCE mediated by the RRSs, resulting in expression of the antigen-binding protein. The number of vectors, the location of the HCF- or LCF-encoding sequence, and the relationship between the RRSs can be designed separately depending on the antigen-binding protein.
[0169] In some embodiments, the system is designed for integration into two expression-enhancing loci and expression of a monospecific antigen-binding protein. In some embodiments, the 5'RRS and 3'RRS of one of the two vectors (i.e., the HCF vector and the LCF vector) are identical to the 5'RRS and 3'RRS of one of the two loci, respectively, and the 5'RRS and 3'RRS of the other vector are identical to the 5'RRS and 3'RRS of the other locus, essentially targeting the HCF nucleic acid and the LCF nucleic acid separately, one at each of the two loci. In other embodiments, the HCF coding sequence and the LCF coding sequence are designed to be integrated together into each of the two loci while on separate vectors. According to these embodiments, the 5'RRS and 3'RRS of the first locus are identical to the 5'RRS and 3'RRS of the second locus, respectively. Each locus further contains an additional RRS (intermediate RRS) between the 5'RRS and 3'RRS. Furthermore, the 5' RRS in the first two vectors is identical to the 5' RRS in the first and second loci, the 3' RRS in the first vector is identical to the 5' RRS in the second vector and the middle RRS in the first and second loci, and the 3' RRS in the second vector is identical to the 3' RRS in both loci. Vectors may be designed with a split promoter and selectable marker format (a promoter on one vector and a selectable marker operably linked to the promoter on another vector). Vectors may also be designed with a split selectable marker format or a split intron format to facilitate selection of transfectants with proper integration. Furthermore, systems may be designed such that the relative positions of the LCF and HCF coding sequences differ after integration. In some embodiments, systems are designed with the LCF coding sequence integrated upstream of the HCF coding sequence. In other embodiments, systems are designed with the LCF coding sequence integrated upstream of the HCF coding sequence.
[0170] In some embodiments, the system is designed for integration into two expression-enhancing loci and expression of bispecific antigen binding proteins.
[0171] In some embodiments, in addition to the HCF vector (encoding the first HCF) and the LCF vector (encoding the first LCF), the system further includes a nucleotide sequence encoding a second HCF that is different from the first HCF. The nucleotide sequence encoding the second HCF may be contained, for example, in the LCF vector or in a separate vector, i.e., the second HCF vector. In some embodiments, the sequence encoding the second HCF is contained in the LCF vector between the 5' RRS and 3' RRS on the LCF vector, in which case the system contains an HCF vector and an LCF / HCF vector. The system, particularly the RRS, may be designed to integrate the HCF coding sequence into one of two loci and integrate sequences encoding both the HCF and LCF into the other locus. In other embodiments, the nucleotide sequence encoding the second HCF is on a separate vector and is flanked by the 5' RRS and 3' RRS, in which case the system contains two HCF vectors and one LCF vector. In these other embodiments, the RRSs of the system may be designed so that an LCF coding sequence can be "linked" to one of the HCF coding sequences via RMCE by a common RRS, where the common RRS between the 5' RRS and 3' RRS of the loci is also present at one of the two loci, and the other HCF coding sequence is integrated into the other of the two loci. For example, the 3' RRS of the LCF vector may be identical to the 5' RRS of one HCF vector and also to the middle RRS of one of the two loci. This design results in an LCF-HCF sequence after integration into a locus with a middle RRS. In another example, the 3' RRS of the HCF vector may be identical to the 5' RRS of the LCF vector and to the middle RRS of one of the two loci, thereby resulting in an HCF-LCF sequence after integration at a locus with a middle RRS. In some embodiments, the consensus RRS is designed in a split selectable marker format, or a split intron format, as described herein above.
[0172] In some embodiments, in addition to the HCF vector (encoding a first HCF) and the LCF vector (encoding the first LCF), and a nucleotide sequence encoding a second HCF different from the first HCF, the system further includes a nucleotide sequence encoding a second LCF. That is, the system includes four separate coding sequences: two encoding HCF and two encoding LCF. The two LCFs may be the same or different. The four coding sequences may be arranged in the vectors in different designs. In some embodiments, the four sequences are arranged in two vectors: LCF / HCF and LCF / HCF. The LCF is either upstream or downstream of the HCF in either vector. The system (RRS) may be designed so that one vector sequence is integrated into one locus and the other vector sequence is integrated into the other locus. In some embodiments, the four sequences are arranged in three vectors: LCF, HCF, and LCF / HCF (LCF is either upstream or downstream of the HCF). The RRS of the system may be designed so that sequences in the LCF / HCF vectors are integrated into one locus, and the LCF coding sequence in the LCF vector and the HCF coding sequence in the HCF vector are integrated into the other locus, by utilizing a common RRS shared by the LCF vector, the HCF vector, and the other locus. Similarly, the common RRS may be designed in a split marker format or a split intron format. In some embodiments, the four sequences are arranged in the following four vectors: LCF, HCF, LCF, and HCF. The RRS of the system may be designed so that the LCF coding sequence in one LCF vector and the HCF coding sequence in one HCF vector are integrated together into one locus by using the common RRS, and the LCF coding sequence in the other LCF vector and the HCF coding sequence in the other HCF vector are integrated together into one locus by using the common RRS.
[0173] In various embodiments of the systems provided herein, the nucleotide sequence encoding an HCF or LCF may encode amino acids from a constant region, e.g., an amino acid or domain, or may encode the entire constant region. In certain embodiments, the nucleotide sequence encoding an HCF or LCF may encode one or more constant domains, e.g., CL, CH1, CH2, CH3, or a combination thereof. In some embodiments, the nucleotide sequence encoding an HCF domain may encode a CH3 domain. For example, a nucleotide sequence encoding a first HCF may encode a first CH3 domain, and a nucleotide sequence encoding a second HCF may encode a second CH3 domain. The first and second CH3 domains may be identical or may differ by at least one amino acid. Differences in the CH3 domain or in the constant region can take any of the forms for the bispecific antigen-binding proteins described herein, such as differences that result in different protein A binding properties or differences that result in a "knob-and-hole" format. Regardless of any amino acid sequence differences, two HCF-encoding nucleotide sequences may differ in that one of the two nucleotide sequences has a codon modification.
[0174] In various embodiments of the systems provided herein, each HCF- or LCF-encoding nucleotide sequence is independently operably linked to a transcriptional control sequence, including, for example, a promoter. In some embodiments, the promoters directing transcription of the two HCF-containing polypeptides are the same. In some embodiments, the promoters directing transcription of the two HCF-containing polypeptides and the promoter directing transcription of the LCF-containing polypeptide are all the same (e.g., a CMV promoter, an inducible promoter, a repressible promoter, or any other suitable promoter described herein).
[0175] In some embodiments, the systems of the invention further comprise nucleotide sequences encoding recombinases that recognize one or more RRSs, which may be included in one of the vectors encoding the variable regions or may be provided in separate vectors.
[0176] The systems disclosed herein are designed to enable efficient vector construction and rapid isolation of desired clones, and the use of two expression-enhancing loci is also important for generating stable cell lines. In some embodiments, the systems are designed to use negative selection to identify transformants with the intended site-specific integration (e.g., lack of fluorescence resulting from one or more fluorescent marker genes in the host genome that are removed after RMCE). A single round of negative selection may require only two weeks, but the efficiency of isolating clones with the intended recombination is limited (approximately 1%). Combining negative and positive selection based on a new selectable marker provided by the integrated nucleic acid (e.g., a new fluorescent marker or resistance to a drug or antibiotic, e.g., in a split-second format) can significantly improve the efficiency of isolating clones with the intended recombination (approximately 40%, up to approximately 80%). The systems may also include additional components, reagents, or information, such as protocols for introducing the system's vectors into the cells of the system by transfection. Non-limiting examples of transfection methods include chemical transfection methods, such as the use of liposomes, nanoparticles, calcium phosphate (Graham et al. (1973) Virology 52(2): 456-67, Bacchetti et al. (1977) Proc Natl Acad Sci USA 74(4): 1590-4 and Kriegler, M (1991) Transfer and Expression: A Laboratory Manual. New York: W.H. Freeman and Company. pp. 96-97), dendrimers, or cationic polymers such as DEAE-dextran or polyethyleneimine. Non-chemical methods include electroporation, sonoporation, and optical transfection.Particle-based transfection methods include the use of gene guns and magnetically assisted transfection (Bertram, J. (2006) Current Pharmaceutical Biotechnology 7, 277-28). Viral methods can also be used for transfection. mRNA delivery includes methods using TransMessenger™ and TransIT® (Bire et al. BMC Biotechnology 2013, 13:75). One commonly used method for introducing heterologous DNA into cells is the calcium phosphate precipitation method, as described, for example, in Wigler et al. (Proc. Natl. Acad. Sci. USA 77:3567, 1980). Polyethylene-induced fusion of bacterial protoplasts with mammalian cells (Schaffner et al., (1980) Proc. Natl. Acad. Sci. USA 77:2163) is another useful method for introducing heterologous DNA. Electroporation can also be used to directly introduce DNA into the cytoplasm of host cells, as described, for example, by Potter et al. (Proc. Natl. Acad. Sci. USA 81:7161, 1988) or Shigekawa et al. (BioTechniques 6:742, 1988). Other reagents useful for introducing heterologous DNA into mammalian cells have been reported, such as Lipofectin™ reagent and Lipofectamine™ reagent (Gibco BRL, Gaithersburg, Maryland, USA). Both of these commercially available reagents are used to form lipid-nucleic acid complexes (or liposomes), which, when applied to cultured cells, facilitate the uptake of nucleic acids into the cells.
[0177] Methods for producing antigen-binding proteins The present disclosure further provides methods for producing bispecific antigen-binding proteins. Using the methods disclosed herein, desired antigen-binding proteins can be produced at high titers and / or high specific productivity (pg / cell / day). In some embodiments, the antigen-binding proteins are produced at titers of at least 1 g / L, 1.5 g / L, 2.0 g / L, 2.5 g / L, 3.0 g / L, 3.5 g / L, 4.0 g / L, 4.5 g / L, 5.0 g / L, 10 g / L or more. In some embodiments, the antigen-binding proteins are produced at specific productivity of at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 picograms / cell / day (pcd) or more, as determined based on total antigen-binding protein (pg) produced per cell per day. In some embodiments, bispecific antigen-binding proteins are produced with a ratio of bispecific antigen-binding protein titer to total antigen-binding protein titer of at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 50%, 60% or more.
[0178] In one embodiment, the method utilizes the system disclosed herein and introduces a vector of the system into cells of the system by transfection. Transfected cells that have properly integrated the exogenous nucleic acid into the two expression-enhancing loci of the cells via RMCE may be screened and identified. In some embodiments, transfected cells are identified by negative selection against one or more selectable markers that were present in the host cells prior to transfection. In other embodiments, transfected cells are identified by a combination of negative selection against one or more selectable markers that were present in the host cells prior to transfection and positive selection based on one or more selectable markers provided by the nucleic acid in the vector that they are designed to integrate. HCF-containing polypeptides and LCF-containing polypeptides can be expressed from the integrated nucleic acids, and the antigen-binding protein of interest can be obtained from the identified transfected cells and purified using known methods.
[0179] In another embodiment, the method simply utilizes a cell as described herein above, which contains exogenous nucleic acids integrated at two expression-enhanced loci, the exogenous nucleic acids encoding antigen binding proteins, and which express the antigen binding proteins from the cell, each cloned expression cassette being contiguous within each particular integration site.
[0180] The present specification is further illustrated by the following examples, which should not be construed as limiting. All cited references (including literature references, issued patents, and published patent applications cited throughout this application) are hereby expressly incorporated by reference. [Example]
[0181] Example 1: Expression (via site-specific integration) of monospecific antibodies (Abs) at two specific expression-enhanced loci The Ab chains (AbC1, AbC2) were cloned into vectors in which RSS sites flanked the Ab expression cassette and an expression cassette for a selectable marker, as illustrated in Figure 1. The two Ab chains may be cloned into separate vectors or may be combined into one vector in which the two expression cassettes are arranged in tandem in any one of the following possible orders: AbC1, AbC2, and selectable marker. For example, AbC1 is equivalent to a conventional LC and AbC2 is equivalent to a conventional heavy chain.
[0182] Briefly, DNA encoding the VH and VL domains can be isolated directly from single-antigen-positive B cells by PCR. The heavy and light chain PCR products were cloned into SapI-linearized antibody vectors containing the IgG heavy chain constant region and kappa light chain constant region, respectively. The heavy chain plasmid (AbC2) contains RRS3 and RRS2 sites flanking the heavy chain expression cassette. Additionally, the heavy chain plasmid contains a split selectable marker gene immediately downstream of RRS3 (see, e.g., US7582298). The light chain plasmid contains RRS1 and RRS3 sites flanking the light chain expression cassette. Additionally, the light chain plasmid contains a strong promoter immediately preceding the ATG at RRS3. Therefore, upon integration of the RRS3-proximal promoter into the host cell locus, the initiation ATG from the light chain plasmid is positioned adjacent to the selectable marker gene in the heavy chain plasmid in the correct reading frame, allowing transcription and translation of the selectable gene. The purified recombinant plasmid containing the heavy chain variable region sequence and the purified recombinant plasmid containing the light chain variable region sequence from the same B cell are then mixed and transfected into an engineered CHO host cell line along with a plasmid expressing a recombinase. The engineered CHO host cell line contains the appropriate RSS and selection markers at the loci of SEQ ID NO: 1 (EESYR®, locus 1) and SEQ ID NO: 2. The engineered CHO host cell line contains four different selection markers at two transcriptionally active loci. As a result, when the selection markers are different fluorescent markers, producer CHO cells can be isolated by flow cytometry for positive-negative combinations, representing the desired cell recombinants. When the recombinant plasmids expressing the heavy and light chain genes are both transfected along with a plasmid expressing a recombinase, site-specific recombination mediated by the recombinase results in the integration and replacement of the antibody plasmid at each chromosomal locus containing the RRS. Thus, recombinant cells expressing monospecific antibodies were isolated and subjected to a 12-day fed-batch culture, after which they were harvested and subjected to an Octet titer assay using immobilized protein A. The cells were confirmed to be isogenic and stable.Overall titers in small shake flasks showed increased expression of monospecific antibodies when the two-locus integration method was used, with antibody B producing a significant increase, nearly doubling the titer (Figure 2).
[0183] Example 2: Expression (via site-specific integration) of bispecific antibodies (BsAbs) at two specific expression-enhanced loci For bispecific antibody expression, three antibody chains and two selectable markers were cloned into a plasmid similar to that of Example 1, such that AbC1, AbC2, and selectable marker 1 are flanked by RRS sites compatible with the first locus or integration site (EESYR®, SEQ ID NO: 1, Locus 1), and AbC1, AbC3, and selectable marker 2 are flanked by RRS sites compatible with the second locus or integration site. Our observations suggest that AbC1, as a conventional LC, does not require two gene copies for proper expression. One or two plasmids were generated for each site. Any of the three expression cassettes were arranged in tandem, or in two plasmids. Two expression cassettes were cloned into one vector, and the remaining expression cassette was cloned into the second vector. See Figure 3.
[0184] When recombinant plasmids expressing heavy and light chain genes were co-transfected with a plasmid expressing a recombinase, site-specific recombination mediated by the recombinase resulted in the integration and replacement of the antibody plasmid at each chromosomal locus containing the RRS. Recombinant cells expressing the bispecific antibody were isolated and subjected to a 12-day fed-batch culture, after which they were harvested and subjected to an Octet titer assay using immobilized anti-Fc and a second anti-Fc* (an engineered Fc-detecting antibody; see US 2014-0134719 A1, published May 15, 2014). The cells were observed to be isogenic and stable. Overall titers in small shake flasks were significantly increased by 1.75- to more than 2-fold compared to those obtained using the two-locus integration method (Figure 5).
[0185] Example 3: Large-scale production of bispecific and monospecific antibodies after site-specific synthesis Host cells (CHO-K1) were generated analogously to Example 1 as described above (see also Figure 3 for bispecific antibodies and Figure 1 for monospecific antibodies). Host cells capable of RMCE of gene cassettes in EESYR® (Locus 1) and SEQ ID NO:2 (Locus 2) were compared to host cells capable of RMCE of gene cassettes into only one integration site (Locus 1 / EESYR®). Vectors carrying the antibody light and heavy chains (AbC1, AbC2, AbC3) and the necessary RRS and selectable marker nucleic acids (see Figure 3) were transfected into the production cell line (RSX 2BP ) to generate host cells expressing Ab E, Ab F, Ab G, and Ab H. Each bispecific antibody host cell thus expresses one common light chain and two heavy chains that bind different antigens, one of which is engineered in its CH3 domain to separately bind Protein A (as described in U.S. Pat. No. 8,586,713, incorporated herein by reference).
[0186] For monospecific antibodies, the antibody light and heavy chains (AbC1, AbC2) are cloned into a vector in which RSS sites flank the Ab expression cassette (which also provides a selectable marker gene), as illustrated in Figure 1. Recombination occurs and the host cells expressing Ab J and Ab K (RSX 2 ) was created.
[0187] The CHO-K1-derived antibody-producing cell line RSX was cultured in 2L, 15L, or 50L bioreactors. 2 A seed culture of 10 ...
[0188] Total IgG antibody (titer) was determined after Protein A / Protein G chromatography. For bispecific antibodies, total IgG antibody and each of the three antibody species, including bispecific (heterodimer Fc / Fc*), homodimer with wild-type heavy chain (Fc / Fc), and homodimer with modified heavy chain (Fc* / Fc*), were measured to determine the ratio of the desired bispecific antibody species. Total titer was determined by HPLC using a Protein A / Protein G column with the elution method described in U.S. Patent No. 8,586,713. Briefly, the three bioreactor species were bound to the column during sample loading, and the bispecific species (Fc / Fc*) was first eluted from Protein A using a pH step gradient in the presence of an ionic modifier. The bispecific species was collected during the first elution step, followed by the elution of the two homodimeric species.
[0189] Table 1 shows that using host cells expressing antibodies via two integrated loci significantly improved overall (total) IgG titers and bispecific antibody titers (Figure 6A) in pilot large-scale production cultures.
[0190] [Table 1]
[0191] Bispecific titers were determined as described above. As can be seen in Table 2, bispecific antibody titers as a percentage of total IgG titers produced by the cells are significantly higher in production cultures of host cells with two integration site structures (see Figure 6B). Indeed, it was unexpected that bispecific percentages of 50% or greater were consistently obtained.
[0192] [Table 2]
[0193] To determine the improvement in overall IgG titer, monospecific antibodies were expressed using a two-integration site method. Table 3 illustrates that a significant increase in overall titer in large-scale bioreactors was observed, from 0.6-fold to 1.3-fold. See also Figure 7. In production bioreactors used in manufacturing, particularly those with culture volumes from 500 L up to 10,000 L, increased titers were observed with the use of these improved cell lines, equating to a significant increase in production volume per batch.
[0194] [Table 3]
[0195] Although the foregoing invention has been described in some detail for purposes of illustration and example, those skilled in the art will readily understand that certain changes and modifications can be made to the teachings of this invention without departing from the spirit or scope of the appended claims. array [Table 4-1] [Table 4-2] [Table 4-3] [Table 4-4] [Table 4-5] [Table 4-6] [Table 4-7] [Table 4-8] [Table 4-9] Table 4-10 Table 4-11 Table 4-12
Claims
[Claim 1] The invention as described in the drawings.