Intein based split glutamine synthetase engineering

The split intein-based biomarker selection system in mammalian cells addresses the challenge of multiple selection pressures by reconstituting GS protein through intein splicing, ensuring efficient and balanced expression of multispecific molecules.

WO2025250814A1PCT designated stage Publication Date: 2025-12-04AMGEN INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/031460
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-31
Filing Date
2025-05-29
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing recombinant multi-chain therapeutic molecule expression in mammalian cells faces challenges due to multiple simultaneous selection pressures from antibiotic resistance and metabolite converting genes, leading to poor recombinant product quality, especially for multispecific molecules.

Method used

A split intein-based biomarker selection system using a pair of plasmids that express N-terminal and C-terminal portions of glutamine synthetase (GS) protein, allowing for the reconstitution of a full functional GS protein through intein splicing, enabling simultaneous incorporation of multiple transgenic fragments with a single selection pressure.

Benefits of technology

The system facilitates efficient and balanced cell clone screening, achieving comparable cell recovery and recombinant product quality by using a single selection pressure, thereby overcoming the limitations of traditional multi-marker systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025031460_04122025_PF_FP_ABST
    Figure US2025031460_04122025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to recombinant expression of proteins within cells, including expression of therapeutic molecules. In some aspects, the present application relates to an intein-based split biomarker selection system allows using a single selection pressure for the incorporation of multiple transgenic fragments simultaneously.
Need to check novelty before this filing date? Find Prior Art

Description

INTEIN BASED SPLIT GLUTAMINE SYNTHETASE ENGINEERINGREFERENCE TO SEQUENCE LISTING SUBMITTED ELECTRONICALLY

[0001] The present application is being filed along with a Sequence Listing in electronic format. The Sequence Listing is provided as an XML file entitled “10558-W001- SEC_SequenceListing'’, created May 7, 2025, which is 378 KB in size. The information in the electronic format of the Sequence Listing is incorporated herein by reference in its entirety.TECHNICAL FIELD

[0002] The disclosure relates to recombinant expression of proteins within cells, including expression of therapeutic molecules. In some aspects, the disclosure relates to selection and expression systems in mammalian cells.BACKGROUND

[0003] The successful and robust recombinant multi-chain therapeutic molecule expression in mammalian cells often relies on balanced incorporation and translation of all the transgenic components. The incorporations of transgenes are achieved via tying them with antibiotic resistance and / or metabolite converting genes; however, the antibiotic selection poses significant challenges in large scale manufacturing, and multiple simultaneous selection pressures often adversely impact the stability and health of the transgenic cells.

[0004] DHFR methotrexate (MTX) selection and GS methionine sulfoximine (MSX) selection are commonly -leveraged CHO transgene incorporation selection systems for CHO cells. However, encoding all transgenic components (i.e., multiple polypeptide chains) with one full biomarker often results in poor recombinant product quality7, especially for multispecific molecules.SUMMARY

[0005] This application relates to selection and expression systems in mammalian cells. Such systems are useful to facilitate efficient recombinant expression of proteins within cells, including expression of therapeutic molecules. In some aspects, the disclosure relates to use of a pair of plasmids that each contain a component of the selection system. These plasmids may contain an intein-based split biomarker selection system that allows using a single selection pressure for the incorporation of multiple transgenic fragments simultaneously. For example, the system may utilize an intein system in which two plasmids each facilitate expression of an intein component fused to a portion of a split glutamine synthetase (GS) protein, and the intein activity reconstitutes whole GS only in cells that express the fusion from both plasmids. This application relates to the systems and methods of performing such selection, and the polynucleotides that can be used in such systems and methods.

[0006] In some embodiments, this application relates to a pair of poly nucleotides, wherein the first polynucleotide encodes a first fusion of an N-terminal portion of GS and a first intern component; wherein the second polynucleotide encodes a second fusion of a C-terminal portion of GS and a second intein component; and wherein the intein components of the first and second fusions are capable of splicing the N-terminal and C- terminal portions of GS together into a spliced GS protein. Either or both of these polynucleotides may be part of distinct plasmids for expressing the first or second fusions.

[0007] In some embodiments, the spliced GS protein is a full-length GS. In some embodiments, the spliced GS protein is at least 90% identical to SEQ ID NO: 1. For example, the spliced GS protein may comprise SEQ ID NO: 1. In some embodiments, the spliced GS protein, when expressed in cells, is capable of increasing the viability of cells treated with methionine sulfoximine (MSX).

[0008] Use of a split GS system involves selection of a split site to divide the GS protein for use as part of separate intein fusions. In some embodiments, the amino acid sequence of the N-terminal portion of GS and the amino acid sequence of the C-terminal portion of GS align with the sequences of a pair of fragments of a native GS formed by splitting the native GS adjacent to a cysteine residue within the native GS sequence. In some embodiments, the amino acid sequence of the N-terminal portion of GS and the aminoacid sequence of the C-terminal portion of GS align with the sequences of a pair of fragments of a native GS formed by splitting the native GS between two non-cysteine residues within the native GS sequence. Exemplary split sites within SEQ ID NO: 1 using a native cysteine are between amino acids 52 and 53, between amino acids 98 and 99, between amino acids 1 16 and 117, between amino acids 162 and 163, between amino acids 251 and 252, or between amino acids 345 and 346. Exemplary split sites within SEQ ID NO: 1 that do not use a native split site are between amino acids 101 and 102, between amino acids 104 and 105, between amino acids 109 and 110. between amino acids 113 and 114. between ammo acids 125 and 126. or between amino acids 264 and 265. Split sites that were demonstrated to be particularly effective included splitting SEQ ID NO: 1 between amino acids 52 and 53, between amino acids 98 and 99, between amino acids 101 and 102, between amino acids 104 and 105, between amino acids 109 and 110, between amino acids 116 and 117, between amino acids 125 and 126. between amino acids 162 and 163, between amino acids 208 and 209, between amino acids 251 and 252, between amino acids 264 and 265, or between amino acids 345 and 346.

[0009] The plasmids for expressing the intein fusions also may contain polynucleotides for expressing additional polypeptide chains. For examine, in some embodiments, the first plasmid further comprises a polynucleotide encoding a first polypeptide chain, and wherein the second plasmid further comprises a polynucleotide encoding a second polypeptide chain. In some of these embodiments, the first polypeptide chain and the second polypeptide chain each are components of a multicomponent polypeptide. In some embodiments, the multicomponent polypeptide is an antibody. Thus, in some embodiments, the first polypeptide chain and the second polypeptide chain each are antibody chains or a fragment of an antibody chain. In some embodiments, the first polypeptide chain comprises an antibody heavy chain and the second polypeptide chain comprises an antibody light chain.

[0010] The pair of polynucleotides of this application may utilize a variety of intein systems. For example, in some embodiments, the intein components are part of the NpuDnaE, SspDnaE, AceL-TerL, SspDnaBmini, or SspGyrBmini intein system.

[0011] The intein and GS components of the fusions may be connected directly, or they may be connected by a linker. Therefore, in some embodiments, the N-terminal portion ofGS and / or the C-terminal portion of GS and their respective intein components are connected by a linker.

[0012] In some embodiments, the pairs of polynucleotides of this application are incorporated into methods of selecting cells expressing two polypeptide chains or methods of expressing multicomponent proteins. In some embodiments, such methods comprise introducing a pair of polynucleotides into a population of cells and treating the population of cells with methionine sulfoximine (MSX). In some embodiments, the pair of polynucleotides is introduced into the population of cells by transfection. The pair of polynucleotides may be introduced concurrently or sequentially. In some embodiments, MSX is used at a concentration of 10 pM MSX. In some embodiments, the methods utilize known cell lines or types. The cells may be CHO cells. The cells may be human cells such as HEK cells.

[0013] In some embodiments of the methods, the first polypeptide chain and the second polypeptide chain each are components of a multicomponent polypeptide. For example, in some embodiments, the first polypeptide chain and the second polypeptide chain each are antibody chains or a fragment of an antibody chain. The first polypeptide chain may comprise an antibody heavy chain and the second polypeptide chain may comprise an antibody light chain.BRIEF DESCRIPTION OF THE FIGURES

[0014] FIG. 1A shows a schematic depiction of a split intein-mediated protein splicing reaction. FIG. IB shows an exemplary representation of split intein-mediated splicing of split GS, in which the intein halves were directly fused to the GS halves without any linkers. The intrinsic interaction between the split intein halves brings the split N-terminal and C-terminal GS halves together, and by splicing intein itself out. the process completes the reconstitution of the full functional GS.

[0015] FIG. 2 depicts potential split points for the GS protein. The grey rectangle represents the full-length of the GS protein, from its N terminus to C terminus. The text above GS lists native cysteine sites that may be used as part of a split intein strategy (placed within GS by the grey lines), and the text below GS lists some potential additional non-cysteine split sites (placed within GS by the black lines).

[0016] FIG. 3 shows the results from a comparison of cell viability using different split intein strategies to reconstitute GS. A range of splice sites within GS, combined with five different intein systems, were tested for their ability to facilitate cell survival in the presence of 10 pM MSX 12 days after transfection. These data show that inteins could be tolerated at multiple sites in GS for splicing, and that the NpuDnaE and SspDnaE intein split GS variants were particularly effective. In addition, these data show that native GS cysteine residues can assist “scarless” GS reconstitution within cells.

[0017] FIG. 4 shows the results from a time course measuring cell viability, when using different split sites to reconstitute GS with the NpuDnaE intein system. The cell viability percentage and viable cell density (VCD) were counted every other day until the population reached 95% viability. The error bars represent SD from duplicated experiments. These data show a dip in viability around day 10 (i.e., seven days after MSX treatment) and recovery to near complete viability similar to control cells by days 14-18.

[0018] FIG. 5 is a Western blot of FLAG-tagged protein (or GAPDH control expression) of cell lysates 18 days after transfection of cells with plasmids for expressing split intein- GS constructs using the NpuDnaE or SspDnaE intein systems. This blot shows that the split intein systems were able to facilitate complete or near-complete reconstruction of the full GS protein in cells.DETAILED DESCRIPTION

[0019] Here, we describe “Intein based split glutamine synthetase engineering” as a broadly applicable selection system. The successful and robust recombinant multi-chain therapeutic molecule expression in mammalian cells often relies on balanced incorporation and translation of all the transgenic components. The incorporations of transgenes are achieved via tying them with antibiotic resistance and / or metabolite converting genes; however, the antibiotic selection poses significant challenges in large scale manufacturing, and multiple simultaneous selection pressures often adversely impact the stability and health of the transgenic cells. Although the DHFR methotrexate (MTX) selection and the GS methionine sulfoximine (MSX) selection are still the two most leveraged CHO transgene incorporation selection systems for CHO cells, encoding all transgenic components (multiple polypeptide chains) with one full biomarker often resulted in poor recombinant product quality', especially for multispecific molecules;therefore, using a single but split selectable marker system could be beneficial to allow faster and balanced cell clone screening with one biomarker: (1) each transgene incorporation is linked with a piece of the split biomarker; and (2) only the cells with all split biomarker components balanced and incorporated would survive the selection process.

[0020] To allow minimal disruption on the GS native sequence, we took advantage of intein trans splicing to create split GS. The intein halves were directly fused to the GS halves without any linkers. The intrinsic interaction between the intein halves bring the split N-terminal and C-terminal GS halves together, and by splicing intern itself out, the process completes the reconstitution of the full functional GS (Figure 1). When native GS cysteine residue is used at the GS split site, the reconstituted GS will be “scarless” without leaving any foreign sequences. Based on the intein action mechanism, additional GS split sites other than native cysteine residues could be included with cysteine insertion or substitution to allow expanded split GS flexibility.

[0021] Using a single but split selectable marker system could be beneficial to allow faster and balanced cell clone screening with one biomarker: (1) each transgene incorporation is linked with a piece of the split biomarker; and (2) only the cells with all split biomarker components incorporated would survive the selection process.

[0022] As an exemplary application of these methods, the Examples below demonstrate that the disclosed intein-based split biomarker selection system allows using a single selection pressure for the incorporation of multiple transgenic fragments simultaneously. More specifically, the results showed that (1) the inteins could be well tolerated at multiple sites in GS for trans splicing; (2) native GS cysteine residues could assist “scarless” GS reconstitution; and (3) the high splicing rate of the intein split GS system achieved comparable cell recovery in CHO cells to that of the full GS system.Definitions of general terms and expressions

[0023] In order that the present disclosure can be more readily understood, certain terms are first defined. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure is related. As used in this application, except as otherwise expressly provided herein, each of the following terms shall have the meaning set forth below. Additional definitions are set forth throughout the application.

[0024] Units, prefixes, and symbols are denoted in their Systeme International de Unites (SI) accepted form.

[0025] As used in the present disclosure and claims, the singular forms “a,” "an." and “the” include plural forms unless the context clearly dictates otherwise. Unless specifically stated or obvious from context, as used herein, the term “or” is understood to be inclusive. The term “and / or” as used in a phrase such as “A and / or B” herein is intended to include both “A and B,” “A or B,” “A,” and “B.” Likewise, the term “and / or” as used in a phrase such as “A. B, and / or C” is intended to encompass each of the following embodiments: A, B. and C; A. B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone); B (alone); and C (alone).

[0026] It is understood that wherever embodiments are described herein with the language “comprising,” otherwise analogous embodiments described in terms of “consisting of’ and / or “consisting essentially of’ are also provided. In this disclosure, “comprises,” “comprising,” “containing” and “having” and the like can have the meaning ascribed to them in U.S. Patent law and can mean “includes,” “including,” and the like; “consisting essentially of’ or “consists essentially” likewise have the meaning ascribed in U.S. Patent law and the terms are open-ended, allowing for the presence of more than that which is recited so long as basic or novel characteristics of that which is recited is not changed by the presence of more than that which is recited, but excludes prior art embodiments.

[0027] The terms “about” or “comprising essentially of’ refer to a value or composition that is within an acceptable error range for the value or composition as determined by one of ordinary skill in the art, which will depend in part on how the value or composition is measured or determined, i.e., the limitations of the measurement system. For example, “about” or “comprising essentially of' can mean within 1 or more than 1 standard deviation per the practice in the art. Alternatively, “about” or “comprising essentially of’ can mean a range of up to 20%. Furthermore, particularly with respect to biological systems or processes, the terms can mean up to an order of magnitude or up to 5-fold of a value. When values or compositions are provided in the application and claims, unless otherwise stated, the meaning of “about” or “comprising essentially of' should be assumed to be within an acceptable error range for that value or composition.

[0028] As used herein, the term “glutamine synthetase” refers to an enzyme that is capable of catalyzing the conversion of glutamate and ammonia into glutamine. The abbreviation “GS” is used herein to refer to glutamine synthetase. An exemplary full- length glutamine synthetase amino acid sequence (SEQ ID NO: 1; murine GS) is shown below.MATSASSHLNKGIKQMYMSLPQGEKVQAMYIWVDGTGEGLRCKTRTLDCEPKCV EELPEWNFDGSSTFQSEGSNSDMYLHPVAMFRDPFRKDPNKLVLCEVFKYNRKP AETNLRHICKRIMDMVSNQHPWFGMEQEYTLMGTDGHPFGWPSNGFPGPQGPYY CGVGADKAYGRDIVEAHYRACLYAGVKITGTNAEVMPAQWEFQIGPCEGIRMGD HLWIARFILHRVCEDFGVIATFDPKPI PGNWNGAGCHTNFSTKAMREENGLKCI EEAIDKLSKRHQYHIRAYDPKGGLDNARRLTGFHETSNINDFSAGVANRGAS IR I PRTVGQEKKGYFEDRRPSANCDPYAVTEAIVRTCLLNETGDEPFQYKN

[0029] “Methionine sulfoximine” (commonly abbreviated as MSX or MSO) is an analog of glutamate that is an inhibitor of GS. MSX can therefore be used as an agent to put selection pressure on cells based on their ability to produce glutamate. MSX is commonly added to culture media that lacks glutamate, as part of strategies to select for cells containing plasmids that facilitate GS expression.

[0030] An “intein” is a protein segment that is capable of facilitating the covalent joining (z.e., ligating or splicing via a peptide bond) of flanking protein sequences called “exteins” into a new protein, while excising the intein itself. The intein mediated joining process is referred to as “protein splicing.” An intein can operate either in cis or in trans. A cis-splicing intein splices adjacent exteins from the same polypeptide chain, while a trans-splicing intein splices exteins from two different polypeptide chains together.

[0031] “Split inteins” are a variety of trans-splicing inteins, which operate as a complementary pair of two different split inteins. A split intein pair includes an N-intein and a C-intein. The “N-intein” is the split intein that is expressed as a fusion to the C- terminus of the N-extein, which is the extein that will become the N-terminal portion of the spliced protein. The “C-intein” is the split intein that is expressed as a fusion to the N- terminus of the C-extein. which is the extein that will become the C-terminal portion of the spliced protein. A split intein pair operates only when the N-intein and the C-intein bind together to form a catalytically active enzyme. The active enzyme then catalyzes thejoining of the N-extein to the C-extein, while also joining the two inteins together. An example depiction of the split intein fusion process is depicted in FIG. 1A. A variety of different split intein pairs are known in the art, including the NpuDnaE, SspDnaE, AceL- TerL. SspDnaBmini, and SspGyrBmini systems.

[0032] The term "antibody" means an immunoglobulin molecule that recognizes and specifically binds to a target, such as a protein, polypeptide, peptide, carbohydrate, polynucleotide, lipid, or combinations of the foregoing. As used herein, the term "antibody" encompasses polyclonal antibodies, monoclonal antibodies, chimeric antibodies, humanized antibodies, fully human antibodies, recombinant antibodies, multispecific antibodies, and bispecific antibodies. The different classes of antibodies have different and well-known subunit structures and three-dimensional configurations. For example, a common configuration for an antibody has two full length antibody heavy chains and two full length antibody light chains.

[0033] As used herein, the term “antibody heavy chain” refers to an antibody heavy chain, consisting of a variable region and a constant region as defined for a full-length antibody. A full-length antibody heavy chain is a polypeptide consisting in N-terminal to C- terminal direction of an antibody heavy chain variable domain (VEI), an antibody constant heavy chain domain 1 (CHI), an antibody hinge region (HR), an antibody heavy chain constant domain 2 (CH2), and an antibody heavy chain constant domain 3 (CH3), abbreviated as VH-CH1-HR-CH2-CH3. Heavy chain amino acid sequences are known in the art.

[0034] As used herein, the term “antibody light chain” refers to an antibody light chain, consisting of a variable region and a constant region as defined for a full-length antibody. A full-length antibody light chain is a polypeptide consisting in N-terminal to C- terminal direction of an antibody light chain variable domain (VL), and an antibody light chain constant domain (CL), abbreviated as VL-CL. Light chain amino acid sequences are known in the art.

[0035] The term "antibody fragment" refers to a portion of an intact antibody. An "antigen-binding fragment," "antigen-binding domain," or "antigen-binding region," refers to a portion of an intact antibody that binds to an antigen. An antigen-binding fragment can contain the antigenic determining regions of an intact antibody (e.g., the complementarity determining regions (CDR)). Examples of antigen-binding fragments ofantibodies include, but are not limited to Fab, Fab', F(ab’)2, and Fv fragments, linear antibodies, and single chain antibodies. An antigen-binding fragment of an antibody can be derived from any animal species, such as rodents (e.g., mouse, rat, or hamster) or humans, or can be artificially produced.

[0036] The term “multispecific antibody” means that an antigen binding protein is capable of specifically binding to two or more different antigens. A subcategory of multispecific antibodies is "bispecific antibodies," which are capable of specifically binding to two different antigens. As used herein, an antibody “specifically binds” to a target antigen when it has a significantly higher binding affinity for, and consequently is capable of distinguishing, that antigen, compared to its affinity for other unrelated proteins, under similar binding assay conditions.

[0037] The terms "variable region" or "variable domain" are used interchangeably and are common in the art. The variable region typically refers to a portion of an antibody, generally, a portion of a light or heavy chain, typically about the amino-terminal 110 to 120 amino acids or 110 to 125 amino acids in the mature heavy chain and about 90 to 115 amino acids in the mature light chain, which differ extensively in sequence among antibodies and are used in the binding and specificity of a particular antibody for its particular antigen. The variability in sequence is concentrated in those regions called complementarity determining regions (CDRs) while the more highly conserved regions in the variable domain are called framework regions (FR). From N-terminus to C-terminus, naturally occurring light and heavy chain variable regions both typically conform with the following order of these elements: FR1, CDR1, FR2, CDR2, FR3, CDR3 and FR4. Without wishing to be bound by any particular mechanism or theory, it is believed that the CDRs of the light and heavy chains are primarily responsible for the interaction and specificity' of the antibody with antigen.

[0038] The terms "VL" and "VL domain" and "VH region" are used interchangeably to refer to the light chain variable region of an antibody.

[0039] The terms "VH" and "VH domain" and "VH region" are used interchangeably to refer to the heavy chain variable region of an antibody.

[0040] The terms "constant region" and "constant domain" are interchangeable and have their common meaning in the art. The constant region is an antibody portion, e.g., a carboxyl terminal portion of a light and / or heavy' chain which is not directly involved inbinding of an antibody to antigen, but which can exhibit various effector functions, such as interaction with the Fc receptor. The constant region of an immunoglobulin molecule generally has a more conserved amino acid sequence relative to an immunoglobulin variable domain.

[0041] As used herein, the terms "Fc region” and "Fc domain" refer to a C-terminal region of an IgG heavy chain; in case of an IgGl antibody, the C-terminal region comprises -CH2-CH3 (see above).

[0042] The term "chimeric antibody" refers to an antibody wherein the amino acid sequence is derived from two or more species. Typically, the variable region of both light and heavy chains corresponds to the variable region of antibodies derived from one species of mammals (e.g., mouse, rat, rabbit, etc.) with the desired specificity, affinity, and capability' while the constant regions are homologous to the sequences in derived from another (usually human) to avoid eliciting an immune response in that species.

[0043] A "humanized antibody" refers to a chimeric antibody comprising amino acid residues from non-human CDRs and amino acid residues from human framework regions and constant regions. A humanized antibody may comprise substantially all of at least one, and typically two, variable domains, in which all or substantially all of the CDRs correspond to those of a non-human antibody, and all or substantially all of the FRs correspond to those of a human antibody. A humanized antibody optionally may comprise at least a portion of an antibody constant region derived from a human antibody. A "humanized form" of an antibody, e g., a non-human antibody, refers to an antibody that has undergone humanization. Typically, humanized antibodies are human immunoglobulins in which residues from the CDRs are replaced by residues from the CDRs of a non-human species (e.g., mouse, rat, rabbit, hamster) that have the desired specificity', affinity, and capability. Accordingly, humanized antibodies are also referred to as "CDR grafted" antibodies. Early examples of methods used to generate humanized antibodies are described in U.S. Pat. 5,225,539; Roguska et al.. Proc. Natl. Acad. Sci., USA, 91(3):969-973 (1994), and Roguska et al., Protein Eng. 9(10): 895-904 (1996). Many additional examples and methods relating to humanization of antibodies have subsequently been published.

[0044] A "human antibody" refers to an antibody having vanable regions in which both the FRs and CDRs are derived from human germline immunoglobulin sequences.Furthermore, if the antibody contains a constant region, the constant region also is derived from human germline immunoglobulin sequences. The human antibodies of the disclosure can include amino acid residues not encoded by human germline immunoglobulin sequences (e.g., mutations introduced by random or site-specific mutagenesis in vitro or by somatic mutation in vivo). However, the term "human antibody," as used herein, is not intended to include antibodies in which CDR sequences derived from the germline of another mammalian species, such as a mouse, have been grafted onto human framework sequences. The terms "human antibodies" and "fully human antibodies" and are used synonymously.

[0045] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to polymers of amino acids of any length. The polymer can be linear or branched, it can comprise modified amino acids, and it can be interrupted by non-amino acids. The terms also encompass an amino acid polymer that has been modified naturally or by intervention; for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation or modification, such as conjugation with a labelling component. Also included within the definition are, for example, polypeptides containing one or more analogs of an amino acid (including, for example, unnatural amino acids), as well as other modifications known in the art.

[0046] As used herein, the term "host cell" can be any type of cell, e.g., a primary cell, a cell in culture, or a cell from a cell line. In specific embodiments, the term "host cell" refers to a cell transfected with a nucleic acid molecule and the progeny or potential progeny of such a cell. Progeny of such a cell may not be identical to the parent cell transfected with the nucleic acid molecule, e.g., due to mutations or environmental influences that may occur in succeeding generations or integration of the nucleic acid molecule into the host cell genome.Intein Systems

[0047] An "‘intein” is a protein segment that is capable of facilitating the covalent joining (z.e., ligating or splicing via a peptide bond) of flanking protein sequences called “exteins” into a new protein, while excising the intein itself. A variety of different intein systems are known in the art, many of which were found in many natural organisms such as bacteria, fungi, and plants. Information regarding intein systems in their use can befound, for example in W02014 / 004336A2 and Wang, H., et al.. Front. Bioeng. Biotechnol. 10: Art. 810180 (2022).

[0048] Some intein systems utilize “split inteins,’' which are a variety of trans-splicing inteins that operate as a complementary pair of two different split inteins. The intein components, when bound together, can catalyze the joining of their respective exteins. An example depiction of the split intein fusion process is depicted in FIG. 1A. A variety of different split intein pairs are known in the art, including the NpuDnaE, SspDnaE, AceL- TerL. SspDnaBmini, and SspGyrBmini systems.

[0049] In some embodiments, the methods and systems disclosed herein can utilize any split intein system capable of facilitating intein activity within mammalian cells. These intein components may be utilized as components of GS split intein fusion proteins, as described below.GS Split Intein Fusion Proteins

[0050] A variety of GS proteins can be used in accordance with the disclosure herein. In some embodiments, the GS protein is a mammalian GS. In some embodiments, the GS comprises a protein that has 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with SEQ ID NO: 1. In some embodiments, the GS comprises SEQ ID NO: 1.

[0051] In a split intein system, the GS is split into two pieces that are fused with the two intein components, such that the intein activity’ can join those two pieces back together. That split site may be at a position within GS that results in both smaller fragments lacking GS enzymatic activity. That split site may be at a position within GS that results in fragments of GS that both fold into stable structures independent of the other portion of GS.

[0052] Split intein pair sequences disclosed in the art may be fused with GS split fragments according to the disclosures in this application. Some exemplary split intein pairs that may be used are provided in Table 1 below, which includes intein pairs from the NpuDnaE, SspDnaE, AceL-TerL, SspDnaBmini, and SspGyrBmini systems. Each system includes an N-terminal intein sequence (“N-seq”), which is attached to the C- terminus of an N-terminal portion of GS. Each system also includes a complementary C- terminal sequence C’C-seq"). which is attached to the N-terminus of a C-terminal portion of GS.

[0053] In some embodiments, the intein and GS polypeptides are joined without a linker. In some embodiments, the intein and GS polypeptides are joined with a linker. A linker may be any peptide connecting the split GS sequence and an intein. In some embodiments, the linker is a short peptide. In some embodiments, the linker is a single residue, such as a cysteine or serine residue. In some embodiments, the linker is a linker known in the art to facilitate flexible linkage betw een polypeptides.

[0054] Some intein systems require utilize a cysteine residue at the terminus of an extein to facilitate the protein splicing reaction. For example, the NpuDnaE, SspDnaE, and AceL-TerL systems use such a cysteine. Therefore, some sites within SEQ ID NO: 1 that may be used as a split / splice site are those portions of the sequence w here a native cysteine is present. Exemplary7split / splice sites with a native cysteine are betw een amino acids K52 and C53, between amino acids L98 and C99, between amino acids 1116 and Cl 17, between amino acids Y162 and C163, between amino acids G251 and C252, or between amino acids N345 and C346.

[0055] Potential splice sites are not limited to those portions of the sequence w ith a native cysteine, even when such a cysteine is required by the intein system. In some embodiments, a cysteine is inserted into the terminus of the GS fragment's sequence at the splice site when constructing a GS -intein fusion. In some embodiments, the native residue at the terminus of the GS’s fragments sequence at the splice site is mutated to a cysteine when constructing a GS-intein fusion. Therefore, those portions of SEQ ID NO: 1 that may be used as a split / splice site are not limited to those portions of the sequence where a native cysteine is present. Exemplary split / splice sites without a cysteine are between amino acids VI 01 and Fl 02, between amino acids Y 104 and N 105, between amino acids A109 and El 10, between amino acids LI 13 and R114, between amino acids SI 25 and N126, or between amino acids E264 and N265. In some embodiments utilizing these split sites, exemplary terminal mutations include Fl 02C, N105C, El 10C, R114C, N126C, and N265C.

[0056] Some intein systems require utilize a serine residue at the terminus of an extein to facilitate the protein splicing reaction. For example, the SspDnaBmini and SspGyrBmini use such a serine. Therefore, some sites within SEQ ID NO: 1 that may be used as a split / splice site are those portions of the sequence where a native senne is present. Despite this, potential splice sites are not limited to those portions of the sequence with anative serine, even when such a serine is required by the intein system. In some embodiments, a serine is inserted into the terminus of the GS fragment’s sequence at the splice site when constructing a GS-intein fusion. In some embodiments, the native residue at the terminus of the GS’s fragments sequence at the splice site is mutated to a serine when constructing a GS-intein fusion. Therefore, those portions of SEQ ID NO: 1 that may be used as a split / splice site are not limited to those portions of the sequence where a native serine is present. Exemplary’ split / splice sites without a serine are between amino acids V101 and Fl 02, between amino acids Y104 and N105, between amino acids Al 09 and E110, between amino acids LI 13 and R114, or between amino acids E264 and N265. In some embodiments utilizing these split sites, exemplary terminal mutations include F102S, N105S, E110S, R114S, and N265S.

[0057] In some embodiments, the split intein-GS fusion pairs comprise the sequences listed in Table 2 and Table 3. In some embodiments, the split intein fusion pairs comprise pairs of sequences selected from SEQ ID NOs: SEQ ID NOs: 14 and 15, SEQ ID NOs: 14 and 16, SEQ ID NOs: 17 and 18, SEQ ID NOs: 17 and 19, SEQ ID NOs: 20 and 21, SEQ ID NOs: 20 and 22, SEQ ID NOs: 23 and 24, SEQ ID NOs: 25 and 26, SEQ ID NOs: 27 and 28, SEQ ID NOs: 29 and 30, SEQ ID NOs: 31 and 32, SEQ ID NOs: 33 and 34, SEQ ID NOs: 33 and 35, SEQ ID NOs: 36 and 37, SEQ ID NOs: 36 and 38, SEQ ID NOs: 39 and 40, SEQ ID NOs: 39 and 41, SEQ ID NOs: 42 and 43, SEQ ID NOs: 42 and 44, SEQ ID NOs: 45 and 46, SEQ ID NOs: 45 and 47, SEQ ID NOs: 48 and 49, SEQ ID NOs: 48 and 50, SEQ ID NOs: 51 and 52, SEQ ID NOs: 51 and 53, SEQ ID NOs: 54 and 56. SEQ ID NOs: 55 and 56, SEQ ID NOs: 57 and 58, SEQ ID NOs: 57 and 59, SEQ ID NOs: 60 and 61 , SEQ ID NOs: 60 and 62, SEQ ID NOs: 63 and 65, SEQ ID NOs: 64 and 65, SEQ ID NOs: 66 and 67, SEQ ID NOs: 66 and 68, SEQ ID NOs: 69 and 70, SEQ ID NOs: 71 and 72, SEQ ID NOs: 73 and 74, SEQ ID NOs: 75 and 76, or SEQ ID NOs: 77 and 78.

[0058] In some embodiments, the split intein-GS fusions may also include a tag or marker. Such markers may7be used for purification, detection, or visualization purposes. For example, a split intein-GS fusion may include a FLAG tag, a His tag, or Myc tag. Exemplary amino acid sequences of FL AG-tagged intein constructs are provided as SEQ ID NOs: 149-213.Nucleotides encoding GS Split Intein Fusion Proteins

[0059] In certain aspects, provided herein are polynucleotides comprising a nucleotide sequence encoding a split intein-GS fusion described herein. Polynucleotides provided herein can be, e.g., in the form of RNA or in the form of DNA. DNA includes cDNA, genomic DNA, and synthetic DNA, and DNA can be double-stranded or single-stranded. If single stranded, DNA can be the coding strand or non-coding (anti-sense) strand. In certain embodiments, the polynucleotides are isolated. In certain embodiments, the polynucleotides are substantially pure. In certain embodiments, a polynucleotide is purified from natural components.

[0060] The nucleic acids comprising a nucleotide sequence encoding a split intein-GS fusion can be cloned into vectors for expression in host cells. Such vectors are useful as, for example, methods in which the vectors are introduced into cells as part of selection methods. A wide range of vectors are known in the art that are suitable for mammalian expression, and may be applied for encoding and expressing the polypeptides disclosed herein.

[0061] Exemplary nucleic acids encoding split intein-GS fusions are provided as SEQ ID NOs: 79-145, which encode split intein-GS fusions with the amino acid sequences listed in SEQ ID NOs: 12-78, respectively. Similarly, exemplar}' nucleic acids encoding FLAG- tagged split intein-GS fusions are provided as SEQ ID NOs: 215-281, which encode split intein-GS fusions with the amino acid sequences listed in SEQ ID NOs: 147-213, respectively.Transgenic Expression in Mammalian Cells

[0062] A host cell, when cultured under appropriate conditions, synthesizes proteins that may be purified if so desired. For example, methods are known in the art by which host cells express and synthesize antibodies, which can subsequently be collected from the culture medium (if the host cell secretes them into the medium) or directly from the host cell producing it (if they are not secreted). In most circumstances, a population of host cells is utilized to synthesize and purify the protein or proteins at the desired scale.

[0063] The selection of an appropriate host cell will depend upon various factors, such as desired expression levels, polypeptide modifications that are desirable or necessary foractivity (such as glycosylation or phosphorylation) and ease of folding into a biologically active molecule.

[0064] Mammalian cell lines available as hosts for expression are known in the art and include, but are not limited to, immortalized cell lines available from the American Type Culture Collection (ATCC) and any cell lines used in an expression system known in the art can be used to make the recombinant polypeptides of the application. In general, host cells are transfected with a recombinant expression vector that comprises DNA encoding one or more desired polypeptides (e.g., antibody chains). Among the host cells that may be employed are eukaryotic cells, including insect cells and established cell lines of mammalian origin. Examples of suitable mammalian host cell lines include Chinese hamster ovary (CHO) cells or their derivatives such as Veggie CHO and related cell lines which grow- in serum-free media (Rasmussen et al., 1998, Cytotechnology 28: 31), the COS-7 line of monkey kidney cells (ATCC CRL 1651) (Gluzman et al., 1981, Cell 23: 175), L cells. Cl 27 cells, 3T3 cells (ATCC CCL 163), HeLa cells, BHK (ATCC CRL 10) cell lines, and the CV1 / EBNA cell line derived from the African green monkey kidney cell line CV1 (ATCC CCL 70) as described by McMahan et al., 1991, EMBO J. 10: 2821. human embryonic kidney (HEK) cells such as 293, 293 EBNA or MSR 293, human epidermal A431 cells, human Colo205 cells, other transformed primate cell lines, normal diploid cells, cell strains derived from in vitro culture of primary tissue, primary explants, HL-60, U937, HaK or Jurkat cells.

[0065] In some embodiments, mammalian cells are transfected with a pair of polynucleotides encoding a split intein-GS fusion pair. In some embodiments, the polynucleotides are in the form of a pair of plasmids that each promote expression of a split intein-GS fusion. In some embodiments, the polynucleotides are in the form of a pair of plasmids that each promote expression of both a split intein-GS fusion and an additional polypeptide.Cell Line Selection with GS

[0066] Glutamine synthetase is an enzyme that is capable of catalyzing the conversion of glutamate and ammonia into glutamine. The abbreviation “GS” is used herein to refer to glutamine synthetase. The amino acid sequences the GS enzyme is highly conserved across species, and an exemplary full-length glutamine synthetase amino acid sequence is provided as SEQ ID NO: 1 (murine GS).

[0067] GS is commonly used as a selection marker in cells that do not produce sufficient glutamine for cell grow th. Therefore, selection occurs for cells that increased GS activity, and accordingly higher levels of glutamine.

[0068] In some embodiments of this application, GS-based selection is dependent on cells containing both polynucleotides (e.g., vectors) of a split intein-GS pair. The intein reaction, which requires both polynucleotides, is necessary for the full GS and its enzymatic activity7to be present.

[0069] In some embodiments, MSX is applied to cells to increase the selection pressure for GS on those cells, as part of strategies to select for cells containing plasmids that facilitate GS expression. In some embodiments, MSX is added to culture media that lacks glutamate, as part of strategies to select for cells containing plasmids that facilitate GS expression. In some embodiments, an MSX concentration between 1 and 100 pM, between 1 and 75 pM, between 1 and 50 pM, between 1 and 25 pM, between 1 and 20 pM, between 5 and 20 pM, between 10 and 20 pM, or between 10 and 15 pM is included in culture media to apply selection pressure. In some embodiments, an MSX concentration of 1 pM, 1.5 pM, 2 pM, 2.5 pM, 3 pM, 3.5 pM, 4 pM, 4.5 pM, 5 pM, 5.5 pM, 6 pM, 6.5 pM, 7 pM, 7.5 pM, 8 pM, 8.5 pM, 9 pM, 9.5 pM, 10 pM, 10.5 pM, 11 pM, 11.5 pM, 12 pM, 12.5 pM, 13 pM, 13.5 pM, 14 pM, 14.5 pM, 15 pM, 15.5 pM, 16 pM, 16.5 pM, 17 pM, 17.5 pM, 18 pM, 18.5 pM, 19 pM, 19.5 pM, 20 pM, 25 pM, 30 pM, 35 pM, 40 pM, 45 pM, 50 pM, 55 pM, 60 pM, 65 pM, 70 pM, 75 pM, 80 pM, 85 pM, 90 pM, 95 pM, or 100 pM, is included in culture media te apply selection pressure.PARTICULAR EMBODIMENTS

[0070] Particular embodiments of the invention include the following.1. A pair of polynucleotides, wherein the first polynucleotide encodes a first fusion of an N- terminal portion of glutamine synthetase (GS) and a first intein component; wherein the second polynucleotide encodes a second fusion of a C-terminal portion of GS and a second intein component; and wherein the intein components of the first and second fusions are capable of splicing the N -terminal and C-terminal portions of GS together into a spliced GS protein.The pair of polynucleotides of embodiment 1 , wherein the first polynucleotide is part of a first plasmid for expressing the first fusion; wherein the second polynucleotide is part of a second plasmid for expressing the second fusion; or wherein both the first polynucleotide is part of a first plasmid for expressing the first fusion, and the second polynucleotide is part of a second plasmid for expressing the second fusion. The pair of polynucleotides of embodiment 1 or embodiment 2, wherein the spliced GS is a full-length GS. The pair of polynucleotides of any of embodiments 1-3, wherein the spliced GS protein is at least 90% identical to SEQ ID NO: 1. The pair of polynucleotides of embodiment 4, wherein the spliced GS comprises SEQ ID NO: 1. The pair of polynucleotides of any of embodiments 1-5. wherein the spliced GS, when expressed in cells, is capable of increasing the viability of cells treated with methionine sulfoximine (MSX). The pair of polynucleotides of any of embodiments 1-6, wherein the amino acid sequence of the N-terminal portion of GS and the amino acid sequence of the C-terminal portion of GS align with the sequences of a pair of fragments of a native GS formed by splitting the native GS adjacent to a cysteine residue within the native GS sequence. The pair of polynucleotides of any of embodiments 1-6. wherein the amino acid sequence of the N-terminal portion of GS and the amino acid sequence of the C-terminal portion of GS align with the sequences of a pair of fragments of a native GS formed by splitting the native GS between two non-cysteine residues within the native GS sequence. The pair of polynucleotides of any of embodiments 1-6, wherein the amino acid sequence of the N-terminal portion of GS and the amino acid sequence of the C-terminal portion of GS align with the sequences of a pair of fragments formed by splitting SEQ ID NO: 1 between amino acids 52 and 53, between amino acids 98 and 99, between amino acids116 and 117, between amino acids 162 and 163, between amino acids 251 and 252, or between amino acids 345 and 346.The pair of polynucleotides of any of embodiments 1 -6, wherein the amino acid sequence of the N-terminal portion of GS and the amino acid sequence of the C-terminal portion of GS align with the sequences of a pair of fragments formed by splitting SEQ ID NO: 1 between amino acids 101 and 102. between amino acids 104 and 105. between amino acids 109 and 1 10, between amino acids 1 13 and 114, between amino acids 125 and 126, or between amino acids 264 and 265. The pair of polynucleotides of any of embodiments 1-6, wherein the amino acid sequence of the N-terminal portion of GS and the amino acid sequence of the C-terminal portion of GS align with the sequences of a pair of fragments formed by splitting SEQ ID NO: 1 between amino acids 52 and 53, between amino acids 98 and 99, between amino acids 101 and 102, between amino acids 104 and 105, between amino acids 109 and 110, between amino acids 116 and 117. between amino acids 125 and 126. between amino acids 162 and 163, between amino acids 208 and 209, between amino acids 251 and 252, between amino acids 264 and 265, or between amino acids 345 and 346. The pair of polynucleotides of embodiment 2, wherein the first plasmid further comprises a polynucleotide encoding a first polypeptide chain, and wherein the second plasmid further comprises a polynucleotide encoding a second polypeptide chain. The pair of polynucleotides of embodiment 12, wherein the first polypeptide chain and the second polypeptide chain each are components of a multicomponent polypeptide. The pair of polynucleotides of embodiment 13, wherein the first polypeptide chain and the second polypeptide chain each are antibody chains or a fragment of an antibody chain. The pair of polynucleotides of embodiment 14, wherein the first polypeptide chain comprises an antibody heavy chain and the second polypeptide chain comprises an antibody light chain. The pair of polynucleotides of any of embodiments 1-15, wherein the intein components are part of the NpuDnaE, SspDnaE, AceL-TerL, SspDnaBmini, or SspGyrBmini intein system.The pair of polynucleotides of embodiment 16, wherein the intein components are part of the NpuDnaE or SspDnaE intein system. The pair of polynucleotides of any of embodiments 1-17, wherein the N-terminal portion of GS and the first intein are connected by a linker. The pair of polynucleotides of any of embodiments 1-17, wherein the C-terminal portion of GS and the second intein are connected by a linker. The pair of polynucleotides of any of embodiments 1-17, wherein the N-terminal portion of GS and the first intein are connected by a linker, and wherein the C-terminal portion of GS and the second intein are connected by a linker. A method of selecting for cells expressing two polypeptide chains comprising: a. introducing the pair of polynucleotides of embodiment 1 into a population of cells; and b. treating the population of cells with methionine sulfoximine (MSX). The method of embodiment 21, wherein the pair of polynucleotides is introduced into the population of cells by transfection. The method of embodiment 21 or embodiment 22, wherein each of the pair of polynucleotides is introduced concurrently. The method of embodiment 21 or embodiment 22, wherein each of the pair of polynucleotides is introduced sequentially. The method of any of embodiments 21-24, wherein the population of cells is treated with 10 pM MSX. The method of any of embodiments 21-25, wherein the population of cells comprises CHO cells. The method of any of embodiments 21-25, wherein the population of cells comprises human cells. The method of embodiment 27, wherein the population of cells comprises HEK cells.A method of expressing multicomponent proteins comprising: a. introducing the pair of polynucleotides of any of embodiments 12-15 into a population of cells; and b. treating the population of cells wi th methionine sulfoximine (MSX). The method of embodiment 29, wherein the pair of polynucleotides is introduced into the population of cells by transfection. The method of embodiment 29 or embodiment 30, wherein each of the pair of polynucleotides is introduced concurrently. The method of embodiment 29 or embodiment 30, wherein each of the pair of polynucleotides is introduced sequentially. The method of any of embodiments 29-32, wherein the population of cells is treated with 10 pM MSX. The method of any of embodiments 29-33, wherein the population of cells comprises CHO cells. The method of any of embodiments 29-33, wherein the population of cells comprises human cells. The method of embodiment 35, wherein the population of cells comprises HEK cells. The method of any of embodiments 29-36, wherein the first polypeptide chain and the second polypeptide chain each are components of a multicomponent polypeptide. The method of embodiment 37, wherein the first polypeptide chain and the second polypeptide chain each are antibody chains or a fragment of an antibody chain. The method of embodiment 38, wherein the first polypeptide chain comprises an antibody heavy chain and the second polypeptide chain comprises an antibody light chain.EXAMPLESExample 1: Glutamine synthetase (GS) reconstitution by intein trans splicing

[0071] To demonstrate the broad ability of intein trans splicing to reconstitute split GS sequences, split GS fusions were designed using five different intein systems: NpuDnaE. SspDnaE, AceL-TerL, SspDnaBmini, and SspGyrBmini. Table 1, shown below, lists the sequences of the intein pairs used as part of the tests disclosed in this Example. Each system's N-terminal intein sequence (“N-seq”; to be attached to the C-terminus of an N- terminal portion of GS) was used in conjunction with its complementary C-terminal sequence (“C-seq”: to be attached to the N-terminus of a C-terminal portion of GS).TABLE 1

[0072] For screening purposes, a series of plasmids were constructed using the inteins systems in Table 1, with each plasmid containing a single cassette for expressing an intein fused with a fragment of GS. The base vector for these plasmids was the mammalian transient expression vector pBMV-SRa. Nucleotides synthesized by Twist Bioscience were assembled into plasmids using the Golden Gate assembly method. (See Engler C et al., PLoS One 3:e3647 (2008), which is incorporated herein by reference in its entirety) After sequencing confirmation by Sanger, transfection-grade DNA was prepared using Maxi plasmid purification kits (Qiagen).

[0073] V arious pairs of plasmids were designed such that, for any given pair, one plasmid expresses a first intein-GS fusion that interacts with a second intein-GS fusion expressed by a second plasmid, and the fusions’ interaction is via their intein halves. When the intein fusions are expressed within the same cell, the intrinsic interaction between the intein halves (and subsequent trans splicing) brings the split N-terminal and C-terminal GS halves together. By splicing intein itself out, the process completes the reconstitution of the full functional GS (see FIG. IB). For some plasmid pairs, native GS cysteine residues were used at the GS split site, allowing the reconstituted GS to be "‘scarless” (i.e., without any foreign sequences in the reconstituted GS). For other plasmid pairs, as an active site cysteine residue is required for robust intein splicing, a cysteine insertion / substitution was made at the GS split site. Testing the range of split sites allowed testing of the flexibility of the split GS approach. FIG. 2 shows some potential split points for the GS protein, both those utilizing native cysteine (shown above the protein) and those without a native cysteine (shown below the protein).

[0074] The sequences of the tested intein-GS fusions are provided below in Table 2. In total, a panel of 42 pairs of intein split GS constructs were generated, including a full GS and a split GS control. The split GS control was a leucine zipper-based split GS pair of plasmids, in which a GCN4 leucine zipper drives the non-covalent assembly of GS haves in both homodimeric and heterodimeric formats. The full GS sequence (SEQ ID NO: 1; Mus musculus) was expressed in the same base plasmid. Additional constructs weregenerated that included the FLAG tag conjugated to GS, the intein-GS fusions, or leucine zipper-GS fusions, and their sequences are provide as SEQ ID NOs: 146-213 (not shown in Table 2). Polynucleotides encoding the amino acids of SEQ ID NOs: 12-78 are listed in SEQ ID NOs: 79-145, respectively. Polynucleotides encoding the amino acids of SEQ ID NOs: 147-213 are listed in SEQ ID NOs: 215-281, respectively.

[0075] In each intein construct in Table 2 below, the intein sequence is underlined and bolded and the GS sequence is in plain text.T BLE 2

[0076] The intein sequences in Table 2 were used in pairs, as described in Table 3 below.TABLE 3

[0077] CHO cells that have been genetically engineered to lack the GS gene (z.e. , a GS knockout cell line) were used for testing the functionality of split intein-GS fusion pairs. The CHO cells were cultured and transfected with these split GS pairs, and the selection process was initiated by adding 10 pM methionine sulfoximine (MSX) to the culture media starting at Day 3 post-transfection. The cells were maintained via fresh selection media exchange and cell splitting every two days to avoid overgrown cells. The cell viability percentage and viable cell density (VCD) were counted at Day 12 posttransfection (z.e., nine days after addition of MSX). Trypan blue staining was used to distinguish viable and non-viable cells, and the VCD and cell viability measurements were obtained using a Vi-Cell XR viabi li ty analyzer (Beckman Coulter) following the instrument manual’s protocol.

[0078] As shown in FIG. 3, on day 12 analysis of the full GS positive control (vector comprising SEQ ID NO: 1) showed cells with greater than 90% viability and a cell density of approximately 8x106cells / mL. The zinc finger split GS construct pair (GS split between Y 104 and N 105, fused with the GCN4 Leucine Zipper) allowed for greater than 70% viability and a density just under IxlO6cells / mL. With respect to the intein constructs, this screen revealed that many NpuDnaE and SspDnaE split GS pairs showed similar cell recovery (% viability ) to the full-length GS and better recovery than the split GS control. Most of the NpuDnaE and SspDnaE split GS pairs also showed greater cell density than the split GS control. The inteins in these systems were consistently tolerated at the range of different sites in GS for splicing, including using native GS cysteine residues to facilitate scarless GS reconstitution (z.e., reconstitution of GS with no sequence deviation from the parent GS at the splice site).

[0079] The AceL-TerL, SspDnaBmini, and SspGyrBmini split GS pairs were significantly less effective at facilitating cell survival and grow th at day 12 (see FIG. 3). While populations of viable cells w ere noted, their cell viability percentage did not rise to the same level as those obtained with either the controls or the NpuDnaE and SspDnaE split GS pairs. Also, the VCD of most of the AceL-TerL, SspDnaBmini, and SspGyrBmini split GS pairs remained very low at day 12.

[0080] To gain an understanding of transfected cell status over time, the cell VCD and viability’ percentage were monitored continuously after transfection and MSX selection on 8 NpuDnaE and 2 SspDnaE intein split GS pairs (see FIG. 4). Cells were treating with MTX 3 days after transfection. The intein pairs tested were Y 104_N 105 NpuDnaE (SEQ ID NOs: 14 and 15), E264_N265 NpuDnaE (SEQ ID NOs: 20 and 21), K52_C53 NpuDnaE (SEQ ID NOs: 23 and 24), L98_C99 NpuDnaE (SEQ ID NOs: 69 and 70), I116 C117 NpuDnaE (SEQ ID NOs: 77 and 78), G251_C252 NpuDnaE (SEQ ID NOs: 27 and 28), A109_E110C NpuDnaE (SEQ ID NOs: 73 and 74), S 125_N126C NpuDnaE (SEQ ID NOs: 17 and 19), S125_N126C SspDnaE (SEQ ID NOs: 36 and 37), and E264_N265 SspDnaE (SEQ ID NOs: 39 and 40). These intein pairs w ere compared to the controls of transfecting (1) a plasmid expressing full GS (SEQ ID NO: 1) and (2) a zinc finger split GS construct pair (GS split between Y104 and N105, fused with the GCN4 Leucine Zipper; SEQ ID Nos: 12 and 13).

[0081] The cells were maintained via fresh selection media exchange and cell splitting based on the VCD counts every two days to avoid overgrown cells. As seen in FIG. 4, the cell viability percentage of the split GS control dropped below 80% at Day 10 before recovering to approximately 95% by Day 18. Although the selected intein split GS constructs similarly experienced a cell viability’ percentage decrease around Day 10, the percentage drop was milder to the level betw een 80 and 90%, and most of them could recover above 95% before Day 18. Among these 10 selected intein split GS pairs, the GS E264_N265 (with cysteine insertion) NpuDnaE split pair appeared to have the best performance in that the cells w ere able to recover above 95% viability by’ Day 12. In addition, the native GS cysteine split sites, K52_C53, Il 16 C117, and G251_C252, achieved above 95% viability’ by Day 14 with the NpuDnaE intein.

[0082] To confirm the GS reconstitution by intein trans splicing, the FLAG tag containing intein split GS series were cloned into plasmids and tested in parallel. Theamino acid sequence of the FLAG-tagged intein constructs used for this experiment were SEQ ID NOs: 149-213. Other than the addition of the FLAG tag, the proteins of SEQ ID NOs: 149-213 are the same as those in SEQ ID NOs: 14-78. As controls for experiments using FLAG-tagged intein constructs, a FLAG-GS (SEQ ID NO: 146) and FLAG-tagged zinc finger-split GS (SEQ ID NOs: 147 and 148) sequences were used.

[0083] The GS N- and C-termini are relatively flexible in published GS structures, which suggests that fusing FLAG-tag may not significantly impact the GS, therefore, the two FLAG-tag containing intein split GS halves were constructed as FLAG-N-GS-N-intein and C-intein-C-GS-FLAG. The cells expressing FLAG tag-containing intein split GS were harvested at Day 18 and lysed for assessing the GS reconstitution by Western Blot. The FLAG tag containing split GS in general showed slight decreased cell recovery but still held the same trend compared to the tag-free version (data not shown). FIG.5 shows the results of representative (see FIG. 5). In FIG. 5, lanes 1-12 are lysates from the intein split GS constructs: (1) Full GS control plasmid, (2) Zinc finger split GS control, (3) Y104_N105 NpuDnaE, (4) E264_N265 NpuDnaE, (5) S125_N126C NpuDnaE, (6) K52 C53 NpuDnaE, (7) G251_C252 NpuDnaE, (8) E264_N265 SspDnaE, (9) S125_N126C SspDnaE, (10) L98_C99 NpuDnaE. (11) A1O9_E110C NpuDnaE, and (12) Il 16 C117 NpuDnaE. As expected, the zinc finger-based split GS control showed separate bands for the two expressed proteins because zinc finger systems promote only non-covalent interactions that are disrupted in SDS PAGE gel analysis. In contrast, all of the tested intein split GS constructs had full GS reconstituted by Day 18 with minimal or undetectable unsplit halves. This observation suggested high intein splicing efficiency to reconstitute full GS. This result is consistent with the ability’ of these constructs to enable quick cell recovery by Day 18 (as shown in FIG. 4 and discussed above).

[0084] These data show that an intein-based split biomarker selection system allows using a single selection pressure for the incorporation of multiple transgenic fragments simultaneously. Such approach enables lower cost and faster timeline for identifying clones with high productivity. More specifically, the results showed that (1) the inteins could be well tolerated at multiple sites in GS for trans splicing; (2) native GS cysteine residues could assist “scarless” GS reconstitution; and (3) the high splicing rate of the intein split GS system achieved comparable cell recovery in CHO cells to that of the full GS system.

Claims

CLAIMS1. A pair of polynucleotides, wherein the first polynucleotide encodes a first fusion of an N- terminal portion of glutamine synthetase (GS) and a first intein component; wherein the second polynucleotide encodes a second fusion of a C-terminal portion of GS and a second intein component; and wherein the intein components of the first and second fusions are capable of splicing the N-terminal and C-terminal portions of GS together into a spliced GS protein.

2. The pair of polynucleotides of claim 1, wherein the first polynucleotide is part of a first plasmid for expressing the first fusion; wherein the second polynucleotide is part of a second plasmid for expressing the second fusion; or wherein both the first polynucleotide is part of a first plasmid for expressing the first fusion and the second polynucleotide is part of a second plasmid for expressing the second fusion.

3. The pair of polynucleotides of claim 1, wherein the spliced GS protein is a full-length GS.

4. The pair of polynucleotides of claim 1, wherein the spliced GS protein is at least 90% identical to SEQ ID NO: 1.

5. The pair of polynucleotides of claim 4, wherein the spliced GS protein comprises SEQ ID NO: 1.

6. The pair of polynucleotides of claim 1, wherein the spliced GS protein, when expressed in cells, is capable of increasing the viability of cells treated with methionine sulfoximine (MSX).

7. The pair of polynucleotides of claim 1, wherein the amino acid sequence of the N- terminal portion of GS and the amino acid sequence of the C-terminal portion of GS align with the sequences of a pair of fragments of a native GS formed by splitting the native GS adjacent to a cysteine residue within the native GS sequence.

8. The pair of polynucleotides of claim 1, wherein the amino acid sequence of the N- terminal portion of GS and the amino acid sequence of the C-terminal portion of GS alignwith the sequences of a pair of fragments of a native GS formed by splitting the native GS between two non-cysteine residues within the native GS sequence.

9. The pair of polynucleotides of claim 1, wherein the amino acid sequence of the N- terminal portion of GS and the amino acid sequence of the C-terminal portion of GS align with the sequences of a pair of fragments formed by splitting SEQ ID NO: 1 between amino acids 52 and 53, between amino acids 98 and 99, between amino acids 116 and117, between amino acids 162 and 163, between amino acids 251 and 252, or between amino acids 345 and 346.

10. The pair of polynucleotides of claim 1, wherein the amino acid sequence of the N- terminal portion of GS and the amino acid sequence of the C-terminal portion of GS align with the sequences of a pair of fragments formed by splitting SEQ ID NO: 1 between amino acids 101 and 102, between amino acids 104 and 105, between amino acids 109 and 110, between amino acids 113 and 114, between amino acids 125 and 126, or between amino acids 264 and 265.1 1. The pair of polynucleotides of claim 1, wherein the amino acid sequence of the N- terminal portion of GS and the amino acid sequence of the C-terminal portion of GS align with the sequences of a pair of fragments formed by splitting SEQ ID NO: 1 between amino acids 52 and 53. between amino acids 98 and 99. between amino acids 101 and 102, between amino acids 104 and 105, between amino acids 109 and 1 10, between amino acids 116 and 117, between amino acids 125 and 126, between amino acids 162 and 163, between amino acids 208 and 209, between amino acids 251 and 252, between amino acids 264 and 265, or between amino acids 345 and 346.

12. The pair of polynucleotides of claim 2, wherein the first plasmid further comprises a polynucleotide encoding a first polypeptide chain, and wherein the second plasmid further comprises a polynucleotide encoding a second polypeptide chain.

13. The pair of polynucleotides of claim 12, wherein the first polypeptide chain and the second polypeptide chain each are components of a multicomponent polypeptide.

14. The pair of polynucleotides of claim 13, wherein the first polypeptide chain and the second polypeptide chain each are antibody chains or a fragment of an antibody chain.

15. The pair of polynucleotides of claim 14, wherein the first polypeptide chain comprises an antibody heavy chain and the second polypeptide chain comprises an antibody light chain.

16. The pair of polynucleotides of claim 1, wherein the intein components are part of the NpuDnaE, SspDnaE, AceL-TerL, SspDnaBmini, or SspGyrBmini intein system.

17. The pair of polynucleotides of claim 16, wherein the intein components are part of the NpuDnaE or SspDnaE intein system.

18. The pair of polynucleotides of claim 1 , wherein the N-terminal portion of GS and the first intein are connected by a linker.

19. The pair of polynucleotides of claim 1, wherein the C -terminal portion of GS and the second intein are connected by a linker.

20. The pair of polynucleotides of claim 1 , wherein the N-terminal portion of GS and the first intein are connected by a linker, and wherein the C -terminal portion of GS and the second intein are connected by a linker.

21. A method of selecting for cells expressing two polypeptide chains comprising: a. introducing the pair of polynucleotides of claim 1 into a population of cells; and b. treating the population of cells with methionine sulfoximine (MSX).

22. The method of claim 21, wherein the pair of polynucleotides is introduced into the population of cells by transfection.

23. The method of claim 21, wherein each of the pair of polynucleotides is introduced concurrently.

24. The method of claim 21, wherein each of the pair of polynucleotides is introduced sequentially.

25. The method of claim 21, wherein the population of cells is treated with 10 pM MSX.

26. The method of claim 21, wherein the population of cells comprises CHO cells.

27. The method of claim 21, wherein the population of cells comprises human cells.

28. The method of claim 27, wherein the population of cells comprises HEK cells.

29. A method of expressing multicomponent proteins comprising: a. introducing the pair of polynucleotides of claim 12 into a population of cells; and b. treating the population of cells with methionine sulfoximine (MSX).

30. The method of claim 29, wherein the pair of polynucleotides is introduced into the population of cells by transfection.

31. The method of claim 29, wherein each of the pair of polynucleotides is introduced concurrently.

32. The method of claim 29. wherein each of the pair of polynucleotides is introduced sequentially.

33. The method of claim 29, wherein the population of cells is treated with 10 pM MSX.

34. The method of claim 29, wherein the population of cells comprises CHO cells.

35. The method of claim 29, wherein the population of cells comprises human cells.

36. The method of claim 35, wherein the population of cells comprises HEK cells.

37. The method of claim 29, wherein the first polypeptide chain and the second polypeptide chain each are components of a multicomponent polypeptide.

38. The method of claim 37, wherein the first polypeptide chain and the second polypeptide chain each are antibody chains or a fragment of an antibody chain.

39. The method of claim 38, wherein the first polypeptide chain comprises an antibody heavy chain and the second polypeptide chain comprises an antibody light chain.

Citation Information

Patent Citations

  • Recombinant altered antibodies and methods of making altered antibodies

    US5225539A

  • Split inteins, conjugates and uses thereof

    WO2014004336A2

  • Direct selection of cells expressing high levels of heteromeric proteins using glutamine synthetase intragenic complementation vectors

    WO2017197098A1

  • Auxotrophic cells for virus production and compositions and methods of making

    WO2022187546A1

  • Method for selection of cell line expressing high levels of recombinant protein using glutamine synthetase split expression vector

    WO2025037708A1